FR version is available. Content is displayed in original English for accuracy.
Benchmarks how well Harness+models can create a Pac-Man game from a single prompt:
“Create a Pac-Man game in a single HTML page”
Each model gets one shot — no follow-up prompts or fixes.

Discussion (24 Comments)Read Original on HackerNews
The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.
It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.
They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.
That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.
Successful ecosystems are just pedantically lame enough to keep silly folks from going YOLO, but empowering enough to allow people to still have fun.
>which model has the best training code that was Pac-Man.
You mean which model is more cautious about copyright bleed-through of the $9Tn in FOSS and user code they misappropriated though isomorphic plagiarism. =3
Nodejs wasn't easier for people. JS is a horrible language where simple stuff like comparisons, array access, and member access are broken.
I personally have been working on a Three.js project with Opus 5 and 5.5 that I never would have continued with had I needed to dive into documentation by hand.
Seeing immediate results is incredibly motivating.
Remember when people considered you a genius for prompting with "You are a skilled writer....".
https://news.ycombinator.com/item?id=49882889
These results get better and better, but the half-baked pac mans games and pelicans are just funny.
You cannot prompt that level of brokenness. Same for those slop-posters you see everywhere.
It has the advantage of taking good screenshots and letting me know if it’s decent within five seconds of playing.
Buffered controls, different pathfinding AI for each ghost, level transitions, etc.
Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.