ZH version is available. Content is displayed in original English for accuracy.
In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].
Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.
Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

Discussion (45 Comments)Read Original on HackerNews
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
[1] https://dorrit.pairsys.ai/
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
[0] https://en.wikipedia.org/wiki/Goodhart%27s_law
This has been the plan since the start of all this, they regurgitate code in ever better forms but they still aren’t inventing new things yet.
For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
That's pretty generous.
Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.
More-generally, I suspect an influence from how left-to-right languages (i.e. English) affect comic layouts. Overcoming that bias often means using vertical space to exploit the top-to-bottom habit instead. (Consider the rarity of an English-language comic panel where action is from bottom-right to top-left.)
MacBook Pro 3D in SVG for me the most helpful one.
https://youtu.be/w0qDV2QhAlg?si=JyHS6rZLmIGAuTd3&t=46
A monkey, surely?
Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.
The only other one I've spotted is animated is Gemini 3.0's 2025 run of an elephant.
Qwen3.8 is very clearly distilled from Claude models.