DE version is available. Content is displayed in original English for accuracy.
In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].
Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.
Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

Discussion (12 Comments)Read Original on HackerNews
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
[1] https://dorrit.pairsys.ai/
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
[0] https://en.wikipedia.org/wiki/Goodhart%27s_law
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
A monkey, surely?