Back to News
Advertisement
ttkgally about 3 hours ago 12 commentsRead Article on gally.net

DE version is available. Content is displayed in original English for accuracy.

In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].

Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.

Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

[1] https://simonwillison.net/2025/Nov/25/

Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 284 words in the discussion.

Trending Topics

#models#benchmark#octopus#gemini#dorrit#years#seems#general#top#still

Discussion (12 Comments)Read Original on HackerNews

svcrunch31 minutes ago
I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

1. It tests visual reasoning and structured output in a single task.

2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

[1] https://dorrit.pairsys.ai/

vova_hn2about 1 hour ago
Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

BrokenCogs12 minutes ago
Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?
GaggiX8 minutes ago
Gemini 3.8 Flash results are often not very coherent but it does put a lot of shading and details to hide the fact.
neilellis13 minutes ago
Well that benchmark is now saturated, what next. How fast you can hack the pentagon?
samayasharabout 1 hour ago
All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

eddytrex_21 minutes ago
Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?
steinvakt2about 2 hours ago
Feels like google has a different training set than the others?
sajithdilshanabout 1 hour ago
Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others
qiine38 minutes ago
Asking to animate it add an interesting layer of difficulty
sceptic123about 1 hour ago
> An elephant typing on a typewriter

A monkey, surely?