Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

95% Positive

Analyzed from 955 words in the discussion.

Trending Topics

#models#jaw#gemini#svg#flash#human#opus#https#benchmark#habsburg

Discussion (41 Comments)Read Original on HackerNews

NemoNobody10 minutes ago
This is excellent - some of the best research I've seen yet regarding AI, demonstrating immediately what it can't do and what it does do - both are unarguably problematic and have nothing to do with it being self-aware, skynet or the end of all of our jobs.

Still does generate odd looking frogs too - even added a few creative flourishes, based entirely on assumptions, which is very important to note whenever interacting with AI - necessary even if one is going to use them for actual work.

AI, as is, remains one of the best tools humans have ever created but it's really not human and is always delivering to us what it thinks we want as the primary goal in every interaction.

Even if incapable of delivering what was asked of it - it will still answer, confidently, regardless of whether or not the answer is correct, heavily biased or humorously wrong. This is a valuable lesson in a simple, and hopefully self-evident, package.

Seriously, very well done - I will occasionally be sending people to your site for the foreseeable future, I appreciate your work!

hn_throwaway_99about 3 hours ago
I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense.

Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.

viciousvoxelabout 2 hours ago
Perhaps you're thinking of Salad Fingers?
qwertybasedabout 1 hour ago
~I'm leaning more towards the rage face poker face~

edit: nevermind, definitely "monkey-puppet side-eye" vibe.

thebigshipabout 2 hours ago
Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/

also my favorite SVG was def the google/gemini-3.6-flash

edit: ok better now I think

troupoabout 2 hours ago
> also my favorite SVG was def the google/gemini-3.6-flash

That looks like something from Machinarium or Robots :)

riazrizvi22 minutes ago
The secret to great interview questions and challenge tests is keeping them secret. Posting them on HN and getting them onto the front page puts them in jeopardy.
wren6991about 2 hours ago
Opus 5 clearly frogmaxxed.

gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

fasterikabout 2 hours ago
Arguably, a royal portrait is a misinterpretation of the prompt, since it's just asking for a specific facial feature. But I guess you could look at it as a bit of artistic license.
andybakabout 1 hour ago
> frogmaxxed

raninemandibularprognathism-maxxed?

ianberdinabout 1 hour ago
Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code.

https://playcode.io/blog/macbook-svg-benchmark

evan_about 2 hours ago
Hopsburg Jaw
ricardobeatabout 1 hour ago
Gemini 2.5 Pro fails, but has a distinctive art style that is quite nice. It seems to understand shading to a much higher level than all other models.
k1e25 minutes ago
No gpt-5.6 sol and no fable?
rush86999about 2 hours ago
It's opus 5 > Kimi K3 > grok 4.5

That's a pretty good benchmark

getnormalityabout 3 hours ago
This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.
Advertisement
gerdesjabout 1 hour ago
Please could we have a human generated image to compare the AI generated tosh with?
throwuxiytayq37 minutes ago
My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.
dehrmannabout 2 hours ago
How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.
linksnapzzabout 2 hours ago
A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.
esaymabout 2 hours ago
stavrosabout 1 hour ago
I'm assuming they mean an SVG.
MiroslavPokornyabout 1 hour ago
My test is to ask AI to pick up all the rubbish at the beach.
csomar25 minutes ago
Here is GLM 5.2 (https://codeinput.com/s/HAO0qTxw2ia) which is still inferior to Opus. I can't find Qwen 3.8 which now is my daily driver replacing GLM. This SVG test matches my experience when working with the different models. The other models can get the details right but their output is structured in a way that makes little or no sense.

I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.

leumonabout 2 hours ago
Can you also try the new deepseek v4 flash?
thebigshipabout 2 hours ago
I will add it to next month's report!
leumonabout 1 hour ago
Thank you!
epolanski43 minutes ago
I don't get the point of these benchmarks, what are they supposed to represent practically?
runarberg39 minutes ago
For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.

Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.

andy9919 minutes ago
No
viccis39 minutes ago
The ability of an LLM to produce something not in its training data set.
epolanskiabout 1 hour ago
Gemini 3.6 flash is crazy good.

Would've wanted to see also DS4 flash.

gpvosabout 1 hour ago
Crazy funny, yes, but not good.
epolanski42 minutes ago
I think that on the rendering side, it's miles ahead of the rest, even if off topic.
troupoabout 2 hours ago
Mine is any variations on mammoths in various situations, or anthropomorphic. Since mammoths are invariably majestically going from one place to another in any of the books, models have hard time imagining anything but that.

Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)

Advertisement
thebigshipabout 3 hours ago
I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.

Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

Mistral returned byte-identical output across separate calls.

Gemini narrates its work in 65 comments; Llama says nothing.

If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

n00bskoolbusabout 2 hours ago
The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.
fwipabout 1 hour ago
I wonder if the setup accidentally hit a cache at some layer.
HPsquaredabout 2 hours ago
Bite-identical?
thebigshipabout 2 hours ago
clearly I missed an amazing copy opportunity, thank you haha
sixtyjabout 3 hours ago
For those who don’t know a Habsburg jaw also known as mandibular prognathism, it is a genetic condition characterized by a protruding lower jaw, which was notably prevalent among members of the Habsburg royal family due to their history of inbreeding. This condition often resulted in significant facial deformities and difficulties with eating and speaking.