HI version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
95% Positive
Analyzed from 955 words in the discussion.
Trending Topics
#models#jaw#gemini#svg#flash#human#opus#https#benchmark#habsburg

Discussion (41 Comments)Read Original on HackerNews
Still does generate odd looking frogs too - even added a few creative flourishes, based entirely on assumptions, which is very important to note whenever interacting with AI - necessary even if one is going to use them for actual work.
AI, as is, remains one of the best tools humans have ever created but it's really not human and is always delivering to us what it thinks we want as the primary goal in every interaction.
Even if incapable of delivering what was asked of it - it will still answer, confidently, regardless of whether or not the answer is correct, heavily biased or humorously wrong. This is a valuable lesson in a simple, and hopefully self-evident, package.
Seriously, very well done - I will occasionally be sending people to your site for the foreseeable future, I appreciate your work!
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
That looks like something from Machinarium or Robots :)
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
raninemandibularprognathism-maxxed?
https://playcode.io/blog/macbook-svg-benchmark
That's a pretty good benchmark
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
Would've wanted to see also DS4 flash.
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.