DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
94% Positive
Analyzed from 821 words in the discussion.
Trending Topics
#models#jaw#gemini#svg#flash#opus#https#benchmark#habsburg#frog

Discussion (40 Comments)Read Original on HackerNews
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
That looks like something from Machinarium or Robots :)
https://playcode.io/blog/macbook-svg-benchmark
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
raninemandibularprognathism-maxxed?
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
That's a pretty good benchmark
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
Would've wanted to see also DS4 flash.
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.