DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
75% Positive
Analyzed from 3491 words in the discussion.
Trending Topics
#models#benchmark#model#writing#tasks#don#fable#glm#generated#article

Discussion (89 Comments)Read Original on HackerNews
I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.
"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...
This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.
I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune
Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent
Look at the later comments, they have substance oriented discussions.
HN Mods - can you please consider a policy against such comments, it's overflowing the site and is diluting discourse and value. If people don't like an article, they can simply ignore it. These articles are reaching the top because enough people consider it of value.
But when an article comes to the top, and the topmost comment and discussion thread is an unqualified witch hunt, it's getting sick.
Can we assume everything we think as AI - must have had a high-quality human pattern behind it, and there is no way to 100% prove which is which - unless the author shows a screencast of them typing the artice?
This is not healthy. The right thing to do is – if someone doesn't like an article, they should ignore it – they shouldn't so confidently brand it AI without any proof at all, just because it fits their mood and style.
Notably, Pangram is very conservative, and it's not difficult to manually get an LLM generated passage of text to turn human-written. So a score of near-100% AI generated means the writer didn't do even very light editing for a large part of the text.
There are extremely good reasons to be skeptical of fully LLM-written content. Our attention spans and our online platforms were built in a time where a long, data-supported article with references was expensive to produce. The time to write it was vastly longer than the time to read it, which means you could usually rely on some good faith, baseline level of accuracy and thinking on the writer's part.
With LLM-generated content, it's very difficult to know if 5 minutes, 5 hours or 5 days went into writing of the content. On the surface, it all looks similar, but the 5 minute version usually communicates very little or very shallow ideas, makes factual errors, and is generally lacking a lot of context. It's fast food writing.
These low effort versions of content take way more to read than they take to write. And combined with the obtuseness of the writing style, it all places undue burden on the reader to figure out the underlying message, because a lot of it has been mangled by the writing process.
I think LLMs are hugely helpful for writing, but to use their proper potential one needs to use them for feedback and engage with them at a level deeper than simply "write an article about X" or "rewrite this paragraph", and the text then doesn't obviously read AI generated as a bonus - I think nobody really has a problem with this.
Per the writing, reading AI writing is like having something taste "chemical". Not very specific, but still a very recognizable and bad taste that makes it hard to enjoy and marks the thing as low quality.
> I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling.
Was that the case? Genuine question. Here the em dash is an actual em dash symbol, where in the other paragraph you used what looks like a minus symbol. It’s also the type of construct used by Claude.
https://reinvently.co.uk/about/
The parent was clearly not stating anything confidently.
> Look at the later comments, they have substance oriented discussions.
There are two sentences in the parent comment. Why is the top response (yours) pointedly ignoring the sentence with substance?
Other than having all the AI tells, what proof can there be? Are you seriously asking that, because there is no way to provide hard proof that we should just stop pointing out obvious AI tells?
In this article, _I_ get unstuck right at the very first paragraph:
> Same driver, same track. The LLM is the star. Seventeen leading models driven round the identical 28-realworld task lap — one harness, same verbatim prompts, deterministic grading — and the results go on the board.
It jsut doesn't make much sense to me. At best, I think it can be glossed as... "I made an arbitrary benchmark which I'm not going to explain, and I plotted the results."
------
Getting my own opinions out, this is blatant slop. It claims to be "deterministic grading", but then almost the _entire_ webpage is editorialization. Examples:
* "If you only run one model, run glm-5.3"
* "opus-5 posts the best rubric on the default panel"
* "deepseek-v4-pro is nominally cheaper still at $0.0029 [...] treat it as a batch-only option."
* " It performed well on what it completed"
I won’t be surprise if 80% of content in social media - including maybe HN - is AI generated as of now.
And this really has to be policed. Once some tipping point is reached and too much of the HN homepage is AI slop, the site is dead.
Further, you ignored the actual contents of my comments, to latch onto a superficial aspect. Please tell me: why do these results contradict existing attempts at benchmarking LLMs, which were designed with considerably more effort? Because the website certainly doesn't explain why in a way that is human-readable, which is why I asked.
EDIT: as explained by another commenter below, its because Fable refused to perform some of the tasks.
Whatever we call 'AI coded' now has an awful lot in common with old TV screenplays where no word of dialogue was wasted. It's all the same style: punchy, plays on words, a bit of smart-ass in there.
But beyond that, can't you see how terrible the writing is? This is unadulterated AI slop.
But, you're right. The prose is miserable Claude-speak, difficult to wade through.
EDIT: You were correct, Fable and Opus reject some of the coding tasks, which is why they score lower. Thanks for explaining.
EDIT2: I believe this benchmark is invalid, my Opus 5 runs the supposedly rejected tasks just fine.
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
Data at https://gertlabs.com/rankings
maybe if you mentioned 3.7-flash it might have been slightly more believable (but still false).
The reason for both of those things is, as you point out, the benchmark is very obviously saturated
Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.
I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.
Some of the other failures like the colicky baby one are also probably soft refusals, it's not clear what the grading criteria are but I'm guessing it got docked for not going anywhere near a possible diagnosis.
It hallucinates more, and in more destructive ways, than other models I've worked with and generates truly atrocious jargon and bizarre inhuman explanations that end up cluttering things. The code it writes is terrible too. Overly complex with a lot of technical debt.
Fable is good in a few very specific domains (graphics programming) but otherwise it's an overhyped model. Far too expensive too. Opus 5 is outright better in every metric.
can't take any generated benchmark seriously. if you produce actual results, then produce actual copy to go with it.
That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.
Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.
Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.
> The only downsides are
The catch's in their revolting terms of service.
Yes, Z.ai demands an unconditional, irrevocable, transferable, sublicensable, perpetual, worldwide license to use, modify, reproduce, adapt, publish, perform, distribute, and create derivative works from your prompts and outputs. The same license extends to your username and profile picture.
Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Prohibitions on "disturbing" or "inappropriate" content, whatever that is. Professional use prohibitions.
Discussing Z.ai is prohibited to the point even my posting this comment is against their terms of service.
And they can of course ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan you won't ever see that money ever again.
Even the US companies aren't this bad.
https://chat.z.ai/legal-agreement/terms-of-service
I don't even know if I could use code generated by them, because they claim the copyright.
Better don't travel to Singapore (wise anyway because someone could slip drugs into your suitcase) or China if you use them.
Additionally I may want to run attack simulations which requires the removal of safeguards. My only option is to use an open model I can run on my own hardware.
These are open weight models (GLM-5.3 soon too). You can run them on the Together AIs or Firework AIs of this world. Use OpenRouter or HF Inference Providers in between and you can effortlessly switch between models and providers.
I have been using GLM and Kimi models the last few months mixed with the latest Anthropic models and for my daily work there is barely a difference anymore (except for pricing).
All of that pretty much means nothing for the orgs that just want to do the AI equivalent of picking IBM.
Am I the only one that knows people in industries like insurance, banking, consultancy, materials, etc? Cause none of them gives two damns about what the leading SOTA is, procurement and compliance matter.
I got their lowest subscription tier and burned a week of quota on trialling it.
I feel that choosing tasks that don't trip the classifier would also be a form of bias towards Anthropic.
In case it's not clear, the coding tasks are really benign things; there's nothing security-, health- or biology- related in there. The classifier being tripped is definitely unreasonable.
- Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers.
- X, not Y
- A, never B
- Tasteless em dashes
- Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.
Please, bloggers, write with your own voice. Don’t let an LLM do it for you.
So I have a feeling a lot of these early claims are not going to pan out in the long run.