FR version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
81% Positive
Analyzed from 2659 words in the discussion.
Trending Topics
#model#models#glm#https#alpha#open#com#better#weights#more

Discussion (88 Comments)Read Original on HackerNews
Very impressive model.
Here are some examples, open-source documented and the data available in HF datasets:
https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone
https://openzot.github.io/arcade/ - https://github.com/openzot/arcade
https://openzot.github.io/machinery/ - https://github.com/openzot/machinery
The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had this problem was Mimo 2.5, which is quite dated at this point. As a result of this, you cannot leave it unattended / not usable for agents.
It sounds like you were using a quant model.
An amusing thought of returning to your workstation to find it as an obsidian block after it gets stuck executing "dd" thousand times.
Didnt know they exist - looks very good, maybe even better than Archive.ph
https://livebench.ai/
while here it outperforms Fable by a significant margin:
https://oxalpha.com/
but if the latter is true, will people still say it was "distilled" from Fable?
It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.
Source: https://twitterwebviewer.com/?tweet=2091116504787935350
Many people and even software engineers fall for this all the time.
Most of these people are from crypto pivoting to AI doing this.
AI has made this easier and cheaper and it is going to get a LOT worse.
Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds.
The public have no chance.
The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
The metric used there is me screaming at my screen per operating hours.
Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3)
And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim.
Z.AI is the only provider for GLM 5.3 on OpenRouter. I don't see 5.3 on Hugging Face. Not sure if this new model is "full GLM" or something smaller, or if they will like Moonshot AI publish weights but put restrictive license [1], which will again leave Z.AI as single GLM model provider on OpenRouter.
[1] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.
2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.
I am leaning towards 1.
Related PR: https://github.com/jeffhajewski/latticedb/pull/5
The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens.
I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.
https://x.com/syneryder/status/2091978367579156569/photo/1
Created in a single turn - but technically not a "one-shot", because I gave it a tool to convert SVG to PNG so it could visualize what it had made. I asked it to keep iterating with tools during the same turn until it was happy.
I've also been using Ox Alpha for tasks that better resemble real work, and I'm really enjoying working with it. I've downgraded my Anthropic account so I can put some budget towards Ox Alpha instead, with the rumors that this one is going to be cheap. Opus & Fable are still better at getting large tasks / features done autonomously, but Ox Alpha can work autonomously too, and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. (As much as I don't want to say that, as someone with Claude /stickers on their laptop.)
That’s very valid, but right now every other model I use is easier to talk to than Opus 5.0
Opus 5.0 has an impenetrable way of communicating. I can parse it, but it takes so much more work than it should.
hard agree. it does not really feel "smart", but the personality is super refreshing
Chinese labs are not releasing all of their model weights. Qwen is known as an open weight model by most, but their top model is not open weight.
Releasing weights is a marketing strategy for newer labs to get their brand out there.
Seems legit.
It's really hard to know how good it is. So much hype around it.
Where? And "Tonight" in which timezone?
Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and told them they missed the timezone :)
In any case it should be releasing in a few hours. Timezones are hard.
https://x.com/davis7/status/2091285712566140986
Wenghi is behind DeepSWE, one of the best benchmarks.
Inference was atrocious in terms of speed and constant timeouts. If it's served fast it will be a delight to use.
It took me a year talking about it until my wife knew that ChatGPT and Gemini are two different things.
PS: some answers, especially if you do a deep dive on comment history, make it very clear about the joint effort from some entities to drum up support for Chinese models. This has been clear on HN as anything even mildly critical of Chinese tech gets downvoted unnaturally quickly. One can just wonder what's behind the effort...
It took me a year talking about it until my wife knew that Kimi K3 and GLM 5.3 are two different things.
But why does that matter? End users (I believe, feel free to correct) do not really contribute all that much revenue-wise. They're certainly not the SOTA target audience.
The professional market doesn't need a household name. They need the most sensible tool for the job, and the CN models right now tick many boxes when it comes to that.
Fast follower persona clusters around emerging zeitgeist across the tellings. At the moment, arguably that's mostly Qwen for everyday hobbyists, and GLM for those that can run 512GB to 1.5TB of memory. This persona is seeking viable applied results: "I have frontier at home".
The early majority pick things up after models are curated into apps like LM Studio or one's platform app of choice, usually at least one major release behind because it takes that long to choose and package into mass distribution.
This is the step where early majority persona "has no idea" what the parade of weird names is about, they care about qualia of the conversations they try to have.
This persona is, at present, very under-served, and likely to remain so until mass devices can perform feeling like 27B at Q4 large quality better, or workplace devices can achieve a pragmatic utility like 135B at Q8 or better.
Harnesses that work where the workplace persona lives bridge this. This persona doesn't care the Chinese model name, they care "does it code?" For that, the applied harness and model take time to be matched, as JetBrains did harnessing a tailored Qwen 3.6 in the IDE. More efforts like https://www.jetbrains.com/junie/ are needed for the majority persona to perceive value from changing their workflow again.
HN's "job" is better outcomes with less friction at each persona.
I would be surprised if the specialist that knows that various Chinese models exist and/or that a user might choose a harness and model separately are a "majority" even of the early variety... in terms of revenue, humans, tokens, or any metric.
(Happy to be proven wrong)
Ox is just GLM. And z.ai is the maker of GLM.
The main players in the openweight model market have been known for a while.
And they already have significant user penetration.
This reminds me a lot of media horse-race reporting, saying that "candidate X has no chance unless they" and "candidate Y has a strong showing in", and it's very thinly cover for the publication liking Y and disliking X, avoiding talking about actual policy, and trying as much as they can to make their predictions self-fulfilling.
Not to go off topic but I am pleased to see open model support from US companies like Poolside.ai, NVIDIA, IBM, Google, etc.
on toy benches it made quite a few mistakes but was able to fix all of them on its own
(meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)
[0]: https://openrating.io/blog/current-state-of-ai-model-fingerp...