DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
75% Positive
Analyzed from 4314 words in the discussion.
Trending Topics
#meta#model#don#models#right#more#spark#data#https#here

Discussion (146 Comments)Read Original on HackerNews
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"?
There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics.
It is at least somewhat remote from the bit that is doing the actual business things.
They want to position AI as an insurmountable threat in order to regulate away any future competitors. They’re trying to speedrun regulatory capture.
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
[1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
That's obviously not the issue with that -- you don't see those comments on Google's AI announcements.
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
You might consider following your own advice.
Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :)
That:
- like all models it was trained on stolen data
- additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of
> If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Why should anyone stay silent?
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
They all suck. Pick your poison.
Here's the order, from best to worst.
Amodei
Google
SamA
Zuck
Elon
I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI?
* great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall vision, and managing them properly to ship amazing stuff.
* sane and reasonable takes on AI/LLM stuff. I can't really argue with "pursuit of truth" as the guiding principle. Grok talks normally without "Claudlish", has a balanced score on political bias unlike other models, has a low hallucination rate, is the best at dealing with latest news (unlike ChatGPT that refuses to believe new developments and gaslights the user), and they "never silently downgrade intelligence or fall back to other models."
In contrast, while Dario is doubtless a super smart pioneer in the AI space, his sanctimonious "We know what's good for you" attitude and extreme censorship is really offputting. The lengths to which he tries to ban or hamstring open models seems like an underhanded way to defeat competition. If he were to succeed, it would be a big setback to the thriving ecosystem of open models and hamper the development of the entire industry.
> Meta announces they have a new model, demonstrating its capabilities.
> Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“
That's pretty much 90% of HN these days.
Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.
Microsoft releases a new version of Windows? Here come the gripes about Azure.
Google changes something in GMail? Play Store!
It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...
I plan to update it with more pelicans from all the models released since.
(Spoiler alert: They haven't improved much since then).
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
https://news.ycombinator.com/item?id=49538333
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
Bird knees bend same way human ones do
Thank you for doing this, I love your benchmark the most!
Also 3X token use vs. 1.2
Definitely an upgrade over 1.2
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?
I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
Definitely shows how important a user data flywheel is for RL and model improvement.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
I don't see the wiggle room at all.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
Lmao. And their benchmark table only shows max reasoning.
https://news.ycombinator.com/item?id=49541149