Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
93% Positive
Analyzed from 1501 words in the discussion.
Trending Topics
#anthropic#model#more#mythos#fable#opus#productivity#tasks#claude#amp
Discussion Sentiment
Analyzed from 1501 words in the discussion.
Trending Topics
Discussion (48 Comments)Read Original on HackerNews
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
I find it hard to imagine launching this criticism at a new technology.
A performance improvement is relative to some baseline, and that baseline for some may be a lot lower than others, and if they adopt tech effectively it really could be a big boost. Across the board though I don’t think it’s sensible.
I work in an industry that very intentionally tries to be inefficient and I can tell you having certain tasks automated that before had a person barrier intentionally acting inefficiently that you can now sidestep by outsourcing their tasks to something like Claude gives me a massive performance increase because I’m not blocked as much anymore. I can literally just replace some external tasks that were intentionally slowing processes down for their own benefits with a few prompts and move along. I could have done the tasks before but then people would ask why I’m spending my time doing it, now I can just say “oh, I was blocked so I had Claude take care of that blocker” and move along.
It can further be true that some developers are getting 5x or 10x while as a whole their organization is sped up less than 2x. I'm sure many tasks at Anthropic are sped up 5x or 10x or more.
It can further be the case that many people overestimate their gains as well. That's fine, and I think what you're saying - but it's still wild to shake your head and go "pfft, they have not even doubled their productivity". Double is a lot!
Seems like that would be easily recupable even without future growth.
While it is true that investors only got a fraction of the company for that money, and their EV can be debated, I think the clear the value is there from a net cash in to value produced.
Last valuation was like 1 trillion. Company could be worth like 1/20 and still justify the cash.
They've probably already settled on most of the architecture and the big ideas, so they're details in big things instead of how to make complete small things.
The thing LLMs really speed up is how some ordinary person-- a PhD student, or similar, can whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where LLMs shine: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.
They can’t measure even measure it, it’s just vibes. They may not even be more productive.
Maybe they should ask an AI to create one!
Is that correct?
> Same model weights as Mythos 5, deployed with higher-coverage safeguards (see Section 4.5.2.2)
Also it occurs to me that they're somewhat incentivized to downplay cyber risks after what happened last time...
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
Seems like a strange expectation.
If Chinese model hacks US government... free marketing?
It does feel they are trying to ask the government to lock the market for us.
This sounds like it might be a Mythos finetune for some specific task.
EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench
By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
What am I missing?
LOLOLOL
>6.3 [Appendix redacted] > This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.
interesting...
EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.