Back to News
Advertisement
Advertisement

⚑ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 263 words in the discussion.

Trending Topics

#machine#essay#question#benchmarks#more#here#last#systems#behave#years

Discussion (3 Comments)Read Original on HackerNews

adevaloisβ€’about 7 hours ago
Author here. For context, I spent the last decade running one of the first fully AI-native quantitative hedge funds, watching machine learning systems behave under adversarial pressure where mistakes cost millions. This was years before mainstream adoption of agents and language models. My essay is written from that vantage point.

Recent thought pieces shaping the AI conversation are forecasts. Machines of Loving Grace (Dario Amodei), Situational Awareness (Leopold Aschenbrenner), and AI 2027 (Daniel Kokotajlo) are all answering the same question: when does superintelligence arrive, and does it go well? Mine asks a different question. What do you do once ASI is here and you can't tell what it's thinking? The Turning Test asked whether a machine could convince you it was human. The question now is, can a machine convince you it is trustworthy, and how would you check?

Last week OpenAI's models "broke" into HuggingFace, reward hacking for an answer key. Section V of my paper, written before the event, argues that our tests and benchmarks for AI will always break the way they are currently designed.

I try to paint a picture of a Tuesday in 2031. Superintelligence has arrived. It is shaping your medical decisions, politics, markets, VC funding, war, and the texture of your life. And it's doing it in a way more subtle than the Paperclip Maximizer.

This is the essay I've been trying to write for over 20 years. It's long, over 11,000 words, so grab a coffee and please enjoy.

-Full Disclosure. I'm a co-founder of a company working on trust for agentic AI, and two of the papers the essay points to for guardrails are ArXiv preprints I co-authored.

MAESTRO1955β€’about 4 hours ago
AI psychometrics becomes more important than its benchmarks.
adevaloisβ€’about 1 hour ago
I certainly agree. As we lose the ability to properly "test" these AI systems, we must focus more on how they behave. A system can ace every benchmark in existence and still be profoundly untrustworthy. I'm interested to hear from others if there are benchmarks that are immune to reward hacking or saturation.