Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

73% Positive

Analyzed from 3087 words in the discussion.

Trending Topics

#truth#sentence#true#more#llm#model#belief#self#https#false

Discussion (68 Comments)Read Original on HackerNews

kstenerudabout 6 hours ago
> It might seem absurd to you to even suggest superhuman AIs could function as a truth-oracle (it certainly does to me), but there are two reasons to take it seriously. First, it is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results, and I’ve had many discussions end with people delegating final authority on the truth to an AI.

There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.

igleriaabout 3 hours ago
> people delegating final authority on the truth to an AI

I can't avoid looking down on this attitude. But speaks more about people than it speaks about AI: those people want to win an argument, nothing more and nothing else

> There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.

namely grok.

nnevatieabout 2 hours ago
> those people want to win an argument

Indeed - X, specifically, is infested with "Grok is this true?" -type discourse.

cyanydeezabout 2 hours ago
well lets be fair, X is also infested with bullshit, so in that arena, it's more often more true than whatever is on X.

argument still valid but keep in mind the context.

smallnixabout 3 hours ago
> The set-up assumes that the game and life are the same thing, and such is the pervasive nature of the idea of the game within the society that just by believing that, they make it so.

(The player of Games, banks)

TeMPOraLabout 2 hours ago
Not sure what you're quoting, but what the quote seems to describe is known by other names, including "intersubjectivity", and is the very thing on which all our society and civilization stands. "Family", "loyalty", "society", "money", "market", "government", "law", "limited liability corporations", etc. are all abstract ideas made real through power of shared belief.

More specifically (and this is IMO critical), what pins such concepts down and makes them real, is not belief in the concept - but everyone's belief that everyone else believes in those concepts. E.g. it's not my belief in value of money that makes money valuable for me - it's my belief that the bank and the shopkeeper and the taxman all believe the money is worth something, which they do because they believe everyone else agrees too. This recursion allows us to expect certain things, such as me expecting to get food from store in exchange for money, and as long as enough people believe everyone else believes, the concept is as good as real.

Point being: this is not some gotcha or trickery, but the very stuff our lived reality is made off.

smallnixabout 2 hours ago
It's from a sci-fi novel, I forgot to add the quote and edited it in.

While I understood the idea of shared beliefs for particular cases (god, money), being presented with the full concept was very interesting (that was for me when reading hararis 'Sapiens').

The gotcha really is, that our massive use of LLMs creates new weird shared beliefs en passent

ykonstantabout 3 hours ago
We can call it AI Engine Optimization and it will become alright...
baqabout 5 hours ago
Title is a bit clickbaitish, but the content is well worth reading - came in with my pitchfork ready and left agreeing with basically all of it, with questions like ‘what if the probe could return 3 dimensions: truthfulness, knowledge confidence and decidability?’

Also the observation that people treat LLMs like oracles when they’re everything but is spot on, something I’ve also been thinking about and it’s quite a bit scary.

lstoddabout 3 hours ago
> what if the probe could return 3 dimensions: truthfulness, knowledge confidence and decidability?

can't be done because the other two dimensions depend on the first. or in other words it's all the same dimension only with a different name.

baqabout 2 hours ago
'I know this is true' is different than 'This is true' is different than 'It is impossible to say if it's true or false' is different than 'I've no idea', right?
lstoddabout 2 hours ago
no, and that's the point. hard to wrap your mind about this bit, yes.
Legend2440about 9 hours ago
I think this article pushes the premise farther than is reasonable.

The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.

lesostepabout 2 hours ago
As I understand it, this argument holds exactly the same for any type of truth: both objective and subjective (assuming subject is logically sound).

Whatever vector we choose, there always would be contradictions.

And while people also have internal conflicts, they are capable of adapting their belief system, while adjusting a vector would just give an LLM new self-contradictory classifier.

mike_hearnabout 3 hours ago
Not only that, but nobody really cares about whether truth vectors can work correctly under carefully constructed paradox edge cases. Well, maybe mathematicians do, but nobody else.
TeMPOraLabout 2 hours ago
Mathematicians, philosophers, cognitive scientists, and others.

The latent space encoded in the weights of a model is itself an object of study, just as interesting as the model itself.

cyanydeezabout 2 hours ago
someone on HN about a year ago was trying to argue that an LLM is not a cultural artifact for an anthropologist to study.

There's definitely people around that have no idea about the "soft" sciences and how an LLM while constructed from a "hard" science is almost entirely a "soft" science object.

samlinnferabout 5 hours ago
Ah but is the model's belief of the statements 1. complete, 2. consistent, 3. decidable?
Terr_about 6 hours ago
> The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true.

I think even that's too-optimistic: The LLM is a document-extender, so its "belief" is whether a token seems like it would statistically fit-next in a partial document, based on prior documents. This is usually not the kind of analytic truth we're interested in, and we've already figured out how to constantly extract it.

If we peek at vectors and weights, we'll we'll probably end up measuring the moods and styles for whatever tokens are about to get emitted next, whether that's dialogue for a fictional character (of various kinds), a narrator, or an impersonal memo conclusion paragraph. We'll be measuring "earnestness and conviction", on the same level as "loquaciousness" or "pleading" or "talking like a pirate."

So is Truthiness [0] what we really want? Probably not. If our document described the character as Yoda, then The Force connecting all existence ends up truthy. Using "a really gullible person" can repeat anything you supply as truthy. Even if we set things up as "a respected encyclopedia article" or "a relentlessly logical super-genius", we're really changing the influence mix of styles and biases, rather than creating a logical mind independent of text inside the LLM.

[0] https://en.wikipedia.org/wiki/Truthiness

FeepingCreatureabout 2 hours ago
I half-disagree: this is exactly the kind of analytic truth humans are usually interested in. "Truth" as perceived by humans is based on a massive system of prefiltering, narrowing, preprocessing and situational awareness.

A drop of water falls on your hand. Is it raining? Depends. Are you painting a watercolor picture outside? Then probably yes. Or are you going to the store? Then probably no. So truth is inseparable in practice from usecase. Is a whale a fish? I don't know, are you a geneticist or a poet? It's all mood and usecase. I'm not convinced there's any difference in kind between the LLM's speaker-selection and a human's choice of research field.

That is to say, it's not that I think you're wrong, it's that there is no other pursuit of truth than what you describe. If anything, the LLM's pursuit of next-token prediction is unusually honest for a truth-seeker.

TeMPOraLabout 2 hours ago
Worth remembering what the goal function behind the next token prediction is. It's what makes it go beyond moods and styles, and work with concepts of fact the same way we do.
TZubiriabout 9 hours ago
If anything, the whole vector space is the LLM's truth.
zarzavatabout 8 hours ago
LLMs are not optimized only for truth they are optimized for a more complicated objective that includes e.g. humans liking their output. It is a universal truth that to get humans to like you, you have to lie to them.
TeMPOraLabout 2 hours ago
Yes. Much like our own communication. Thus is the difference between a research paper and the poem. People like the former for objectivity, the latter for beauty.
elendilmabout 6 hours ago
<It is a universal truth that to get humans to like you, you have to lie to them.>

You don't represent all of us buddy. Some of us love truth above all else.

NitpickLawyerabout 7 hours ago
> It is a universal truth that to get humans to like you, you have to lie to them.

Thus said HAL9000

aesthesiaabout 10 hours ago
Fun, though as hinted at the end, the point of LLM "truth" probes is to measure the model's internal judgment of truthfulness. There's no reason this judgment, even if measured with 100% accuracy, couldn't be mistaken or logically inconsistent.
tomaskafkaabout 2 hours ago
> Second, some people do really believe in a kind of platonic representation space that all models converge on, and that represents the “true” state of the world. If truth is indeed an objective part of the world, then you might expect such a universal truth direction to emerge as models get better. This

This really sounds like all those scifi stories with people trying to find god in the computer. How could it be there, in the jumbled mirror of internet scraped texts.

inigyouabout 1 hour ago
I've heard some weird definitions of God. I always thought God was supposed to be a giant man sitting in the clouds, but some people apparently equate God with the meaning of life, or human civilization, or morality, or the sum of all human knowledge or experience.
abc123abc123about 1 hour ago
Easy! It is... emergent!
orduabout 1 hour ago
> This resolved the most basic liar paradox, but not every diagonal attack, since not all functions on [0, 1] have fixed points. To make this work in general, we can for example allow only continuous functions on [0, 1] (which always have a fixed point by Brouwer’s fixed-point theorem). But that restriction comes at the cost of expressivity: "This sentence has truth score less than 0.5" is not a continuous function of the truth score of the sentence.

Ok, I get it. Either expressiveness or completeness, but the question arises: did mathematicians explore systems with limits on expressiveness? In a field of computer programming there is Rust with limited expressiveness that doesn't solve all the problems, but still makes things much simpler. How about a mathematics with limited expressiveness and some unsafe blocks here and there?

layer820 minutes ago
Of course this has been explored: https://en.wikipedia.org/wiki/G%C3%B6del%27s_incompleteness_... The bar is rather low, however (like Robinson arithmetic). Basically, you’d have to forgo integer arithmetics with multiplication.
Xmd5aabout 2 hours ago
H: All right, all right... and the limitation of human intelligence?

M: Mr Gödel is telling you that the theorem limits formal systems. So you see your intelligence as a formal system, as a machine?

H (pompously): Indeed. I have the impression that everything I do ought to be done by a machine, which could moreover speak just as well in my place.

G: From where I am, it is difficult to tell whether you exist or whether you are the virtual creation of a GAT - a Generator of Automatic Truisms. Intelligence does not exist without error, perhaps even without obstinacy in error; but who would take the risk of giving a computer that kind of psychology? As for the incompleteness theorem, it certainly did not foresee bad-tempered theories...

---

Gödel's Theorem, or an Evening with Mr Homais Jean-Yves Girard

Translation: https://files.catbox.moe/kac0wu.pdf

Original: https://perso.ens-lyon.fr/pierre.lescanne/ENSEIGNEMENT/LOGIQ...

ziofillabout 9 hours ago
A direction that is 99.99% accurate survives this argument completely. For all practical purposes one does not need totality.
customguyabout 5 hours ago
For a one time coin flip, sure. For a lot of them, depending on the stakes, that can be very, very wrong.

https://www.google.com/search?q=why+99.99+accuracy+is+not+en...

It doesn't take a lot of thought experiments to realize this. Imagine if every bite of food we eat had a 0.01% chance to turn into something instantly lethal in our mouth. Average lifespans would be reduced measurably, and apart from anxiety, we'd develop all sorts of strategies and laws around that. E.g. absolutely NO eating for airplane pilots. You wouldn't go on a date to have dinner, dancing and sex, you'd go dancing and have sex, and then have breakfast. People would modify their jaws and stomachs so they could eat less, but bigger chunks of food. It would be a whole thing!

And that's not even talking about water changing on us, or a tiny chance of getting sucked into the toilet whenever we use it, and a lot of other things where going from damn near 100% to 99.99% would change everything for the worse, by so much.

Ukvabout 1 hour ago
ziofill's claim was that "A direction [in an LLM's embedding vector space] that is 99.99% accurate" is fine for practical purposes, not that 99.99% is fine for the chance of any given bite of food not killing you or similar hypotheticals - you'd want a few more 9s there.

To justify relevance of inability to correctly answer liars-paradox-type questions ("what won't your response to this be?"), the article suggested the way LLMs are used in practice is dependant on them being entirely accurate truth oracles:

> > as a truth-oracle [...] is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results

But for the replacement to make sense they just need to be more accurate than what they're replacing (ignoring other factors like convenience and cost) - in this case standard Google search results and knowledge box which were obviously not 100.0% accurate.

voxlabout 7 hours ago
A basic course in statistics will inform you of why a 99.99% accurate test should be looked at with skepticism when diagnosing a rare disease. Yet we see the fancy 9s and think somehow this many 9s is enough.
ziofillabout 7 hours ago
Sure, but that’s not what we are talking about
ulrikrasmussenabout 4 hours ago
Diagnosis of rare diseases is a subset of "true" statements, no?
elendilmabout 6 hours ago
Really? Laws of probability works against your argument.

Such a basic course in statistics needs to be scrutinized.

maxbondabout 3 hours ago
If 1 in 10,000 people have a disease, then a "test" which always reports the patient doesn't have the disease will be correct 99.99% of the time. "99.99% accuracy" should be "looked at with skepticism" in that it doesn't tell you what you need to know to understand the quality of a a test for a rare disease (a classifier under conditions of severe class imbalance); at a minimum, you would want to understand it's false positive and false negative rate, not (just) it's overall error rate.

See example "A": https://en.wikipedia.org/wiki/Base_rate_fallacy

californicalabout 4 hours ago
Every breath you take, there is a 99.99% chance that everything is normal, and a 0.01% chance that you breathe mild acid which horribly burns and causes a massive coughing fit.

It probably don’t cause long term damage unless that breath happened to be more important than normal, like while driving right as a child runs into the road.

Would you act differently knowing that you had one of these occasional acid breaths? Even if it only happened once per day on average (0.001%)

Taniwhaabout 4 hours ago
I think that one of the main problems with LLMs is that we're shoveling everything in to them, without any guidance as to what is "real"/"true"/"factual" - crazy anti-vaxxer-cooker stuff is is sitting in their along with science without any concept of scientific reality and no guidance to what is real and what is insane minds spinning on each other
cyanydeezabout 2 hours ago
I think this is closer to "true". The troubling bit is they're also trained with fiction bits. They're role players and you can get them to switch into arbitrary modes, including sychophants who will tell you whatever you want to hear is true.
Taniwha19 minutes ago
Yes exactly, sure they need to know about "Pride and Prejudice" but they have no way to tell it didn't happen, nor any way to tell that all people of that age lived that way (most were dirt poor scrabbling for a living)
gdiamosabout 6 hours ago
I wish I could get a model to state its assumptions.
reichsteinabout 5 hours ago
Models do not have assumptions. They have probabilities for what the next token should be. With enough context, in the context window and built into the model, that next token isn't completely random, it's correlated with something someone might choose to write.

But people write all kinds of crap manually. So far, the data people have been writing has tended to be denser around what people could agree on (there are many lies, but only one truth), so the model is more likely to go there.

If we start putting AI generated text into the training data, it's not clear what that means for the resulting model. It's already clear that some actors are trying to influence models by putting large amounts of content out there that agree with them.

Figuring out which content is safe to train from is the real problem for future model trainers.

Vecrabout 9 hours ago
Someone at MIRI must know how to solve this. Good luck getting them to tell you how!
FeepingCreatureabout 2 hours ago
For context: Logical Induction https://arxiv.org/abs/1609.03543

Sadly the line of research seems to have been abandoned. Kind of understandable with the ascent of LLMs: there is no time left for a multi-decade research program.

Advertisement
saithoundabout 9 hours ago
It looks like Zach Weinersmith predicted this exact line of research 11 years ago [1], when he suggested testing the liar sentence using fMRI.

The same analysis applies: the probe tells us what the LLM thinks about the truth value if the sentence, not the truth value of the sentence. I don't think anyone claimed that these probes were truth oracles.

[1] https://smbc-comics.com/index.php?id=3657

scotty79about 5 hours ago
> this sentence has no proof

As a programmer I was never impressed in such paradoxes. For me it was kinda obvious that in any sufficiently complex language you can create eqivalent of buggy infinite loop/recursion.

layer89 minutes ago
The issue is that you can’t generally determine whether a statement is “buggy” in that way, because under the assumption that you could, you can construct another paradox.
TZubiriabout 9 hours ago
Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth?

I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false" paradoxes, I feel that the sentence is paradox and therefore it's neither false nor true.

I do agree that it's a truth vector sounds like a silly panacea fantasy, though. But more logical formality is not the counter argument that would convince me of it, rather I believe that there's less formal and rigorous ways to get closer to truth.

galaxyLogicabout 8 hours ago
Interesting point about "self-referential sentences". I tend to agree. In my view a sentence saying something like "This sentence ..." does not have valid semantic meaning. It says nothing because, what "This" in "This sentence" means is ill-defined.

If terms we use are not well-defined, then sentences using such terms can not have meaning.

But for the sake of argument let's explore, what could the "this" in (so called) "self-referential" sentences refer to?

Do they refer to a specific encoding of the sentence you are reading, as some bits in computer memory perhaps?

That would require that those bits somehow have a unique "identity" and the "this" in a self-referential sentence would have to refer to those bits in specific addresses of a specific memory-chip.

But of course the "this" does not specify which memory chip, which specific (concrete) encoding of its (purported) meaning it is referring to. And if it did, then it would be talking about that specific set of bits in that specific memory-chip, not of "itself".

The fallacy is that what we perceive as a "sentence we read" is somehow "speaking" of something. But no, the sentence is not a subject, a sentence can not speak, and THEREFORE it specifically can not speak of itself.

A sentence can not speak, only actors, only subjects, like humans and AI, can "speak". And their speech must be encoded in some physical medium. A written sentence like "This sentence is ..." gives us the false impressions that somehow the SENTENCE IS SPEAKING of itself!

But speech can not speak, speech is the product of speaking.

Hence, in my view, "self-referential sentences" do not have any meaning and whatever paradoxes they might seem to create are results of confusion between ontological levels of "Subject" vs. "Speech".

aesthesiaabout 6 hours ago
Quines produce similar issues to self-referential sentences without being directly self-referential. e.g. "'Yields falsehood when preceded by its quotation' yields falsehood when preceded by its quotation."
TZubiriabout 6 hours ago
> "'Yields falsehood when preceded by its quotation' yields falsehood when preceded by its quotation."

Here be the limits of my brain.

It does feel (vibes) like an obfuscation of the self-referential canonical 'This sentence is false' example, like a sum of obfuscation techniques that are designed to confuse, but don't materially change the nature of the phenomenon:

1- reference is made implicit, instead of explicit

2- References another object, necessitating (at least) two instances of the same sentence, with one referencing the other.

3- Maybe some unnecessarily complex language that could be made simpler?

So what I'm getting out of it is that it feels like a mental trap that is hard to compute and understand because it was designed that way (or because it evolved that way), not necessarily because of it containing a fundamentally useful knowledge, it's difficulty to parse IS itself the interesting property of the sentence.

Maybe meta analysis like looking into the history of this sentence and seeing how many people went crazy going down that rabbit hole would be more enlightening than engaging in it in good faith. Which might only be useful if you have a high enough IQ that it doesn't confuse you any longer.

erehwebabout 6 hours ago
As the article notes, the sentence "This sentence is written in English" is well understood, and true. "This sentence is written in French" is also well understood and false.
TZubiriabout 6 hours ago
Right, my original comment wasn't about self-reference, but paradoxes, and not all self-references result in paradoxes.

Not sure what the difference between the language of a sentence and the truthiness of a sentence is. Maybe it has to do with the fact that the language refers to syntactical features that can be determined at 'compile' time, while truthiness refers to a semantic quality feature that can only be determined at 'runtime'. (See en.wikipedia.org/wiki/colorless_green_ideas_sleep_furiously for at the very least a funny canonical example of a syntactically valid, but semantically nonsensical sentence.)

The fact that the language function can also be mapped to reduced parts of the sentence, might also be relevant, you can deduce that "This-" and "This sentence is-" are English, while you cannot partially compute that "This sentence is-" is true. A single negation bit being enough to change the result of the operation.

roywigginsabout 8 hours ago
There is no lack of more elaborate logics:

https://en.wikipedia.org/wiki/Non-classical_logic

codechicago277about 7 hours ago
I’m not a logician either but believe this is what Tarski’s definition of truth solves for. In order to make a statement about the statement itself, you have to introduce a new meta language. Then a statement in the meta language is only true if the underlying statement is true.

Much more rigorous explanation: https://plato.stanford.edu/entries/tarski-truth/

UltraSaneabout 9 hours ago
There are a LOT of different logical system.

Paraconsistent Logic

Intuitionistic Logic

Dialetheism

cgioabout 8 hours ago
Fuzzy logic?