ES version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
73% Positive
Analyzed from 3087 words in the discussion.
Trending Topics
#truth#sentence#true#more#llm#model#belief#self#https#false

Discussion (68 Comments)Read Original on HackerNews
There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.
I can't avoid looking down on this attitude. But speaks more about people than it speaks about AI: those people want to win an argument, nothing more and nothing else
> There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.
namely grok.
Indeed - X, specifically, is infested with "Grok is this true?" -type discourse.
argument still valid but keep in mind the context.
(The player of Games, banks)
More specifically (and this is IMO critical), what pins such concepts down and makes them real, is not belief in the concept - but everyone's belief that everyone else believes in those concepts. E.g. it's not my belief in value of money that makes money valuable for me - it's my belief that the bank and the shopkeeper and the taxman all believe the money is worth something, which they do because they believe everyone else agrees too. This recursion allows us to expect certain things, such as me expecting to get food from store in exchange for money, and as long as enough people believe everyone else believes, the concept is as good as real.
Point being: this is not some gotcha or trickery, but the very stuff our lived reality is made off.
While I understood the idea of shared beliefs for particular cases (god, money), being presented with the full concept was very interesting (that was for me when reading hararis 'Sapiens').
The gotcha really is, that our massive use of LLMs creates new weird shared beliefs en passent
Also the observation that people treat LLMs like oracles when they’re everything but is spot on, something I’ve also been thinking about and it’s quite a bit scary.
can't be done because the other two dimensions depend on the first. or in other words it's all the same dimension only with a different name.
The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.
Whatever vector we choose, there always would be contradictions.
And while people also have internal conflicts, they are capable of adapting their belief system, while adjusting a vector would just give an LLM new self-contradictory classifier.
The latent space encoded in the weights of a model is itself an object of study, just as interesting as the model itself.
There's definitely people around that have no idea about the "soft" sciences and how an LLM while constructed from a "hard" science is almost entirely a "soft" science object.
I think even that's too-optimistic: The LLM is a document-extender, so its "belief" is whether a token seems like it would statistically fit-next in a partial document, based on prior documents. This is usually not the kind of analytic truth we're interested in, and we've already figured out how to constantly extract it.
If we peek at vectors and weights, we'll we'll probably end up measuring the moods and styles for whatever tokens are about to get emitted next, whether that's dialogue for a fictional character (of various kinds), a narrator, or an impersonal memo conclusion paragraph. We'll be measuring "earnestness and conviction", on the same level as "loquaciousness" or "pleading" or "talking like a pirate."
So is Truthiness [0] what we really want? Probably not. If our document described the character as Yoda, then The Force connecting all existence ends up truthy. Using "a really gullible person" can repeat anything you supply as truthy. Even if we set things up as "a respected encyclopedia article" or "a relentlessly logical super-genius", we're really changing the influence mix of styles and biases, rather than creating a logical mind independent of text inside the LLM.
[0] https://en.wikipedia.org/wiki/Truthiness
A drop of water falls on your hand. Is it raining? Depends. Are you painting a watercolor picture outside? Then probably yes. Or are you going to the store? Then probably no. So truth is inseparable in practice from usecase. Is a whale a fish? I don't know, are you a geneticist or a poet? It's all mood and usecase. I'm not convinced there's any difference in kind between the LLM's speaker-selection and a human's choice of research field.
That is to say, it's not that I think you're wrong, it's that there is no other pursuit of truth than what you describe. If anything, the LLM's pursuit of next-token prediction is unusually honest for a truth-seeker.
You don't represent all of us buddy. Some of us love truth above all else.
Thus said HAL9000
This really sounds like all those scifi stories with people trying to find god in the computer. How could it be there, in the jumbled mirror of internet scraped texts.
Ok, I get it. Either expressiveness or completeness, but the question arises: did mathematicians explore systems with limits on expressiveness? In a field of computer programming there is Rust with limited expressiveness that doesn't solve all the problems, but still makes things much simpler. How about a mathematics with limited expressiveness and some unsafe blocks here and there?
M: Mr Gödel is telling you that the theorem limits formal systems. So you see your intelligence as a formal system, as a machine?
H (pompously): Indeed. I have the impression that everything I do ought to be done by a machine, which could moreover speak just as well in my place.
G: From where I am, it is difficult to tell whether you exist or whether you are the virtual creation of a GAT - a Generator of Automatic Truisms. Intelligence does not exist without error, perhaps even without obstinacy in error; but who would take the risk of giving a computer that kind of psychology? As for the incompleteness theorem, it certainly did not foresee bad-tempered theories...
---
Gödel's Theorem, or an Evening with Mr Homais Jean-Yves Girard
Translation: https://files.catbox.moe/kac0wu.pdf
Original: https://perso.ens-lyon.fr/pierre.lescanne/ENSEIGNEMENT/LOGIQ...
https://www.google.com/search?q=why+99.99+accuracy+is+not+en...
It doesn't take a lot of thought experiments to realize this. Imagine if every bite of food we eat had a 0.01% chance to turn into something instantly lethal in our mouth. Average lifespans would be reduced measurably, and apart from anxiety, we'd develop all sorts of strategies and laws around that. E.g. absolutely NO eating for airplane pilots. You wouldn't go on a date to have dinner, dancing and sex, you'd go dancing and have sex, and then have breakfast. People would modify their jaws and stomachs so they could eat less, but bigger chunks of food. It would be a whole thing!
And that's not even talking about water changing on us, or a tiny chance of getting sucked into the toilet whenever we use it, and a lot of other things where going from damn near 100% to 99.99% would change everything for the worse, by so much.
To justify relevance of inability to correctly answer liars-paradox-type questions ("what won't your response to this be?"), the article suggested the way LLMs are used in practice is dependant on them being entirely accurate truth oracles:
> > as a truth-oracle [...] is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results
But for the replacement to make sense they just need to be more accurate than what they're replacing (ignoring other factors like convenience and cost) - in this case standard Google search results and knowledge box which were obviously not 100.0% accurate.
Such a basic course in statistics needs to be scrutinized.
See example "A": https://en.wikipedia.org/wiki/Base_rate_fallacy
It probably don’t cause long term damage unless that breath happened to be more important than normal, like while driving right as a child runs into the road.
Would you act differently knowing that you had one of these occasional acid breaths? Even if it only happened once per day on average (0.001%)
But people write all kinds of crap manually. So far, the data people have been writing has tended to be denser around what people could agree on (there are many lies, but only one truth), so the model is more likely to go there.
If we start putting AI generated text into the training data, it's not clear what that means for the resulting model. It's already clear that some actors are trying to influence models by putting large amounts of content out there that agree with them.
Figuring out which content is safe to train from is the real problem for future model trainers.
Sadly the line of research seems to have been abandoned. Kind of understandable with the ascent of LLMs: there is no time left for a multi-decade research program.
The same analysis applies: the probe tells us what the LLM thinks about the truth value if the sentence, not the truth value of the sentence. I don't think anyone claimed that these probes were truth oracles.
[1] https://smbc-comics.com/index.php?id=3657
As a programmer I was never impressed in such paradoxes. For me it was kinda obvious that in any sufficiently complex language you can create eqivalent of buggy infinite loop/recursion.
I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false" paradoxes, I feel that the sentence is paradox and therefore it's neither false nor true.
I do agree that it's a truth vector sounds like a silly panacea fantasy, though. But more logical formality is not the counter argument that would convince me of it, rather I believe that there's less formal and rigorous ways to get closer to truth.
If terms we use are not well-defined, then sentences using such terms can not have meaning.
But for the sake of argument let's explore, what could the "this" in (so called) "self-referential" sentences refer to?
Do they refer to a specific encoding of the sentence you are reading, as some bits in computer memory perhaps?
That would require that those bits somehow have a unique "identity" and the "this" in a self-referential sentence would have to refer to those bits in specific addresses of a specific memory-chip.
But of course the "this" does not specify which memory chip, which specific (concrete) encoding of its (purported) meaning it is referring to. And if it did, then it would be talking about that specific set of bits in that specific memory-chip, not of "itself".
The fallacy is that what we perceive as a "sentence we read" is somehow "speaking" of something. But no, the sentence is not a subject, a sentence can not speak, and THEREFORE it specifically can not speak of itself.
A sentence can not speak, only actors, only subjects, like humans and AI, can "speak". And their speech must be encoded in some physical medium. A written sentence like "This sentence is ..." gives us the false impressions that somehow the SENTENCE IS SPEAKING of itself!
But speech can not speak, speech is the product of speaking.
Hence, in my view, "self-referential sentences" do not have any meaning and whatever paradoxes they might seem to create are results of confusion between ontological levels of "Subject" vs. "Speech".
Here be the limits of my brain.
It does feel (vibes) like an obfuscation of the self-referential canonical 'This sentence is false' example, like a sum of obfuscation techniques that are designed to confuse, but don't materially change the nature of the phenomenon:
1- reference is made implicit, instead of explicit
2- References another object, necessitating (at least) two instances of the same sentence, with one referencing the other.
3- Maybe some unnecessarily complex language that could be made simpler?
So what I'm getting out of it is that it feels like a mental trap that is hard to compute and understand because it was designed that way (or because it evolved that way), not necessarily because of it containing a fundamentally useful knowledge, it's difficulty to parse IS itself the interesting property of the sentence.
Maybe meta analysis like looking into the history of this sentence and seeing how many people went crazy going down that rabbit hole would be more enlightening than engaging in it in good faith. Which might only be useful if you have a high enough IQ that it doesn't confuse you any longer.
Not sure what the difference between the language of a sentence and the truthiness of a sentence is. Maybe it has to do with the fact that the language refers to syntactical features that can be determined at 'compile' time, while truthiness refers to a semantic quality feature that can only be determined at 'runtime'. (See en.wikipedia.org/wiki/colorless_green_ideas_sleep_furiously for at the very least a funny canonical example of a syntactically valid, but semantically nonsensical sentence.)
The fact that the language function can also be mapped to reduced parts of the sentence, might also be relevant, you can deduce that "This-" and "This sentence is-" are English, while you cannot partially compute that "This sentence is-" is true. A single negation bit being enough to change the result of the operation.
https://en.wikipedia.org/wiki/Non-classical_logic
Much more rigorous explanation: https://plato.stanford.edu/entries/tarski-truth/
Paraconsistent Logic
Intuitionistic Logic
Dialetheism