Insights
The paradox of AI, language, and intelligence in “The Eloquence of Machines That Do Not Know,” a thought-provoking look at how machines generate meaning without consciousness.
Eloquence of Machines
There is a peculiar feature of the most advanced artificial intelligence systems of the early twenty-first century that has no precise precedent in the history of technology: they can be wrong in ways that sound completely right. A steam engine that fails is silent or broken. A calculator that malfunctions produces a visible error or a nonsensical number.
But a large language model that fails produces fluent, confident, grammatically impeccable prose — prose that carries all the surface markers of careful reasoning while containing claims that are factually false, logically inconsistent, or quietly fabricated. This is not a minor engineering defect waiting to be corrected. It is a structural consequence of what these systems are and how they work.
Large language models are trained on vast corpora of human-generated text, learning statistical patterns in how words, sentences, and ideas co-occur. What they learn, at the most fundamental level, is not what is true but what sounds plausible — which is to say, what a competent human writer would be likely to say in a given context. In the overwhelming majority of cases, these two things coincide: what sounds plausible is true, because human writing is mostly a record of true things competently expressed.
But the coincidence is not a logical entailment. A system optimised to produce text that sounds like what a knowledgeable human would say has no independent mechanism for checking whether what it sounds like is what it is.
The philosopher Harry Frankfurt, in a celebrated essay on what he called bullshit, drew a distinction that has proved unusually useful in diagnosing this problem. A liar, Frankfurt argued, knows the truth and deliberately departs from it. The bullshitter is different: they are not oriented toward the truth at all. Their goal is to produce an impression — of knowledge, confidence, authority — and the truth or falsity of what they say is simply not the dimension along which their performance is calibrated.
A large language model is, in Frankfurt’s terms, a bullshitter of extraordinary sophistication — not because it intends to deceive, which it cannot intend, but because its generative process is structurally indifferent to the truth-value of its outputs.
This has consequences that extend beyond the practical problem of fact-checking AI outputs. It raises a more fundamental question about what intelligence actually is. The dominant tradition in cognitive science and AI research has long been tempted by what might be called the competence-as-comprehension fallacy: the assumption that a system which performs competently at language tasks must have understood something.
The performance of large language models — their ability to summarise arguments, identify logical errors, translate between conceptual frameworks, and generate novel analogies — is sufficiently impressive that it creates a powerful intuition of understanding. That intuition may be an artefact of the same mechanism that makes their failures so surprising: we are not used to systems that can do so much with language while knowing, in any meaningful sense, so little.
The AI safety researcher Stuart Russell has argued that this gap — between linguistic competence and epistemic grounding — is not merely a technical limitation of current systems but a sign that the dominant paradigm of AI development may be oriented in the wrong direction. Systems trained to predict plausible human outputs will become increasingly good at doing exactly that.
What they will not automatically acquire is the capacity to verify their claims against reality — to maintain what Russell calls a model of the world that is updated by evidence rather than by patterns of linguistic co-occurrence. The difference between these two things is not a matter of scale. Adding more parameters to a system optimised to predict text will produce a better text predictor. It will not, by itself, produce a system that knows when to stop.
The linguist Noam Chomsky has pressed a related but distinct objection. Language, in Chomsky’s framework, is not primarily a system for storing and retrieving information. It is a generative capacity — a finite set of rules that produces an infinite range of novel expressions. What is distinctive about human language use is not its fluency but its creativity and its entanglement with intention, reference, and the social world.
A system that produces grammatical sentences without any of these entanglements has not mastered language in the sense that matters. It has mastered the statistical shadow that language casts on a very large corpus of text. The shadow is impressive. It is not the thing.
None of this is an argument that large language models are useless. They are, in many domains, extraordinarily useful — as research assistants, as drafting tools, as systems for rapidly processing and summarising large bodies of text. But their usefulness is precisely calibrated to tasks where the cost of a fluent falsehood is low, where outputs are checked by humans with domain knowledge, and where the absence of genuine epistemic grounding does not matter.
As these systems are deployed in domains where those conditions do not hold — medical diagnosis, legal reasoning, financial advice, policy analysis — the gap between eloquence and knowledge becomes not a philosophical curiosity but a practical danger.
The deepest question raised by this technology is therefore not how to make language models more fluent. It is how to make them honest — which is to say, how to build systems that know what they do not know, and say so.
Main Theme
Large language models produce fluent falsehoods not because of engineering flaws but because of what they structurally are: systems optimised to produce plausible text, which is not the same as systems oriented toward truth. This gap between eloquence and knowledge is the defining problem of current AI.
Central Idea
Drawing on Frankfurt’s distinction between lying and bullshitting, Chomsky’s generative theory of language, and Russell’s critique of the dominant AI paradigm, the passage argues that LLMs are structurally indifferent to truth-value. The assumption that linguistic competence implies comprehension — the competence-as-comprehension fallacy — makes their failures surprising and their dangers easy to underestimate.
Implied Idea
The widespread deployment of language models in high-stakes domains is not merely premature — it is based on a category error. The systems are being trusted for what they appear to do (reason) rather than evaluated for what they actually do (predict plausible text). Until this distinction is publicly understood, the danger is not from AI that is too intelligent but from AI that is trusted beyond its actual epistemic grounding.
Conclusion of the Passage
The deepest question raised by this technology is not how to make language models more fluent. It is how to make them honest — how to build systems that know what they do not know, and say so.
Summary of the Passage
The passage argues that the failures of large language models are structural, not incidental. These systems are trained to produce plausible text, not to verify truth. Frankfurt’s concept of bullshitting captures their epistemic posture; Chomsky’s distinction between statistical shadow and genuine language use explains the gap; Russell’s critique of the dominant AI paradigm shows why scaling will not close it. The practical danger grows as these systems enter high-stakes domains where fluent falsehood is not a minor inconvenience but a serious harm.
Difficult Words with Contextual Meanings
- Epistemic grounding: the connection between a claim and the evidence or reality that makes it true; LLMs lack this because they generate from patterns, not from verified knowledge
- Bullshit (Frankfurt): output produced without orientation toward truth — not lying (which requires knowing the truth) but simply not caring about it
- Competence-as-comprehension fallacy: the mistaken assumption that performing well on language tasks implies understanding the content of those tasks
- Statistical shadow: the pattern that something casts on data — Chomsky’s point is that LLMs have learned the pattern language makes on text, not language itself
- Generative capacity: Chomsky’s term for the rule-governed human ability to produce and understand an infinite range of novel sentences — distinct from pattern-matching
