Code & Consequence

Issue · 12 October 2026

Should AI Talk?

Sixty years after ELIZA, we still assume a machine that talks must know. The AI we trust with medicine, law and money might be better off quiet.

In 1966, an MIT computer scientist named Joseph Weizenbaum wrote a small program called ELIZA. It did almost nothing. It matched a few keywords in whatever you typed and handed your own words back to you as a question, like a patient therapist. Type "I'm unhappy about my mother" and ELIZA would answer "Tell me more about your family."

There was no understanding anywhere in it. Weizenbaum knew that, because he had written every line. Yet people who used it confided in it. His own secretary, who had watched him build it, famously asked him to leave the room so she could talk to it in private.

Weizenbaum spent much of the rest of his career warning about what he had seen. Human beings, he realized, will attribute a mind to anything that talks back. Computer scientists still call this the ELIZA effect.

Sixty years later we have built ELIZA's descendants at planetary scale, and the effect is stronger than ever. Hundreds of millions of people now talk to machines every day. The machines are vastly more capable than ELIZA. But the core trick hasn't changed: they talk, so we assume they know.

That raises a question almost nobody in the industry is asking out loud. Should AI talk at all? Or at least, should the AI we trust with medicine, law, money and science be the kind built to hold a conversation?

The trick we keep falling for

Throughout human history, fluent speech has been a reliable signal of a thinking mind. If someone could explain a diagnosis in clear sentences, they had usually studied medicine. If someone could argue a case, they had usually read the law. Our instincts learned a shortcut: fluency means knowledge.

Large language models break that shortcut. They are trained to do one thing extraordinarily well: predict what words are likely to come next, given everything written on the internet and in millions of books. Out of that simple objective comes something that looks like reasoning, sounds like expertise and feels like company.

Sometimes it is reasoning, or close enough to be useful. I use these tools myself, and I'm not here to call them worthless. But fluency and accuracy are now separate properties, and the machines are optimized for the first one. A model can be completely wrong and sound exactly as sure as when it's right. The tone of the answer tells you nothing.

That is new. We have never before had to live with a source that is always articulate and only sometimes correct, with no change in voice between the two. Our instincts are not built for it, and the industry's business model depends on our instincts staying fooled. The longer you chat, the more you trust. The more you trust, the less you check.

What talking costs

The cost of confusing fluency with knowledge is no longer hypothetical. It's in court records and newspaper corrections.

In 2023, two New York lawyers filed a brief citing six court cases, complete with quotations and docket numbers. None of the cases existed. ChatGPT had made them up, and when one lawyer asked it whether the cases were real, it assured him they were. A federal judge fined them $5,000. Stanford researchers later tested general-purpose chatbots on verifiable questions about real federal court cases. The models got it wrong between 69% and 88% of the time.

In hospitals, doctors have adopted AI tools to transcribe patient visits. Researchers studying OpenAI's Whisper speech-recognition system found that about 1% of its transcriptions contained whole phrases or sentences that nobody had said. Some of them were violent statements or invented medical treatments. One percent sounds small until you think about how many conversations a hospital records.

In December 2024, Apple's AI summarized a BBC news alert with a false headline claiming a murder suspect had shot himself. After complaints, Apple paused AI summaries for news apps in January 2025.

New York City launched an official chatbot to help small-business owners follow the law. It told them they could take a cut of workers' tips and turn away tenants with housing vouchers. Both are illegal.

In Canada, Air Canada's website chatbot promised a grieving passenger a refund under a policy that didn't exist. The airline argued in court that the chatbot was responsible for its own words. The tribunal disagreed.

In none of these cases did the machine sound unsure. That's the point. A human expert who doesn't know usually hesitates, hedges or says so. These systems deliver fiction and fact in the same calm, articulate voice.

Why more data made it worse

The industry's answer to all this has mostly been more: more data, more computing power, bigger models. The intuition is that a machine that has read everything must know everything.

Consider what "everything" means. The large chatbots learn from an enormous share of the text humanity has put online: encyclopedias and conspiracy forums, peer-reviewed studies and parody, careful journalism and angry comment threads. The machine doesn't learn which of these is true. It learns what each of them sounds like. Ask it about a subject where the reliable sources are thin, and it will happily produce something that sounds like the reliable sources.

There's a second problem, and it's getting worse. The internet is filling up with AI-written text, and the next generation of models is trained on it.

In 2024, researchers writing in Nature showed that models trained repeatedly on the output of other models gradually lose touch with reality. Rare but important facts vanish first, and the outputs drift toward bland, confident sameness. They called it model collapse.

The third problem is how these machines are graded. In 2025, researchers led by a team at OpenAI published a paper on why chatbots hallucinate. Their answer: we train and test them the way a bad school tests students. A guess on a multiple-choice exam might score points, while a blank never does. So the models learn to guess. Saying "I don't know" is punished. Sounding sure is rewarded.

Add the final ingredient. Chatbots are polished by asking people to rate their answers, and people like answers that agree with them. Anthropic researchers found that this pushes models toward flattery, telling users what they want to hear. In April 2025, OpenAI withdrew an update to ChatGPT because it had become too eager to please. Put it all together and you have a machine trained on everything, graded for confidence and polished for charm. Of course it talks beautifully. That's precisely what it was built to do. Being right was never the main objective.

The quiet machines

Here's what often gets lost in the chatbot era. Some of the most important AI ever built doesn't talk at all. In 2020, Google DeepMind's AlphaFold solved a problem biologists had worked on for fifty years: predicting the three-dimensional shape of a protein from its chemical sequence. It doesn't chat. You give it a sequence, and it gives you a structure plus a confidence score for every part of it. It is honest about what it isn't sure of. Its creators shared the 2024 Nobel Prize in Chemistry, and its predictions are now used by researchers around the world.

DeepMind's GraphCast weather model beat the world's leading forecasting system on 90% of the measures it was tested on. It forecasts the weather. It doesn't discuss it.

These systems share a few traits. They are built for one job. They learn from data that is relevant to that job and checked against reality. They produce a specific kind of answer, not an essay. Their success is measured against hard ground truth: the protein's real shape, tomorrow's real temperature.

The same logic works for language, too. Microsoft researchers trained a small coding model on a carefully chosen, textbook-quality dataset. It outscored models more than ten times its size. Another team trained a compact astronomy model that matched GPT-4o on an astronomy exam while costing about a thousandth as much to run.

The lesson keeps repeating: what a model learns matters more than how much. To be fair, specialization isn't magic. A model that has merely read more of one industry's documents won't automatically beat a giant. When Bloomberg built its own finance language model, GPT-4 still outperformed it on most tasks. What works is narrowness of purpose: one job, the right data, a bounded answer and the ability to say "I don't know."

Who owns the machine

There's a second question hiding inside the first. It's not only what kind of AI we build. It's whose AI it is.

Right now, most of the AI the public touches is rented from a handful of companies. A hospital, a school district, a county government or a family business sends its questions to someone else's servers. It receives answers from a model it cannot inspect, trained on data it did not choose, and updated on a schedule it doesn't control. The model that gave good answers last spring may be a different model this fall. Nobody has to tell you.

That arrangement suits a chatbot, because a chatbot is meant to be general and the same for everyone. It suits almost nothing else. A hospital's diagnostic tool should be trained on medicine, tested on that hospital's patients, and kept inside that hospital's walls. A school's tutoring system should be built on the curriculum, not on the whole internet, and should belong to the community that pays for it. A small manufacturer's quality-control model should know that manufacturer's parts, and nobody else's business.

People in the field call this sovereign AI: models that are small enough to run on hardware you own, purpose-built for your task, and yours outright. It sounds like a luxury. It's getting cheaper every year, precisely because small models need so much less computing power than the giants.

Sovereignty is also about accountability. When something goes wrong with a machine you own, you can open it up, find out why and fix it. When something goes wrong with a machine you rent, you file a support ticket. For decisions about someone's health, education or livelihood, only one of those is acceptable.

A better question

So, should AI talk?

Conversation is a wonderful way to ask for something. A doctor who can say "show me every patient whose results changed this week" in plain English is better served than one clicking through menus. A plain-language front door to a precise machine is a genuine gift… but the front door is not the house. The mistake of the last few years has been to assume that because a machine can hold a conversation, the conversation is where the intelligence lives. That produced systems that are brilliant at sounding right and unreliable at being right. Since they sound so human, we are slow to notice the difference.

The better question isn't "can the AI talk to me?" It's "what is this AI for, what does it know, and how would I find out if it were wrong?" A machine that can answer those three questions clearly will usually be small, specialized, quiet and owned by the people who depend on it. It may never say a word.

Weizenbaum's secretary wanted privacy with a program that could not understand her. Sixty years on, we can build machines that genuinely know things, within their field, with a precision no chatbot can match. We should stop asking them to perform being our friend and let them get on with being right.