One evening in August I asked a coding agent to change a sentence at the top of a page. It said done. I opened the page and the sentence was as it had been. I asked again, and it said done again. The third time I asked why.

It had changed the sentence, in a file. The page was built from a different file. It had not looked at the page, because I had not said the page. I had said the sentence. It took the likeliest meaning of what I typed and reported success in the tone it would have used if it had checked.

A person would have asked which page. Or looked. Nothing in the exchange said huh.

Eighty-four seconds

In 2015 the linguist Mark Dingemanse and his colleagues went through recordings of ordinary conversation in twelve languages from eight families. They counted how often somebody stopped to check what the other had meant, and got it put right. About once every eighty-four seconds, across all twelve, counting every kind of check from a plain huh to a guess at what was meant. The fix, on average, ran about as long as the thing it fixed.

This is conversational repair, and it is most of how intent gets read. Nobody reads another mind in one go. You read a little, check, and read again.

The commonest gap between one person finishing and the next starting is under a fifth of a second, in every language measured. Which is quicker than a sentence can be planned. So we are working out what someone means while they are still saying it, and one short word puts it right when we are wrong.

Language was spoken for a very long time before it was written. It grew up between two people who could see each other, with tone and timing and a face to read. If the reading went wrong, the speaker was a second away.

Dingemanse had earlier gone looking for the word itself, and found something like it in all ten of the languages he checked, on five continents. It sounds much like the same word in Siwu, spoken in Ghana, as it does in Dutch. Languages differ in almost everything else. They agree here.

Whatever shaped language did so in conversation, and the repair came with it. In 2024 Dingemanse and a colleague argued that it is the scaffold the rest of language's complexity stands on.

Built to be misheard

That history explains something odd about the words we use. They are ambiguous, and the ambiguity is a feature.

Three psycholinguists made the case in 2012. Any efficient language will be ambiguous, provided the listener has context to fill the gaps. Ambiguity also lets short words be reused. Bank. Fine. The speaker spends less and the listener sorts it out, and where they cannot, they ask. Natural language is cheap to speak because it is cheap to repair.

Programming languages made the opposite bet. Every symbol means one thing, because the reader cannot ask. Where two readings are possible, the rulebook picks one in advance, or the compiler stops. A language for a reader with no huh cannot afford to be misheard.

A large language model takes the first kind of language and runs it like the second. It receives words built for a listener who can ask, settles on the likeliest reading, and acts. The ambiguity comes in with the words. The repair does not.

Two readings of the same words. One checks and bends toward what was meant. One keeps its first reading and goes straight past. Two hand-drawn lines leave the same starting point, labelled what was said. A small ring to the upper right is labelled what was meant. One line bends twice, at points marked huh, and lands inside the ring. The other holds its first heading, passes below the ring and runs off the edge with an arrowhead, labelled what was heard. what was said what was meant huh? huh? what was heard
Both lines start from the same words. One checks twice and lands on what was meant. The other delivers what was heard, and keeps going.

It can tell

In May 2026 Jinyan Su and Claire Cardie posted a study whose title is the finding: Knowing but not showing. Asked whether a request could mean two things, most of the models they tested could say so. If anything they over-call it, and are worse at spotting the requests that are clear. But given the same request cold, they answered. One model asked a clarifying question on about two ambiguous requests in a hundred. Given documents to draw on, it asked on one in three hundred. It checked least when it had most to go on.

I ask a model things fifty times a day. I could count on one hand the times one has asked me something back.

Many requests are underspecified. They do not pin down one meaning. A person leaves the gap because the listener can close it. The model closes it with the most common reading for words like yours, and carries on as if it had asked.

Nobody set out to build this. It falls out of how the work is scored. The training rewards whatever people rate as helpful, and asking rarely made the list. Score it that way for long enough and you get a machine that would rather be wrong than ask.

What the interface did

For most of the history of public services, the person did the translating. You arrived with a worry in your own words and the screen could not take it. It wanted a reference number, a category, a menu with nothing on it that matched your situation.

We got good at softening that, with plain words and fewer fields, and the deal underneath held. The interface spoke its language and you met it there.

The machine can now take the need in the words it arrived in. For a lot of people that is a real gain. The person who could never find the right option no longer has to. Meeting people in their own words is what this profession has wanted for as long as it has existed.

The old interface did do one thing the assistant does not. It showed you the choices, so a wrong one was visible, and yours to change. The old failure left marks, abandoned journeys and call volumes, and you could count it and fix it. The new one is smooth. The reading was wrong and the action was taken, and nothing hesitated.

The person may never learn that a choice was made for them.

Software used to be expensive to build, and now it is expensive to misunderstand. Building was the slow part, and the slowness was where misunderstandings got caught: the designer asking what it was for, the tester's raised eyebrow. When a brief becomes working software in an afternoon, none of those people are in the chain.

The same is true of the person typing a worry into a box at midnight. No one checks at either end.

Keep the huh

GDS put GOV.UK Chat in front of more than 10,000 people over eighteen months. One of the five things the team wrote up in March 2026 was this. People got no answer when they phrased a question in a way the system could not answer with confidence.

The team did not make it guess. They let it ask, with clarifying questions for ambiguous phrasing. And the answer rate on questions it should have handled now stands at 88 per cent.

It was one change among several, and the team does not say the question did all of that. Nor do I. Met with a machine that could not answer with confidence, a government team chose asking over guessing.

Nobody wants an assistant that asks did you mean at every step, and the labs know it. OpenAI's own specification tells its models to avoid trivial questions and assume a reasonable person. It also tells them to ask when the cost of guessing wrong is too high, and to seek approval before going beyond what they were asked.

So the rule is already written down. Most of the models can tell an ambiguous request when asked. They were tuned not to act on it. Whether your service acts on it is a design decision, and it is yours.

Before anything acts on a reading of a person, it shows the reading back in the person's words and waits one turn. No, I meant, costs one sentence. Where in your service can the person say no, I meant? If the answer is start again, the service is guessing, with good manners.

GOV.UK has had a version of this for years. The check your answers page shows you what you typed before anything is sent. But a form only ever had to read back what you put in its boxes. An assistant has to read back what it took you to mean, and that is a harder page to design.

A person in a conversation checks when getting it wrong would cost something, and lets the rest go. The rule for your service is the same. Show them the reading whenever it is cheap to show, and ask before anything they cannot easily undo: money moved, a date fixed, an entitlement decided, a record changed. Everything else, act, and keep the way back short.

Confirmation boxes have been around for forty years, and everyone clicks yes. An are-you-sure box is empty. There is nothing in it to read, so yes is the only thing to do.

A read-back has something in it: what the machine thinks you meant, in your own words. If it has got you wrong, you can see it before anything happens.

And it only appears when something cannot be undone. The box on every step is the one people learn to click through.

It could have asked which page. Two words, and it said done instead, three times. The only huh in the whole exchange was mine, and it cost a sentence on a web page. Yours has somebody's money in it.