You used to reach an answer by making one. Blank page, some constraints, a version that failed, another.
By the time there was anything worth showing, you had judged it forty times on the way up, and most of those judgements you never noticed making.
Now it is there when you sit down. Complete, formatted, ready.
That sounds like less work, and in one obvious way it is. The trouble is what a finished thing does to the person checking it.
Where the machine wins
Some work tells you when it is wrong. Code compiles or it does not. A sum balances. A date matches the record.
Software people call the thing that catches it a verifier, and where you have one, letting the machine draft first is a straight gain. It does the slow part, the check is quick and certain, and you would be daft to insist on typing it out yourself.
I work this way most days and would not go back.
Then there is the other kind of work. A service map has no compiler. A journey has nothing that turns red when it is wrong.
The only check is a person who knows how the thing actually goes.
A journey that made sense
Some of my work has been on the assurance side, reading services against the Service Standard, which mostly means asking teams how they know what they know.
One team had done everything right on paper. Short timescales, so they leant on secondary research: existing reports, prior studies, published findings about people in their situation. All of it sound. A model took it in and produced a user journey, and the journey was clear and coherent, sensible steps in a sensible order.
They built it. It was linear. One sitting, start to finish, with the barest save‑and‑continue, because why would anyone need more than that.
Then beta.
Someone starts on a Monday, gets three screens in, and hits a question they need a document for. They leave to find it. The service has no way to hold what they have entered. They come back on Wednesday with the document and start again from the top.
That was not an edge case. That was how the thing was done: over days, in pieces, around the rest of a life. And often not by one person. A daughter working through it for a parent. A support worker with three of these open at once, none of them her own.
None of that was in the secondary research, and it could not have been. The reports recorded what happened. They did not record what it is like to be four days in with the wrong paperwork.
The research gap is the obvious lesson, and it is not the one I took away.
Nobody on that team decided the journey was linear.
A researcher drawing those steps by hand would have hit the question. You cannot lay out a sequence without asking, somewhere around the third box, whether this happens in one go or over a fortnight. The question is unavoidable when you are the one drawing.
The model never hit it. It returned the most probable shape of a journey given those sources, and the most probable shape is linear, because most written‑up journeys are.
And a journey where somebody asked that question and answered it looks exactly like a journey where nobody asked. Same boxes, same order, same confidence. Nothing on the page is marked unexamined.
Why nobody caught it
A well‑calibrated system is one whose confidence matches how often it is right. Weather forecasting is calibrated: when it says seventy per cent, it rains about seventy per cent of the time.
A language model gives you no such signal. Everything in that journey arrived in the same even register. The steps drawn from twenty solid studies and the step invented to bridge a gap came out looking identical. The model has no way to mark the difference, and no reason to want one. Calibration is the thing it does not carry. There is no mark on the page for the parts it was least sure about.
Our half is the older failure. Psychologists call it source monitoring: keeping track of where a piece of knowledge came from. We are poor at it, and we get worse as time passes and the thing gets handled.
Three weeks on, nobody remembers which parts of the journey came from the reports and which parts the model supplied. It is just the journey now. Provenance falls away first, and what is left is a diagram that everyone treats as equally founded.
So one side produces confident, unmarked work, and the other loses track of where any of it came from. Nothing in that arrangement is designed to fail. It fails anyway, and the failure looks like a perfectly good artefact.
Making got cheap, judging didn't
There was a sound reason to hand the making over. Producing an answer cost more than checking one, so giving away the expensive half looked like the obvious trade.
It was, until you look at what checking properly involves.
To judge that journey you would need to know how people actually get through it: over what span, with whom, holding what. That is the same knowledge you would have needed to draw it yourself.
The making got cheap. The judging did not move.
To check it properly you need everything you would have needed to make it. Nobody budgeted for that.
So teams produce more artefacts, faster, on the same understanding they had before. Nothing shows the gap, because an artefact built on real knowledge and one built on probable shapes arrive looking the same.
Two things follow
One is yours, and it is small.
Before you judge whether a generated thing is any good, write down what it must be assuming to be true. Not what it says. What it takes for granted. For that journey the list was short:
This happens in one sitting.
One person does all of it.
Everything they need is already to hand.
Three lines, and every one is a claim about the world that somebody should have checked. Written down, they stop being invisible, and you can walk into a room and ask which of them we actually know.
This can become a ritual, obviously. Any list can. A team can write three assumptions, nod at them, and carry on exactly as before. That is worse than not writing them, because now there is a document saying you looked.
What stops it being theatre is the second column. Not what we assume, but how we know. If the honest answer for every line is that it seemed reasonable, you have not found your assumptions. You have only written them down.
I am not above any of this. A clumsy draft I will pull apart happily, almost with relish. A good one goes past me fast, and I have caught myself reading quickest exactly when I should have been slowest.
The second thing is not yours alone, and it matters more.
That list belongs in the open, not in your head. Written where the team inherits it, the assumptions stop being your private worry. They become things the work is answerable for, which is the only version that survives you being busy.
And someone has to say the sums out loud. If delivery is counted in artefacts produced, judging loses that argument every time, and goes on losing it for years without anyone raising it.
Capacity was never really about producing the thing. It was about how much a team actually knows, and that has not got faster.
Who decided
Beta caught it, and I want to be straight about that, because it matters. The service did not go live broken. The system worked.
Look at what got caught, though, and when. Beta found the symptom: people leaving halfway through and not coming back. It did not find the assumption, which had been sitting in the artefact since the second week, unmarked, and which nothing in between was built to surface.
It cost a rebuild to learn something a question would have cost nothing.
And nobody in that chain decided anything. The linear journey arrived in the shape of the sources and passed through a model with no reason to question it. The people who built it had no way to see it had been assumed rather than found. Everyone did their job. The artefact was reviewed and approved.
The question was never asked, and nothing on the page showed it missing.