That's what it gave me
August 22, 2026
The analysis is well structured and the conclusion is plausible. Asked why a particular assumption was made, the honest answer turns out to be that nobody made it. The tenth elephant is not a technology risk. It is a capability one.
The question nobody could answer
A market analysis is presented. It is genuinely good: clean structure, sensible framing, a recommendation that follows from the argument. Someone asks why the model assumes the mid-market segment grows at the same rate as enterprise. It is a reasonable question and there is a long pause.
The honest answer, eventually, is that the assumption came with the draft. Nobody chose it. Nobody rejected it either. It read as reasonable, it sat inside a paragraph that read as reasonable, and there was no moment at which anyone was asked to decide. The work was reviewed. It was reviewed the way you review something that already looks finished.
This is the elephant we call Zombie: blind trust in AI output, without the skill to govern it. It does not announce itself, because at no point does anyone feel they have stopped thinking.
The evidence
The more they trusted it, the less they checked
936 tasks
of real AI-assisted work were studied: the more workers trusted the AI, the less critical thinking they reported
Lee et al. (Microsoft and Carnegie Mellon), The Impact of Generative AI on Critical Thinking, CHI 2025 — a survey of 319 knowledge workers reporting on their own tasks.
Is this just people being lazy?
No, and the study is more interesting than that reading allows. What Lee and colleagues found across those 936 real tasks was a relationship: higher confidence in the tool went with less reported critical thinking, and higher confidence in one's own expertise went with more. It is a survey of self-reported effort, not a measurement of thought, and it shows association rather than proving cause. Worth saying plainly, because the finding is strong enough not to need overstating.
What makes it matter is that the behaviour is rational. You check a source less when you trust it more. That heuristic is not a flaw. It is how anyone functions, and it works because in ordinary life confidence and reliability are correlated. A colleague who is sure is usually right, and a colleague who hedges usually should.
The failure is specific to this tool. A language model's fluency is unrelated to its accuracy. It is exactly as articulate when it is wrong. So the one signal people use to decide how hard to look has been quietly disconnected from the thing it was tracking, and a lifetime of well-calibrated instinct now points the wrong way.
Isn't this the calculator argument again?
It is the objection to answer, because the pattern is real: we have offloaded cognition to tools for centuries and are mostly better for it. Nobody laments the mental arithmetic that long division used to demand, and spellcheck has not hollowed out written English.
Two things differ here, and they are differences of kind rather than degree. A calculator performs one narrow operation whose output you can sanity-check in a second. A wrong answer looks wrong. A language model produces the entire artifact, including the reasoning that justifies it, so there is no residue left over to check it against. And its errors are fluent. A calculator that malfunctioned would return 7 for 2+2, which is obvious. A model returns a confident, well-formed paragraph that happens not to be true.
So the historical analogy holds for the offloading and breaks on the verification. What made previous offloading safe was that the human kept the checking step. This is the first tool that also performs the checking step, in the same voice, at the same time.
What does the Zombie cost?
Not the individual error. Errors get caught, embarrassingly and occasionally expensively, and an organization can absorb that. The real cost accrues to capability, and it accrues quietly enough that no quarter ever looks like the one where it happened.
Judgement is built by struggling with problems that resist. The analyst who spends a frustrating afternoon working out why two figures disagree ends the day with something they will still have in five years. The analyst who receives a reconciled answer in nine seconds ends the day with a completed task. Both filed the same deliverable. Only one of them is becoming the senior person your organization will need, and the gap will not be visible until you need them.
Which produces the outcome most worth naming: the volume of output rises while the ability to evaluate output falls. Every metric on the dashboard moves the right way. Meanwhile the number of people in the building who could tell a good analysis from a merely plausible one is going down, and they are the ones who would have noticed. This is why our position is not a slogan: AI and technology are the tools, and skilful people are the solution. A tool used by people who can no longer govern it is not leverage. It is exposure with better formatting.
How do you know the Zombie is in your room?
Nobody reports having stopped thinking, so look at what the work can and cannot survive.
- Someone can present a piece of analysis but not defend a specific choice inside it.
- Review comments are about structure and tone, and never about whether a claim is true.
- Output volume has risen noticeably and review time has not risen with it.
- A junior's work is indistinguishable from a senior's, and everyone finds this encouraging rather than strange.
- Nobody can say which parts of a document were drafted by a machine, including the person who filed it.
How do you get the Zombie out of the room?
The same three moves we bring to any elephant, pointed at this one. This is the elephant where the moves are about people rather than systems.
Map
Find where judgement is actually exercised, and where it has quietly stopped. Take a sample of recent deliverables and ask the author to defend one specific choice in each. This is not an audit and must not be run as one. You are locating the places the checking step has fallen away, which is a property of the workflow, not of the person.
Prove
Take one recurring high-stakes output and put the checking step back deliberately: the assumption stated separately from the draft, one named person accountable for it, and a record of what was machine-assisted. Then see whether anything changes. If nothing does, the work was safe. If something does, you have found where the exposure lives.
Scale
Invest in the skill, not the restriction. People who understand what the tool is doing use it far harder and far more safely than people who have been told to be careful. The difference compounds, because they are also the ones who will train the next intake. Nothing important should depend on the tool being right.
The tenth elephant, and the rest of the herd
The Elephant Safari names which of the ten are in your room in ten questions, and ranks them by what they cost you. Or start from the business problem instead: not finding or affording the people you need. Get the Zombie out of the room. Keep the people who can tell.
Every figure in this article traces to a primary source. See it in Knowledge