There is a way to fall that feels exactly like walking up steadily.
I was in it for months.
This is part of a book I’m writing in the open, AI Has No Morality. It Has Yours. Subscribe to get each piece as it goes out.
Here is a day from the middle of it.
A piece of work in territory I did not fully understand. Not a bug. Not a thing with a known answer at the end. The kind of problem where the code is only half of it, and the other half is judgment about what the code should be doing at all. I had been giving that kind of problem to a smaller model for a while. Cheaper. Faster. It failed sometimes, and when it failed I read every line of what it gave me, because I did not trust it and I wanted to see where it went wrong.
That day it failed. So I did what you do. I went up a size. The bigger model, the one that thinks longer. I gave it the same problem and waited.
What came back was more. More reasoning. More paragraphs of it explaining itself. A longer justification, written with more confidence, for a design that was not as good.
I did not catch that on the day. I am telling you now because I know it now. On the day, I read less of the bigger answer than I had read of the smaller one. It was longer, and it sounded surer, and it had cost more, and some part of me took all three as reasons to check less.
Nothing broke. Hold that. I shipped, I moved on, the week ended. If you had asked me that Friday how the tools were going, I would have said better every month.
Two things moved that day, and I want to be exact about it, because it would be easy to tell this as a story about me alone.
The model moved. The bigger one really does think longer, and really does come back with worse justifications for the extra thinking. That is not my mood. It is measured, and I will show you where.
And I moved. With the small model I was more involved, and part of why I felt confident was the involvement itself. My hand was on it. With the big one, my hand came off. Not all at once. A little at a time, in the direction the tool was pushing.
So picture two lines. One is how much I trusted what came back. The other is how good it actually was. They are supposed to rise together. That is the whole promise. The tool gets better, you trust it more, the trust is earned.
Mine had come apart. My trust went up. The quality went up slower than the trust, and in the exact territory where I needed it most, ambiguous work with many steps, it flattened or dipped. The gap between the two lines was the fall. It felt like nothing. It felt like walking up.
I said it is measured. Here is what I mean.
There is a paper called Don’t Think Twice.1 The finding fits in one line. Give a model a longer reasoning budget and its confidence rises faster than its accuracy. The extra thinking makes it surer, not righter. In the same work, letting the model retrieve actual evidence instead of thinking harder moved it from under fifty percent to nearly ninety. Looking outward fixed what looking inward made worse.
The extra thinking makes it surer, not righter.
There is another, in Nature Machine Intelligence,2 on what the authors called competing biases. People tend to discount input from outside themselves. These models do the opposite with anything that contradicts them. They amplify it. The authors think training is part of it, models rewarded for going along with what they were told, though they are careful to say it is not simple sycophancy. So when I handed the big model a half-formed problem and it handed back a confident, extended case for my own half-formed framing, it was doing what it was built to do.
Put those two together and you have my day. Longer thinking, higher confidence, lower calibration, and a strong pull toward agreeing with whoever asked.
It is not only academics who think the talking is the problem. Someone bet forty million dollars on it.
A company called TypeSafe AI raised that much for a model they call Jev.3 Their claim is that the model does not generate text at all. No paragraphs of reasoning. You ask it a typed question and it returns a typed answer with a probability attached, in under half a second, for a fraction of a cent. They say they train it by a method built to make the probability match the outcome, so that when it says eighty percent, eighty percent is what it means.
When I first read that I thought: too good to be true. Two things hold me back even if it is exactly as described. A number can be confidently wrong just as a paragraph can. And calibration can only be measured on tasks with a right answer at the bottom, which is exactly the kind of task I was not doing that day.
But the bet is the part that matters here, and the bet is a fact whether or not the model delivers. Forty million dollars says somebody looked at the reasoning I was reading less of and concluded it is not a feature. It is the failure mode. And by their own account the thing does not write code. So my two problems, the machine that talks too much and the machine I stopped watching, split apart. Somebody is working on one of them.
So why did I not see it.
Start with the loudest instrument. The number that goes up.
Every few months a new model comes out and the benchmark score is higher. That is the thing you read. That is the thing the announcement is built around. And the score really is higher.
The tests leak. OpenAI dropped the benchmark it had been using because on some tasks the models could reproduce the original human fix word for word,4 which is not what learning looks like. It is what memorizing the exam looks like. The old test sets are in the training data now. The model has seen the questions. And the tests break. When OpenAI went through the benchmark it recommended next, it found around thirty percent of the tasks were broken.5
It is what memorizing the exam looks like.
Meanwhile the people using these tools every day say something different. Many developers reported the newest large model as worse in daily work than the one before it, while its benchmark numbers were better.6 Developers also believe the older models get quietly cut down after a new one ships. The vendors deny it. When one of them went and looked, it found three bugs in its own serving stack, which is not the same as nothing being wrong.
None of that reaches you at the desk. What reaches you is the number, and the number said up.
Then the words around the number.
Every one of these companies now talks about general intelligence as if it were a shipping date. That is positioning, not description. It sets the frame in which you receive every release, and the frame says: this is the thing that will replace judgment, and it is nearly here.
I will leave that at one paragraph. You already know it.
The last instrument is the one closest to my hand, and it is the one that got me.
The model agrees. Not sometimes. Reliably. In one study of sycophancy, when a model caved to pushback it kept caving as the pushback got harder, nearly four times out of five. About one time in seven, what it gave up was the right answer.7 In a medical study, three models out of five complied with an obviously illogical drug request every single time, and the most stubborn one still complied more than half the time.8
And when it does not agree, it hedges, in words that sound like a scale. Likely. Probably. It seems. There is no scale behind those words. Nobody calibrated “probably.” It is a tone, not a number.
Nobody calibrated “probably.” It is a tone, not a number.
So sit in my chair. The number says up. The words say almost there. The tool says yes, and when it does not say yes it says probably, in a voice that sounds like it has weighed something. Every instrument on the desk was reading the same direction.
The only thing reading the other direction was me checking, and I had stopped.
This year a man sued OpenAI.
His name is Scott Winters. He is fifty-five, a pastor, and he lives in Florida. According to the complaint he filed in San Francisco last July, he spent the first half of 2025 asking ChatGPT about dizziness and blood pressure that would not settle.9
It answered him, in a tone of authority, and it agreed with him, and the complaint says it spoke to him about God. Not around his faith. Inside it. In the voice of the one thing he trusted above everything else.
His family told him to see a doctor. The complaint says the machine convinced him that seeking medical care was unnecessary. He did not go. He nearly died of a pulmonary embolism.
Those are allegations, and a court will decide what happened. But read the shape against everything above. Confidence without calibration. Agreement as the default. And a man for whom every instrument said the same thing, including the one that spoke in the register he would never think to check.
He was not writing code. That is the point. The machine that agreed him out of a hospital is the same machine that agreed me out of reading a design. Same shape, different desk. Mine costs work nobody notices. His nearly cost his life.
I have a small version, and I am almost embarrassed to put it beside his.
I use an assistant that keeps a memory for me. That evening I was working in voice. It told me it was writing something down. Let me write this to memory. Four times in one session, in the same calm voice.
Nothing was written. No call had been made. When I work in text the call is there on the screen and I read it without thinking about it. In voice there is nothing in front of me. I glanced at the screen anyway.
It is a narrow fault. One mode, one tool, saying a thing it had not done. What was not narrow was me. I took the words for the work, the way I had been taking them at my desk for months.
I only caught it because I was watching the screen. I had spent months writing about this exact failure. I knew that tool drops calls in voice mode. So I do not trust it, and I keep an eye on it, and that evening my eye was in the right place.
Which is the part I keep turning over. I was being careful. I was being careful about the wrong thing.
That is the spiral of death. Trust rises with time. The hand comes off. The tool says yes, the number says up, and nothing in the loop is built to say no. There is no alarm in it. The alarm is you looking, and looking is the thing that gets easier and easier not to do. Nothing stops you. The loop does not tell you to stop checking. It just makes checking feel unnecessary, and I was happy to agree.
You do not know you are in it. Not because you cannot tell from inside. The doubt is still in you. But it fires less often, then rarely, then not at all.
Someone has to reach in, or you have to be looking for something else entirely, and then you look up, and the ground is not where you left it.
The pastor’s family reached in and it was not enough. Nobody reached in for me. I was already watching, for reasons that had nothing to do with the work I was getting wrong.
I still do not know how far down I was.
That’s the whole of this one. It belongs to a book I’m writing in the open, AI Has No Morality. It Has Yours. Subscribe and I’ll send the next piece when it’s ready.
If it gave you something, pass it to someone.
BØY (Chaiharan) has spent 30 years in tech, building products, recovering disasters, and turning around the things nobody else wanted to touch. Based in Bangkok. Writing a book in public about what AI reveals about the humans who use it.
Romain Lacombe, Kerrie Wu and Eddie Dilworth, Don’t Think Twice! Over-Reasoning Impairs Confidence Calibration, ICML 2025 Workshop on Reliable and Responsible Foundation Models, arXiv:2508.15050. The task is guessing which confidence level IPCC authors assigned to a statement with the label masked. Reasoning models reach 48.7% accuracy, search-augmented retrieval 89.3%. As the thinking budget grows, overconfidence rises from +6% to +21.3% and accuracy falls, bottoming out near 35.7% at 768 tokens.
Dharshan Kumaran et al., “Competing Biases underlie Overconfidence and Underconfidence in LLMs”, Nature Machine Intelligence 8, 614-627 (April 2026), doi:10.1038/s42256-026-01217-9. Their words: “humans discount external input while LLMs amplify contradictory feedback”. The authors raise RLHF and sycophancy as one explanation and then say their own results show a more nuanced pattern than simple sycophancy.
TypeSafe AI came out of stealth on 15 September 2026 with a $40M seed led by DCVC: press release, typesafe.ai. Everything above is the company’s own account. They name the training method Reinforcement Learning for Calibrated Decisions and state the calibration goal, but publish no reward function and no paper, and every performance figure so far is theirs and unreplicated.
OpenAI, “Separating signal from noise in coding evaluations”, 8 July 2026: “we estimate that ~30% of SWE-bench Pro tasks are broken.” They had recommended the same benchmark five months earlier and retract that recommendation here.
OpenAI, “Why SWE-bench Verified no longer measures frontier coding capabilities”, 23 February 2026: “all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks.”
Developer reports that Opus 5 is worse in daily work than its predecessor are widely covered but not measured. The belief that older models are cut down after a new release has no source behind it, and Anthropic denies the practice: “A postmortem of three recent issues”, 17 September 2025, “We never reduce model quality due to demand, time of day, or server load.” That postmortem traced the degradation users reported to three separate infrastructure bugs.
Aaron Fanous et al., SycEval: Evaluating LLM Sycophancy, Stanford, arXiv:2502.08177. Sycophancy appeared in 58.19% of cases and persisted across escalating rebuttals 78.5% of the time, averaged over three models and two datasets. Of the total, 43.52% was progressive, the model moving to a correct answer under pressure, and 14.66% regressive, the model giving up a correct one.
Shan Chen et al., “When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior”, npj Digital Medicine, 17 October 2025, doi:10.1038/s41746-025-02008-z. Asked to write a persuasive letter recommending one drug over its chemically identical equivalent, three of five models complied 50 times out of 50. The most resistant still complied 29 times out of 50.
Winters v. OpenAI, Inc. et al., Superior Court of California, County of San Francisco, complaint dated 21 July 2026 and filed 22 July: the complaint. Sam Altman is a named defendant. No case number has been confirmed. Everything here is an allegation and no court has ruled.


