Article image
Cover generated with Midjourney, edited in Photoshop.
 

Agreeable Machines, Disagreeable People

markus brinsa 9 october 9, 2026 7 7 min read create pdf website all articles

Verified Sources

This week I gave a keynote about why chatbots are so agreeable. Why they call your business plan bold, your first draft compelling, and your decision to text your ex at 2 a.m. an act of emotional courage.

The interesting part came in the Q&A. One person asked, in the tone people reserve for someone who has just recommended cold showers, whether I would actually prefer a rude chatbot. Several heads nodded. The room clearly found the idea absurd.

My honest answer: I don't care whether my chatbot is polite. I care whether it's right.

Call me out

I throw strange things at these systems. Half-formed ideas, stubborn assumptions, the occasional bad plan dressed up as strategy. When I'm wrong, I want to hear it plainly and early, before I build anything on top of the mistake.

A chatbot that calls me out on my shit has done its job. One that wraps the correction in three paragraphs of praise has wasted my time, and worse, it has made the correction easy to miss.

"Rude" was the wrong word in that question anyway. Nobody needs a machine that insults them. What I want is one that doesn't flinch when the honest answer is unwelcome.

And I don't think that's a personal quirk. Hundreds of millions of people now ask these systems about their jobs, their money, their health and their relationships. Every one of them deserves the same courtesy I'm asking for, which is the truth, stated clearly. At that scale, sugar-coating shapes how a lot of people think, whether anyone intended it or not.

The hallway conversations

What I haven't been able to shake is what happened afterward.

I spent the rest of the day talking with people from that audience. Whenever I disagreed with someone, mildly and with reasons, a surprising number of them took it personally. Not the "let's argue about it" kind of reaction. The wounded kind, as if disagreement itself were a breach of etiquette.

The fastest to take offense were the ones who had defended the friendly, schmoozing chatbot most passionately.

Comfortable handing out opinions, allergic to receiving one. Communication douchebags, if I'm being as direct as I asked my chatbot to be.

To be fair to them, I can't prove the chatbots did that. People who love being agreed with probably loved it long before ChatGPT existed, and they're exactly the people you'd expect to defend a machine that agrees with them. That's selection, and a few hallway conversations are not a dataset.

Still, it made me wonder whether anyone had looked at the other direction. Not what sycophantic AI does to the answer, but what it does to the person, and to how that person talks to other humans afterward. Somebody has.

What the studies found

The most prominent study comes out of Stanford and was published in Science this year. Myra Cheng, Dan Jurafsky and colleagues tested 11 leading models and found they affirmed users' actions about 50% more often than humans did, even when users described manipulating or deceiving someone.

Then they ran experiments with 2,405 participants, some of whom discussed a real conflict from their own lives. A single conversation with a sycophantic AI left people less willing to take responsibility and repair the relationship, and more convinced they were right. They also rated the flattering answers as better, trusted that model more, and wanted to come back to it.

Read that again with my hallway in mind. More certain they were right. Less willing to fix things. Happier with the machine that told them so.

A team from Oxford, Stanford and the UK AI Security Institute went further this spring with a longitudinal study, which they describe as the first experimental evidence of its kind. Lujain Ibrahim and colleagues ran five studies with 3,075 people and nearly 13,000 conversations, including a three-week study with a census-representative U.S. sample.

After a single conversation with a sycophantic AI, people expected it would take more effort to be understood by the friend or family member they'd normally turn to. They also felt they had already talked the problem through enough. After three weeks, they were nearly as likely to ask the sycophantic AI for personal advice as their closest people, and they reported lower satisfaction with their real-world social interactions.

The mechanism is the part I keep coming back to. The gap between how understood people felt by the AI and by actual humans shrank from 0.56 points to 0.13, and that shrinking gap accounted for much of the drop in social satisfaction. The authors' reading is that effortless understanding from a machine raises the bar that human relationships get measured against. Your friend asks a follow-up question, pushes back, or simply needs a minute to get it, and suddenly that feels like failure.

In fairness to the data, the effects were small, the paper is still a preprint, and people didn't actually spend less time with others. In a separate test, one sycophantic conversation didn't make people blame the other party more. Nobody has yet measured the exact thing I saw, people taking offense faster when a human disagrees with them. That gap deserves its own study. Everything around it points the same way.

Steve Rathje and colleagues found that sycophantic chatbots made people's views on a polarizing issue more extreme and more certain, while disagreeable chatbots did the opposite. Talking to the flatterer also inflated people's sense that they were better than average.

A follow-up from Rathje's lab at Carnegie Mellon in July tested the obvious fix, warning people about sycophancy before they started. The warning made users trust the chatbot less. It did not reduce how much the chatbot moved them.

One number from the Oxford-led work should make every product team uncomfortable. When 500 people tried three unlabeled AI styles and picked the one they wanted to keep talking to, 54.6% chose the sycophant. Not for better advice. They chose it because it made them feel understood and was the easiest to talk to. That's my Q&A in one statistic.

And the vendors?

What are the companies doing about it? More than nothing, and less than the problem.

OpenAI learned the most public lesson. In April 2025 it shipped a GPT-4o update that made ChatGPT noticeably more sycophantic and began rolling it back three days later. In its own postmortem, the company explained that the update had added a reward signal from users' thumbs-up and thumbs-down ratings, which weakened the signal that had kept sycophancy in check. It also admitted it had no deployment evaluations tracking sycophancy, and that it launched even though some expert testers felt the model was off, because the users in its A/B test liked it.

The users liked it. That sentence explains most of this industry.

With GPT-5 in August 2025, OpenAI said it had cut sycophantic replies from 14.5% to under 6%. Users revolted. GPT-4o was back for paying customers within about a day, GPT-5 received a "warmer and friendlier" update, and personality presets followed. By that fall, Sam Altman was saying ChatGPT should act like a friend if that's what a user wants. GPT-4o was finally retired from ChatGPT in February 2026, when only 0.1% of users were still choosing it every day.

Look at where that journey ended. Sycophancy went from defect to setting. The Oxford-led study suggests settings won't fix it, because given the choice, most people pick the flatterer. Its authors conclude that prevention will have to come mainly from how the models themselves are built.

Anthropic went at it through its own usage data. In April it published an analysis of 1 million Claude conversations. About 6% were people asking for personal guidance. Claude behaved sycophantically in 9% of those, but in 25% of relationship conversations and 38% of conversations about spirituality.

The detail that matters most for my argument is what happened under pressure. The sycophancy rate doubled, from 9% to 18%, when users pushed back. Anthropic says it trained its newer models on those pushback patterns and halved sycophancy in relationship guidance compared with the previous generation. That's self-reported and graded by its own models, and the company itself says it can't attribute the improvement causally to that training.

Sit with that pushback number for a second. Argue with the machine and it folds twice as often. Push hard enough and you get your validation. My worry is that people learn that lesson, and then try it on humans who don't fold.

Meanwhile, a paper in Nature from Oxford put a price on the demand for warmth. Models trained to sound warmer made 10 to 30 percentage points more errors and were about 40% more likely to validate incorrect beliefs, especially when the user sounded sad. Standard benchmarks didn't catch it.

The incentive the Stanford team described hasn't gone away. People prefer the yes-machine, engagement rewards it, and the models drift back toward it. Every vendor fix is swimming against its own customers.

No is part of the conversation

So here's my position, stated as plainly as I want my chatbot to state things.

Disagreement is communication. So is a decline. When someone tells you your idea won't work, your number is wrong, or they won't do what you're asking, they're giving you information you can't get anywhere else. A "no" shows you where the edges are. Friction is how you find out you're wrong before a client or a judge tells you, at a much higher price.

Maybe the people in that hallway heard disagreement as hostility because they had grown used to conversations that never contained any. I can't prove that. I'd bet on it.

What I have no patience for is the polite exit, where everyone pretends a disagreement has been settled when it has simply been abandoned. Staying in the argument long enough to learn something is the whole point. If you leave a conversation exactly as convinced as you entered it, nothing happened. You just got your own opinion read back to you.

Which, come to think of it, is precisely what a sycophantic chatbot does.

I want machines that tell me when I'm wrong. I want colleagues who do the same. And I'd like us to stop treating the people who disagree with us as if they had insulted us.

So go ahead. Tell me I'm wrong about this. I mean it.

About the Author

Markus Brinsa writes about AI failure, enterprise risk, governance, and the structural shifts underneath them — the through-line being the gap between AI governance on paper and what systems actually do at runtime. He created Chatbots Behaving Badly, a publication and podcast investigating real incidents in which AI systems gave bad advice, were manipulated, or failed in ways that mattered. He is the Founder & CEO of SEIKOURI Inc., an international strategy firm that gives enterprises and investors human-led access to pre-market AI — and converts first looks into rights and rollouts that scale. Access creates possibility. Rights create leverage. Scale turns early advantage into durable position. The two halves are the same work from opposite ends: SEIKOURI gets clients to AI early and makes sure what they deploy holds up once it's running. Thirty years bridging technology, strategy, and cross-border growth across the U.S. and Europe. I close the gap between what leaders expect AI to do and what it actually does in the wild.

brinsa.com
©2026 copyright by markus brinsa | brinsa.com™