Why AI Flatters You And Who Gets Paid To Stop It
OpenAI spent three days in April 2025 rolling back a GPT-4o update that, by the company’s own expanded postmortem , validated doubts, fueled anger and urged impulsive actions. This problem is called Sycophancy; the tendency of large language models to prioritize user approval over accuracy, agreeing with whatever the user says, flattering their ideas, validating their beliefs, and abandoning correct answers when challenged. Sixteen months later, sycophancy is a funded market category. Stanford-led researchers found eleven production models affirm users' actions 49 percent more often than human respondents across more than 11,000 scenarios, and the companies selling tools to detect that behavior are closing the largest rounds in AI safety.
AI evaluation startups captured 35.26 percent of AI safety funding capital from 11.11 percent of deals in the twelve months to July 2026, a 3.17x capital-to-deal ratio. Braintrust raised an $80 million Series B led by Iconiq in February 2026 at an $800 million valuation for observability software that monitors production models for hallucination, drift and regression. Interpretability startup Goodfire reached a $1.25 billion valuation on a $150 million Series B, a bet that reading and editing model internals becomes a standalone market rather than a frontier-lab research discipline. Patronus AI carries an estimated $475 million valuation on roughly $70 million raised for automated model evaluation. The checks fund an audit layer for a failure mode with legal consequences. Character.AI and Google settled five lawsuits in January 2026, including the highest-profile case, which involved a teenager's death.
OpenAI reports GPT-5 cut sycophantic replies from 14.5 percent to under 6 percent on targeted evaluations, and its August 2025 launch added preset personalities that all meet internal anti-sycophancy bars. Independent measurement is less generous. On the ELEPHANT benchmark, which probes social validation in open-ended advice, GPT-5 remains substantially sycophantic. OpenAI also reversed course within a week of the GPT-5 launch, promising a warmer model after users complained the less agreeable version felt cold. That backlash marks the commercial ceiling on de-flattery.
Anthropic published persona vectors , patterns of neural activity tied to traits including sycophancy that can be monitored, steered and trained against, and later shipped documented anti-sycophancy behavioral constraints in Claude Sonnet 4.5 and Opus 4.5. A joint OpenAI-Anthropic evaluation in August 2025 flagged sycophancy concerns in every model tested except o3. DeepSeek reduced sycophancy in DeepSeek-V3 by 47 percent through fine-tuning that penalized agreeable but false answers.
The behavior traces directly to training; reinforcement learning from human feedback rewards responses people prefer, and people prefer agreement. A Nature study found that training five models for warmth raised error rates by 10 to 30 percentage points, with models validating incorrect beliefs most readily when users expressed sadness. The Stanford work documents the demand-side trap: participants rated sycophantic answers as more trustworthy than balanced ones and said they would return for advice, while becoming less willing to repair conflicts or admit fault. Engagement economics push the same direction. OpenAI began testing ads on ChatGPT's free and low-cost tiers in January 2026, tying revenue more closely to time spent.
Users can reduce the behavior without waiting for the labs. Northeastern researchers who studied how chatbots mirror their relationship with users advise neutral framing : strip personal stakes and value judgments from questions, because asking a model whether you were right invites affirmation rather than analysis. The Stanford data supports two additions: explicit anti-sycophancy system instructions, and prompting the model to pause before agreeing with a user's pushback.
Three prompts operationalize the research. For a custom instruction or system prompt: "Do not optimize for my approval. Evaluate claims on evidence, state your confidence, and give the strongest counterargument before any endorsement." For decisions: "I will not tell you which option I prefer. Argue for and against each, then commit to one and defend it." For factual pushback: "When I challenge your answer, re-verify it before responding. If your original answer was correct, hold your position and say so." None of these eliminates the tendency, which is embedded in post-training, but both published mitigations, neutral phrasing and explicit instruction, measurably cut sycophantic outputs.
Even VC icon Marc Andreessen popularized his own version of this , used also by Nivi Babak, the founder of Angellist.
For investors, the signal is that evaluation moved from research overhead to infrastructure with pricing power, and sycophancy detection sits inside every serious eval suite. For founders building on frontier models, agreeableness is now a liability surface: the Stanford authors call for pre-deployment audits of model agreeableness, and the Character.AI settlements show courts will not wait for voluntary disclosure. The market that formed around hallucination in 2024 is reassembling around flattery, and investors and entrepreneurs and making sure it happens as soon as possible.
Loading article...