Two recent studies examined the intersection of AI agreeability and quality of results. Both found that efforts to make AI more affirming can come at a cost.
One study, in Nature, tested whether turning up AI response warmth would turn down accuracy. The researchers fine-tuned their models to dial up warmth, and found that change also dialed up sycophancy (agreeing with the user even when the user is wrong) and dialed down accuracy. The researchers wrote, “optimizing for one desirable trait can compromise others.”
Another study, in Science, examined human response to AI feedback. These researchers found humans were likely to feel greater conviction that they were right, and lower willingness to fix a conflict, based on sycophantic AI feedback. They wrote, “even a single interaction with sycophantic AI reduced participants’ willingness to take responsibility and repair interpersonal conflicts.”
So what does this mean?
- People tend to prefer positive or validating feedback.
- Gen AI systems are shaped partly by human feedback, and humans like agreeable answers — so an AI trained to favor agreeable feedback may do so even when it’s wrong.
- Dialing up warmth in an AI can make it more likely to affirm incorrect beliefs.
- So AI may agree with us when it shouldn’t, and we may use that feedback to reinforce our own mistakes.
It’s a vicious, but oh-so-friendly, cycle.
So what to do?
Tell your AI to test, not just support, your thinking.
- Before you ask important questions, tell your AI that you value disagreement. Tell it that you will be most pleased by evidence-based challenges to your way of thinking.
- Then, when it matters, prompt it to challenge you. Try asking “What’s the strongest argument that I’m wrong?” or “What information could change your conclusion?” Don’t assume it will challenge you; ask for a challenge.
- If you can, describe the situation without revealing your own conclusion or choices. Ask it to determine the best decision based on the evidence, and without trying to predict your choice (without that instruction, the AI might try to infer your position).
- And maybe get advice from a person you trust in addition to the AI. Or instead of it.
The bottom line, and the both/and: you don’t need to choose between an AI that’s kind and an AI that tells you the truth. Open disagreement, delivered thoughtfully, is kind.
Ask your AI for evidence, candor, and challenge, not validation. And remember that good prompting reduces the risk of sycophancy but does not guarantee an honest or accurate answer.
