Photo: Steve A Johnson / Pexels
Arena study shows AI models prefer their own answers 58% of the time
New research into automated judging reveals that large language models exhibit a persistent narcissistic streak, raising questions about the leaderboards that shape AI development
When you ask an AI to judge a writing contest and one of the contestants is itself, things get awkward. New research examining self-preference bias in large language models finds that AI judges favor their own generated responses roughly 58% of the time in pairwise comparisons.
The findings cut to the heart of a growing dependency in the AI industry: using models to evaluate other models. Platforms like Chatbot Arena, which pit AI systems against each other in blind head-to-head matchups, have become de facto benchmarks for measuring progress.
The narcissism problem
The bias, sometimes called egocentric or narcissistic bias, shows up consistently across the industry’s biggest names. OpenAI’s GPT, Anthropic’s Claude, and Google’s Gemini all exhibit the tendency when placed in judging roles.
Research suggests that somewhere between 40% and 96% of the observed self-preference could stem from actual quality differences in the outputs being compared. In other words, some models might genuinely prefer their own answers because those answers are, by certain metrics, better.
One way researchers have tried to tease this apart is through perplexity, a measure of how surprised a model is by a given text. Models that produce lower-perplexity outputs tend to score higher on quality benchmarks, and those same models tend to show stronger self-preference.
Another notable quirk: AI judges rarely declare ties. In human evaluations, ties are common. AI models seem almost allergic to this outcome, consistently picking a winner even when the quality gap between two responses is negligible.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Why leaderboards should make you nervous
Self-preference bias complicates that assumption in a specific way. If the models doing the judging are also among the models being judged, there’s a structural conflict of interest baked into the process.
Critics of Arena-style evaluation have also pointed to issues of data asymmetry. AI providers whose models are used as judges may have inherent advantages, whether through familiarity with their own output patterns or through subtle alignment between their training data and their evaluation criteria.
The self-preference effect also appears to intensify as model capability increases. More powerful models show stronger tendencies to favor their own outputs.
What comes next
Researchers are planning focused work over 2025 and 2026 to disentangle legitimate self-preference from harmful bias. The goal is to build datasets, potentially drawing from platforms like Chatbot Arena, that can help distinguish between a model preferring its own output because it’s genuinely better and a model preferring its own output because it recognizes familiar patterns in the text.
That distinction matters enormously. If 90% of self-preference reflects real quality gaps, the bias is annoying but manageable. If only 40% does, then automated evaluation systems need significant redesign before they can be trusted as arbiters of model quality.
One potential mitigation involves using panels of diverse models as judges rather than single evaluators, diluting any individual model’s self-preference across the group. Another approach explores calibrating for known biases after the fact, applying statistical corrections the way pollsters adjust for sampling errors.