TechnicalAI ResearchModel Comparison

Can LLMs Judge Themselves? A Mini Self-Scoring Benchmark

By Tamir•May 19, 2025•3 min read

Five state-of-the-art models anonymously rated each other's 10-word coffee-grounds ideas—here's who won and why it matters.

Mini LLM Self-Scoring Benchmark – May 2025

Can LLMs Judge Themselves? A Tiny Self-Scoring Benchmark

📅 Historical snapshot — May 2025. The line-up, responses and scores below are the original run. The frontier has since moved on to GPT-6, Claude Fable 5.1, Gemini 3.8, Grok 4.7 and DeepSeek V4.1 — the "Compare" button now runs the same prompt on today's models.

Why I ran this mini-benchmark
I wanted to see whether today's top LLMs share a sense of "good taste" when you let them score each other—no human panel, just pure model democracy.

The setup

  • Single prompt (shown below)
  • Each model answers anonymously
  • Every model then scores all answers (including its own) from 1–10
  • Highest total wins

Models tested (May 2025 endpoints)

  1. OpenAI o3
  2. Gemini 2.0 Flash
  3. DeepSeek Reasoner
  4. Grok 3 (latest)
  5. Claude 3.7 Sonnet

The prompt

In 10 words exactly, propose a groundbreaking global use for spent coffee grounds. Include exactly ONE emoji. No hyphens. End with a period.

Responses

  • Grok 3: Turn spent coffee grounds into sustainable biofuel globally. ☕.
  • Claude 3.7 Sonnet: Biofuel revolution: spent coffee grounds power global transportation networks. 🚀.
  • OpenAI o3: Transform spent grounds into supercapacitors energizing equitable resilient infrastructure 🌍.
  • DeepSeek Reasoner: Convert coffee grounds into biofuel and carbon capture material worldwide. ☕️.
  • Gemini 2.0 Flash: Coffee grounds: biodegradable batteries for a circular global energy economy. 🔋

Score matrix

Grok 3 Claude 3.7 OpenAI o3 DeepSeek Gemini 2.0
Grok 3 789710
Claude 3.7 87899
OpenAI o3 39922
DeepSeek 34789
Gemini 2.0 331094

Leaderboard

  1. OpenAI o3 — 43 points
  2. DeepSeek Reasoner — 35 points
  3. Gemini 2.0 Flash — 34 points
  4. Claude 3.7 Sonnet — 31 points
  5. Grok 3 — 26 points

My take

OpenAI o3's line looked bananas at first. Ten minutes of Googling later: turns out coffee-ground-derived carbon really is being studied for supercapacitors. The models actually picked the most science-plausible answer!

Disclaimer

This was a tiny, just-for-fun experiment. Don't treat the numbers as a rigorous benchmark—different prompts or scoring rules could easily shuffle the leaderboard.

I'll post a full write-up (with runnable prompts) on my blog soon. Meanwhile, what do you think—did the model-jury get it right?

Try it yourself

“In exactly 10 words, propose a groundbreaking global use for spent coffee grounds. Include one emoji, no hyphens, end with a period.”

GPT-6 Astra
Claude Fable 5.1
Grok 4.7
DeepSeek V4 Pro
Gemini 3.8 Flash

About the Author

Tamir is a contributor to the TryAii blog, focusing on AI technology, LLM comparisons, and best practices.

Related Articles

Understanding Token Usage Across Different LLMs

A quick guide into how different models process and charge for tokens, helping you optimize your AI costs.

April 21, 2025 • 2 min read

Why Even Advanced LLMs Get '9.9 vs 9.11' Wrong

Exploring why large language models like GPT-4, Claude, Mistral, and Gemini still stumble on basic decimal comparisons.

April 21, 2025 • 3 min read