LLM As A Judge

Summary of LLM As A Judge

by Ben Jaffe and Katie Malone

30m•September 28, 2026

Overview of LLM As A Judge

This episode of Linear Digressions explains the idea of using large language models (LLMs) to evaluate the outputs of other LLMs or AI systems—known as “LLM as a judge.” The hosts explore why this approach emerged, how it works in practice, what the research initially found, and why it’s useful but far from foolproof.

What “LLM as a Judge” Means

The core problem is that AI systems are hard to evaluate without lots of labeled data, and human labeling is expensive and slow. Instead of asking humans to review every output, you can ask an LLM to judge whether an answer is:

  • helpful
  • accurate
  • coherent
  • safe/harmless
  • aligned with a rubric
  • better than another model’s response

This is especially useful for:

  • chatbot evaluation
  • comparing model outputs head-to-head
  • judging agent behavior, including tool use and reasoning
  • creating benchmarks at scale

The Original Research: Judging LLMs with LLMs

The discussion centers on the 2023 paper Judging LLMs with MT-Bench and Chatbot Arena, which helped formalize the idea of using LLMs as evaluators.

Main evaluation styles covered

  1. Pairwise comparison

    • Show the judge two responses and ask which is better.
    • This is the most common “horse race” setup.
  2. Single-answer grading

    • Give the judge one response and ask how good it is.
  3. Reference-guided grading

    • Provide a rubric, examples, or reference answers to make the judgment more structured.

Example discussed

The hosts walk through a simple economics question where one model gives a vague, incorrect answer and another gives a better, more contextual response. GPT-4, used as a judge, correctly prefers the stronger answer—suggesting that in many cases, LLM judges can align well with human preference.

Key Findings From the Early Benchmarks

The episode highlights several important results from early LLM-judge research:

  • GPT-4 often agrees with humans
  • In clear cases, agreement was roughly 80–90%
  • When you include:
    • ties
    • positional bias
    • ambiguous cases agreement drops to around 60–65%

This matters because evaluation numbers can look much better or worse depending on how you handle ambiguity and bias.

Chatbot Arena and Real-World Data Collection

A major practical example is Chatbot Arena:

  • Users type in prompts
  • Two anonymous models respond
  • Users vote on which response is better

This creates a real-world dataset of preference judgments that can be used to benchmark models. The episode notes that this setup is both:

  • a fun interactive tool
  • a powerful source of evaluation data

It also exposes users to different models and styles they may not otherwise try.

Major Limitations and Biases

The episode spends significant time on the caveats. LLM judges are useful, but they are not neutral or perfectly reliable.

Common issues discussed

  • Positional bias: preferring the first or second answer depending on order
  • Self-preference bias: preferring outputs from the same model or model family
  • Length bias: sometimes favoring longer answers
  • Non-determinism: the same prompt can produce different judgments on different runs
  • Overconfidence: models may act more certain than they should
  • False stability: highly consistent judgments are not always correct

A particularly important point: even if a judge is consistent, that does not mean it is valid.

Human Agreement Is Not Perfect Either

The hosts emphasize that humans themselves often disagree, which means “ground truth” is not always truly ground truth.

This connects to inter-rater reliability:

  • if humans disagree on a task, the task itself may be subjective or ambiguous
  • raw agreement can be misleading because some agreement happens by chance
  • for many tasks, the best you can hope for is approximate alignment with human judgment

Practical Advice for Using LLM Judges

The episode ends with pragmatic guidance for teams considering LLM-based evaluation.

Good practices

  • Measure how much humans agree with each other first
  • Test judge stability by:
    • flipping response order
    • rerunning the same judgment multiple times
  • Be cautious interpreting small differences
  • Treat the judge as noisy, not authoritative
  • Use human labels to calibrate the system

Choosing a judge model

The best choice depends on the use case:

  • For creating training or benchmark labels

    • use the strongest available model if you can
  • For production monitoring

    • a smaller model may be fine if the task is well-specified and checkable
  • For open-ended, holistic judgments

    • stronger models are usually better

In general:

  • if the task is easy to verify, smaller models can work well as judges
  • if the task is hard to judge holistically, use the best judge you can afford

Main Takeaways

  • LLMs can be useful as cheap, scalable evaluators of other AI systems.
  • Pairwise comparison is the most common and often most effective setup.
  • Early research showed strong agreement with humans in many clear cases.
  • But bias, ambiguity, and nondeterminism can significantly distort results.
  • Human disagreement means evaluation tasks are often inherently fuzzy.
  • LLM judges should be treated as noisy approximations, not ground truth.

Bottom Line

LLM-as-a-judge is a powerful tool for scaling evaluation when human labels are scarce, but it is not a drop-in replacement for human judgment. The key is to understand the task, measure noise and bias, and use the right model and evaluation setup for the job.