Overview of Better Know a Benchmark: Humanity’s Last Exam
This episode of Linear Digressions explains Humanity’s Last Exam (HLE), a benchmark designed to be extremely difficult for large language models. The hosts break down what the benchmark measures, how it was assembled, why it initially exposed major model weaknesses, and why later criticism suggested that some answers in the dataset may themselves be flawed.
What Humanity’s Last Exam Is
Humanity’s Last Exam is meant to be a high-difficulty, expert-level benchmark for LLMs.
- Originally framed as “Humanity’s Last Stand,” it was intended to be a test that:
- human experts could answer,
- but models would struggle with.
- It includes about 2,500 public questions plus a private holdout set.
- Questions are closed-form and have one correct answer, often a number or other exact response.
- The benchmark spans many fields, including:
- math
- physics
- biology and medicine
- humanities and social sciences
- computer science / AI
- engineering
- chemistry
- other specialized topics
How the Benchmark Was Built
The benchmark was assembled as a large crowdsourcing and expert-curation effort involving academics and researchers.
Pipeline
- Experts submitted roughly 70,000 candidate questions.
- Questions were run through an LLM difficulty check:
- if models could answer them easily, they were removed.
- if models struggled, they moved forward.
- About 13,000 questions passed that stage.
- Those questions were then reviewed, refined, and peer-reviewed.
- About 6,000 survived further review.
- The final release became 2,500 public questions, with the rest kept private.
Why the private set matters
The hosts emphasize that a private holdout set helps reduce contamination/leakage, since public benchmark questions can end up on the internet and later appear in model training data.
Early Results: Why It Stood Out
At launch, HLE was dramatically harder than existing benchmarks like GPQA and MMLU, which were already becoming saturated.
- GPQA: models were getting roughly 60%–80%
- MMLU: models were at 80%+
- Humanity’s Last Exam: initial model accuracy was only about 3%–13%
The episode frames this as a benchmark that actually accomplished its goal: it was hard enough that even strong models performed very poorly at first.
Why It’s Easy to Grade
Because the benchmark uses single, exact answers, scoring is straightforward:
- either the model gets the answer right,
- or it doesn’t.
That makes HLE much easier to evaluate than open-ended benchmarks where graders must interpret subjective responses.
Current Performance and Why It’s Hard to Compare
The hosts note that model performance has improved a lot since launch. Roughly a year later, top models are said to be reaching around 50%–65% on some leaderboards.
However, they stress that different benchmark sites report different scores for the same model.
Examples mentioned
Across three leaderboards, the same models can differ significantly:
- one model might score around 40% on one leaderboard and 58% on another
- another model might differ by 10+ points depending on the evaluator
Why the numbers differ
Possible reasons include:
- different question subsets used
- public vs. private set differences
- replacement of questions with alternates
- model nondeterminism
- different model versions or release dates
- different effort settings
- other unobserved evaluation differences
The takeaway is that benchmark scores should be interpreted carefully, even when they look precise.
Criticism and Potential Errors in the Benchmark
A major segment of the episode covers a follow-up critique from Future House, a company building AI research agents.
They investigated the chemistry and biology parts of HLE and argued that around 30% of those answers may be wrong.
Example discussed
One question asked about the “rarest noble gas on Earth as a percentage of all terrestrial matter in 2002,” with the answer allegedly being oganesson.
Future House argued that this is likely flawed because:
- oganesson is synthetic and extremely short-lived
- it may not really count as a gas
- it may not be noble in the relevant sense
- it likely wasn’t treated as terrestrial matter in the cited literature
Why this matters
The episode’s key insight is that the benchmark’s LLM difficulty filter may have introduced bias:
- questions that models got “wrong” may have been incorrectly labeled by humans
- in some cases, the model may have been right and the benchmark answer wrong
- this creates a selection effect that favors difficult questions with disputed or erroneous answers
The HLE creators reportedly responded by adding more peer review and releasing a revised version.
Main Takeaways
- Humanity’s Last Exam is a deliberately hard benchmark built to push beyond saturated tests like MMLU and GPQA.
- It uses closed-form, expert-level questions across many disciplines.
- The benchmark initially exposed a major gap between humans and LLMs, with models scoring in the low single digits to low teens.
- Model performance has improved substantially, but reported scores vary a lot across leaderboards.
- A key lesson is that benchmark construction matters:
- filtering for “questions LLMs miss” can unintentionally preserve wrong answers
- even elite benchmarks can contain flaws
- HLE remains important because it is one of the stronger tests of frontier model capability and calibration.
Final Recommendation from the Episode
When you see HLE in a paper or leaderboard, the hosts suggest paying attention to:
- the reported score,
- whether the benchmark used the public or private set,
- and whether the evaluation methodology is clearly described.
In short: HLE is a valuable benchmark, but not a perfect one—and its quirks are part of what make it scientifically interesting.