SciTech Pulse
AI Models

New Tool Checks What AI Benchmarks Are Actually Measuring

Allen Institute for AI's BenchMIRT tool found that some AI safety benchmarks, like BBQ and WMDP, actually track general reasoning more than safety.

Researchers at the Allen Institute for AI (AI2) have introduced BenchMIRT, a new method for checking what AI benchmarks are actually measuring at the level of each individual question. A is a standardized test used to score how well an AI model performs on a specific ability, such as safety or general reasoning. AI2 found that individual questions inside a benchmark can depend on more than its stated goal. Averaging very different kinds of questions into one score can hide what is really driving it.

BenchMIRT is based on , a technique from psychometrics, the field that measures abilities from patterns of test answers. It builds on the idea that not every question reveals the same amount about the model being tested. AI2 extended this into a multidimensional version that can separate several abilities measured within the same set of questions. The researchers trained BenchMIRT on the results of 100 large language models across 16 benchmarks and more than 34,000 questions, without telling it in advance which benchmarks were meant to measure which ability. The tool independently and consistently identified two dominant dimensions on its own: safety and general reasoning.

For most benchmarks, BenchMIRT confirmed their intended purpose. Reasoning benchmarks tracked with reasoning ability, and benchmarks testing jailbreak or harmful-content refusal tracked with safety. But it also complicated the picture for some. BBQ, a benchmark commonly grouped with safety tests because it checks whether models rely on social stereotypes, aligned more strongly with general reasoning, suggesting a low score there may partly reflect difficulty reasoning through a question rather than unsafe behavior. WMDP, a benchmark that tests dangerous dual-use knowledge in fields such as biology, chemistry and cybersecurity, also aligned more with reasoning than safety, because stronger-reasoning models were simply better at recognizing and refusing to answer such questions.

BenchMIRT also found that a single benchmark can mix different signals. Within HarmBench, which tests whether a model will comply with harmful requests, most questions aligned with safety, but its questions about reproducing copyrighted material, such as song lyrics, aligned instead with general reasoning. The researchers said the findings do not mean the benchmarks are flawed, but that a single benchmark score can combine several different signals that BenchMIRT can help separate and make easier to interpret.

Terms explained

The story so far

  1. Robotis Hands Its Humanoid Robots to Korean University Students
  2. Nonprofit Plans Interstellar Launch on a Trajectory an AI Discovered
  3. New Tool Checks What AI Benchmarks Are Actually Measuring
#ai#LLM benchmarks#Allen Institute for AI#AI2
Rate this story

Related stories