Meta introduces GAMUT benchmark to measure factual completeness in AI

Meta introduces GAMUT benchmark to measure factual completeness in AI

The new evaluation framework tests whether AI models actually tell you everything you need to know, not just whether they get individual facts right

Meta AI researchers have released GAMUT, a benchmark designed to measure something most AI evaluations quietly ignore: whether an AI’s answer is actually complete. Not just accurate, not just fluent, but whether it includes all the facts that matter.

The benchmark, short for Grounded Assessment of Multimodal Factuality, was published as an arXiv paper on July 21, 2026. It introduces a structured rubric system that converts “did the AI say everything it should have” into binary, machine-gradable checklists.

What GAMUT actually does

GAMUT tackles the problem of omission with a two-level meta-rubric framework. The system organizes required content hierarchically, meaning it doesn’t just list facts an answer should include. It structures them by importance and category, then translates those requirements into yes-or-no questions that language models can reliably grade.

Advertisement

The benchmark contains 1,813 questions rooted in real wearable imagery across ten distinct domains. Each question comes paired with expert-verified rubrics. Human experts defined what a complete answer looks like before any AI was tested against it.

The results are humbling

Meta evaluated 14 different AI models against the GAMUT benchmark. The top performer was Gemini 3.1 Pro, which scored 58.7%.

The benchmark demonstrated strong discriminative ability across the models tested, meaning it could meaningfully separate better performers from worse ones rather than clustering everything together.

Meta has made the benchmark resources publicly available through Hugging Face.

Why this matters beyond the lab

The multimodal aspect of GAMUT adds another layer. The benchmark’s questions are grounded in wearable imagery, meaning models need to interpret visual inputs and produce comprehensive textual responses.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Meta introduces GAMUT benchmark to measure factual completeness in AI

Meta introduces GAMUT benchmark to measure factual completeness in AI

The new evaluation framework tests whether AI models actually tell you everything you need to know, not just whether they get individual facts right

Meta AI researchers have released GAMUT, a benchmark designed to measure something most AI evaluations quietly ignore: whether an AI’s answer is actually complete. Not just accurate, not just fluent, but whether it includes all the facts that matter.

The benchmark, short for Grounded Assessment of Multimodal Factuality, was published as an arXiv paper on July 21, 2026. It introduces a structured rubric system that converts “did the AI say everything it should have” into binary, machine-gradable checklists.

What GAMUT actually does

GAMUT tackles the problem of omission with a two-level meta-rubric framework. The system organizes required content hierarchically, meaning it doesn’t just list facts an answer should include. It structures them by importance and category, then translates those requirements into yes-or-no questions that language models can reliably grade.

Advertisement

The benchmark contains 1,813 questions rooted in real wearable imagery across ten distinct domains. Each question comes paired with expert-verified rubrics. Human experts defined what a complete answer looks like before any AI was tested against it.

The results are humbling

Meta evaluated 14 different AI models against the GAMUT benchmark. The top performer was Gemini 3.1 Pro, which scored 58.7%.

The benchmark demonstrated strong discriminative ability across the models tested, meaning it could meaningfully separate better performers from worse ones rather than clustering everything together.

Meta has made the benchmark resources publicly available through Hugging Face.

Why this matters beyond the lab

The multimodal aspect of GAMUT adds another layer. The benchmark’s questions are grounded in wearable imagery, meaning models need to interpret visual inputs and produce comprehensive textual responses.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.