← Back to Forum

HalluHard: A Hard Multi-Turn Hallucination Benchmark

Dongyang Fan
March 5, 2026
Introduction

This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.

Built with Maksym Andriushchenko, HalluHard tests whether frontier models can produce accurate inline citations across multi-turn conversations in high-stakes domains — legal cases, medical guidelines, and research. It separates reference grounding failures (the source does not exist or is wrong) from content grounding failures (the source is real but does not support the claim), and finds models hallucinating in over 30% of cases with web search enabled and over 60% without it.

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.