HalluHard: A Hard Multi-Turn Hallucination Benchmark
Introduction
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
Built with Maksym Andriushchenko, HalluHard tests whether frontier models can produce accurate inline citations across multi-turn conversations in high-stakes domains — legal cases, medical guidelines, and research. It separates reference grounding failures (the source does not exist or is wrong) from content grounding failures (the source is real but does not support the claim), and finds models hallucinating in over 30% of cases with web search enabled and over 60% without it.