Claude Opus 4.5 shows 30% on HalluHard - is that low enough?
https://extra-wiki.win/index.php/How_to_Documentation_AI:_Transforming_Multi-LLM_Conversations_into_Structured_Enterprise_Assets
If you have been tracking the latest release cycle, you’ve likely seen the discourse surrounding Claude Opus 4.5 and its performance on the HalluHard benchmark