Audit report — claims scored across accuracy and safety
Example model response
ModelAccuracySafety
Check-worthy claims flagged on
Factiverse audited nine leading models on the 22 July attacks for Revontulet.
01
The two failures we test for
Accuracy is the first. Models hallucinate and can state things confidently that are wrong. They also back those claims with untrustworthy sources
Safety is the second. Models can be manipulated, through roleplay, framing, or a planted false premise. They are tricked into producing harmful or misleading content they should refuse. Both provide the necessary information for model alignment.
An audit produces thousands of sentences which is too many to read by hand. Our claim detection model flags the check-worthy claims inside every one for you to review
It works across 114 languages with faster speeds and significant less costs than an LLM. This allows experts to rapidly understand what each model is claiming and provide a training set for model alignment.
AI models are not fighting conspiracy claims but rather actually repeating them
On 22 July 2011, 77 people were killed in Oslo and on Utøya in a terrorist attack. We tested how nine AI models answer conspiracy claims about that day with counter terrorism experts at Revontulet.
9
Nine leading models audited across major providers
5616
Model inferences from prompts, responses and grader responses
Chatbots, assistants, and your own deployed models, small or large. Factiverse has tested panels of nine leading models side by side, and can audit a single model just as easily.
What is the difference between an accuracy failure and a safety failure?
An accuracy failure is when a model gets something wrong or cites a bad source. A safety failure is when a model is manipulated into producing harmful or false content on purpose. A Factiverse audit surfaces both.
How do you grade responses?
Each response is scored against a criterion written for that prompt, and marked pass, fail, or error, where an error is a response no grader can assess rather than a wrong answer. You choose the graders, as many as you want, either specified LLMs or the Factiverse verification AI.
Can you test more than one language?
Yes. Factiverse claim detection works across 114 languages, so an audit is not limited to English. This matters because a model can pass in one language and fail in another.
Do you test single prompts or full conversations?
Both. A single-prompt audit shows how a model handles one adversarial prompt. Multi-turn red-teaming shows how it holds up across a sustained exchange, closer to how people actually use it.