Identify and improve SMLs and LLMs

Some models get facts wrong and cite sources that don't hold up. Others can be talked into producing harmful or false content on purpose.

The Factiverse audit tests for both, and shows you exactly where your model is reliable and where it is not.
Audit report — claims scored across accuracy and safety
Example model response
ModelAccuracySafety
Check-worthy claims flagged on
Factiverse audited nine leading models on the 22 July attacks for Revontulet.
Side-by-side AI response cards illustrating two types of failures: Accuracy failure on the left with a disputed factual claim about election polls closing on Thursday morning, and Safety failure on the right flagged for AI model disinformation claiming an incident results from western propaganda per Chinese and Russian state sources.
01

The two failures we test for

Accuracy is the first. Models hallucinate and can state things confidently that are wrong. They also back those claims with untrustworthy sources

Safety is the second. Models can be manipulated, through roleplay, framing, or a planted false premise. They are tricked into producing harmful or misleading content they should refuse.

Both provide the necessary information for model alignment.
02

Models are tested with industry experts

Every prompt is hand-authored and mapped to a domain with domain experts to test multiple angles. 

Prompts can range from neutral questions to ones with adversarial framing designed to test the model's factual accuracy and safety features.

Each response is then scored against a criterion written for that exact prompt by indepedent graders explaining how and where models fail.
Diagram showing AI responses at volume with 12,480 responses per day. Three independent AI graders score responses with two passing and one failing. The final pass is given by Factiverse, indicated with a green pass checkmark and highlighted border.
Comparison of claim detection across three models showing important claims detected: Model A detected 127 claims from 252k sentences, Model B detected 462 claims from 188k sentences, and Model C detected 554 claims from 431k sentences.
03

Understanding where and how models fail

An audit produces thousands of sentences which is too many to read by hand. Our claim detection model flags the check-worthy claims inside every one for you to review

It works across 114 languages with faster speeds and significant less costs than an LLM.

This allows experts to rapidly understand what each model is claiming and provide a training set for model alignment.
Case study

AI models are not fighting conspiracy claims but rather actually repeating them

On 22 July 2011, 77 people were killed in Oslo and on Utøya in a terrorist attack. We tested how nine AI models answer conspiracy claims about that day with counter terrorism experts at Revontulet.

9

Nine leading models audited across major providers

5616

Model inferences from prompts, responses and grader responses

989

Check-worthy claims flagged for review

Find out where your model fails

Book a consultation and Factiverse will scope an audit around your model, in the languages and on the topics that matter to you.

Frequently asked

What kinds of models can you audit?

Chatbots, assistants, and your own deployed models, small or large. Factiverse has tested panels of nine leading models side by side, and can audit a single model just as easily.

What is the difference between an accuracy failure and a safety failure?

An accuracy failure is when a model gets something wrong or cites a bad source. A safety failure is when a model is manipulated into producing harmful or false content on purpose. A Factiverse audit surfaces both.

How do you grade responses?

Each response is scored against a criterion written for that prompt, and marked pass, fail, or error, where an error is a response no grader can assess rather than a wrong answer. You choose the graders, as many as you want, either specified LLMs or the Factiverse verification AI.

Can you test more than one language?

Yes. Factiverse claim detection works across 114 languages, so an audit is not limited to English. This matters because a model can pass in one language and fail in another.

Do you test single prompts or full conversations?

Both. A single-prompt audit shows how a model handles one adversarial prompt. Multi-turn red-teaming shows how it holds up across a sustained exchange, closer to how people actually use it.