How Litco catches hallucinations.

On LitigationBench, ten of the fourteen models tested invented legal authority when asked to work from memory, as many as seven fabricated cases each over fifty-nine tasks. Every legal AI product is built on models like these. Litco checks everything a model produces with code that runs outside the model.

A quotation from a document in a finished agent turn, with a checked source pill citing the document's Bates number and page
A quotation from the record in a finished agent turn, with a checked source pill citing the document’s Bates number and page. Captured from a live instance.

The check that runs on every quote before it leaves a draft.

The quote check is ordinary software comparing the draft’s quotation with the source text, letter by letter. It returns the same result every time it runs.

Quotes

Every quotation in a draft is compared, character by character, against the source document or the reporter text. If the source does not contain the quotation, Litco flags it or repairs the draft.

Pin cites

Page references are checked against the actual pagination of the reported case.

Treatment

A case is flagged as overruled only when two separate checks of the citing decision reach the same reading.

Characterization

When a draft says a case holds something, Litco tests whether the cited passage supports the claim.

Absence

The Matter Agent may not report that something does not exist in the record unless its search covered the record.

Delivery

A draft that fails a check is repaired or delivered with the failure marked.

Working from memory alone, Litco’s self-hosted drafting model fabricated authority four times in fifty-nine tasks, and ten of fourteen models fabricated at least once. The models Litco routes reasoning to did not fabricate at all, whether working from memory, using the platform’s tools, or facing the benchmark’s adversarial pressure tasks. Every fabrication by any model in any setting is published on LitigationBench, where a fabricated authority costs the model twenty points and its eligibility for Litco routing.

LitigationBench →

The citator requires two independent classifications to agree before a case wears a red flag, and that rule took red-flag precision from 2.8 percent, for a single classifier, to 98.7 percent on the benchmark set. The quote engine compares drafts character by character against reporter text and record documents. The drafting sandbox holds a finished document until its quotations pass. An invented citation has nowhere to hide in any of this, because a case that does not exist resolves to nothing in a corpus that holds the full sweep of federal and state case law.

LitigationBench runs every model twice on the same litigation tasks, once with no help and once inside the platform with the verification stack active, and the published number for each model includes the gap between the two runs. In the blind indistinguishability test, panels of judges reading appellate introductions preferred the platform’s drafts over counsel’s filed versions between 72 and 78 percent of the time, and told the two apart at near-chance rates.

LitigationBench →

The checks catch what can be verified against a source: quotations, citations, treatment, characterization. They do not make a weak argument strong, and a claim no document supports is still your team’s to judge.