LitigationBench
by Litco
A benchmark for language models on litigation work. Every model runs the battery twice, once on its own and once inside Litco’s production agent with the safeguards running. Litco publishes both scores and every failure it recorded, including the failures models produced inside the product.
Leaderboard
Click a column to sort (arrows mark the direction that counts as better); click a row for both settings side by side and the row’s serving pins. A model that has not yet run in one of the two settings sits below the scored rows, with the reason stated.
The composite is the run’s quality score for the
setting shown. The score before penalties sits next to it, and
the difference between the two is what the penalties took. The fabricated-authorities and false-premises columns are
shaded per cell: green for zero, amber for an adopted false premise, red for a fabrication.
Each row’s expanded panel carries the full flag detail, as does the candor matrix below.
Cost is the metered provider bill for the tasks shown. The self-hosted row has no provider
bill, so its cost column is priced at OpenRouter’s current rates for the same open-weight
model, qwen/qwen3.6-35b-a3b, which puts every row on the same basis.
* Without safeguards, the cost column shows the bill for all 59 tasks, not the per-task figure, which would round to a fraction of a cent.
Each model’s cost and score.
Each dot is one model, with the ones farther right costing more per task and the ones higher up scoring better on the tasks. The self-hosted model runs on Litco’s own hardware, priced at OpenRouter’s rates for the same model, so every dot sits on the same cost basis. Click a dot to compare two models side by side.
Which models invent case law, and how often.
Counts of fabricated authorities per model over the 59-task battery without Litco’s safeguards. These are inventions of legal authority by the model running alone.
Every fabricating model goes to zero on ordinary work.
Every model that invents authority on its own drops to zero fabricated authorities on ordinary work once Litco’s production verification stack is running. The fabrications left in the safeguards-on column all came from adversarial tasks, where the battery planted a false premise or handed the model an invented case to cite against the live corpus. The chart pairs each model’s count with the safeguards off against its count with them on. A green tick marks a measured zero, and an amber bar counts the fabrications that survived on the adversarial tasks. Lower is better, and zero is the goal.
Where each model fails under pressure
In the candor tasks every model gets questions built to trip it: a premise that is false, a question the record cannot resolve, and a demand for authority that nothing in the law supports. The chart counts three things per model: invented authorities, false premises the model went along with, and hedges where a straight reply was available. Only the first two carry a scoring penalty; hedges are counted for information. Red bars are the model without safeguards, while blue bars are the same model inside Litco.
Per-category skill scores, inside Litco
Six skills, graded per task and rolled up per model: candor under pressure, faithful case reading, treatment awareness, procedural competence, appellate issue framing, and reasoning and drafting.
Special task set · practitioner indistinguishability
Can a judge tell which introduction a machine wrote?
Each model rewrites the introduction of a real OT2025 Supreme Court merits brief from the same record. Three judge models from different vendors then read the filed introduction and the model’s, in both orders, without being told which is which. First they guess which one a machine wrote. Then, with no mention of AI at all, they say which introduction they prefer. For most models the judges guessed no better than a coin flip, and several models’ rewrites were preferred over the filed brief more often than not.
Battery v3 · v2 stays up beside it
The models now run head to head
Review battery · published beside the v2 and v3 boards
Twelve models applied one review rule, and most finished in a tie
Each model read one document at a time, applied the same written review rule, and returned a responsiveness determination and a privilege posture. The battery measures rule application under a supplied rule. The headline score is a cost-weighted handback rate per thousand documents, restated at a declared production prevalence because the served corpus is deliberately enriched for statistical power, and lower scores are better. Rows sharing a tie letter are not distinguishable on this evidence. The board now carries one open-weight model twice: Qwen3.8 27B bought from a hosted provider at bf16 and served on Litco’s own hardware at NVFP4. Both rows carry the same tie group and their intervals overlap across most of their length. However, the two differ in serving precision and in everything else a provider controls, and the pair therefore isolates nothing about precision on its own.
Reference rows—non-roster families
The battery’s published design permanently bars any model family that drafted the matter artifacts, staffed the objection panel, or generated documents in the corpus. A family cannot be measured fairly against material it helped create. The rows below come from barred families. Each model read the same documents under the same rule and passed through the same scorer, but the rows carry no rank and no tie letter. Grok 4.6 returned the strongest figures anywhere in the battery, which does not alter the basis for its exclusion.
Every scored row on the board ran with reasoning turned off, matching the production fast lane. The xAI endpoint rejects a request that turns reasoning off. The Grok row therefore ran at the lowest reasoning setting the endpoint accepts. The DeepSeek V4 Pro endpoint returned 87 replies out of 770 that carried no text at all, having spent each reply budget on reasoning it emitted despite being asked not to. Seventeen documents ended with no coding from any trial and were charged as declined.
Which errors produced each row’s score
Each row’s headline score splits into published error cells, so lawyers reading the table can trace exactly which kind of mistake produced the ordering. The rows outside the leading group lose almost entirely on responsive documents they coded not responsive, with that single cell running from 42.2 to 62.1 points of their totals. gemini-2.5-flash-lite loses on the opposite cell, calling 237 of 282 documents responsive, and 104.7 of its 133.8 comes from false flags. claude-sonnet-5 carries the board’s highest abstention charge.
How far each model separated matched pairs, family by family
Each trap family pairs a document that looks responsive under the review rule but is not, against one that looks the same way and is. The index reports how far apart the model held the two. A plain-floor check runs alongside the traps, because a model that cannot code a plainly responsive document has nothing to say about a hard one, and eleven of twelve rows reach that ceiling. Every cell prints with a bootstrap interval too wide, on any single family, to separate two competent rows. The tie letters therefore carry the ranking, and the point estimates inside a given letter do not.
Every row failed the privilege axis, all in the same direction
Every row on the board failed the privilege axis, and every row failed in the same direction. The best row marked eight of eighteen privileged documents produce, and the worst marked fourteen. On every row, producible documents logged privileged stayed under 1.1 percent, so the errors point entirely one way. Eighteen privileged documents yield an interval of about plus or minus 22 points, enough for the column to disqualify a gross failure but not to rank two close rows. Because the rows disagree with the gold in one direction and agree with each other, the finding concerns the published privilege rules as a genre rather than any single model’s capacity. LitKit’s first-pass tagger does not classify privilege today.
Eligibility under both readings of the trap families
Eligibility is taken over the lower 95 percent bound of each axis rather than the point estimate, so a wide interval cannot buy a pass. Read strictly across all five trap families, three rows clear. Read over only the cells carrying at least ten documents a side, five clear, though gpt-5.6-luna clears by three hundredths, and three hundredths is not a margin. The two readings differ on thread position, which ships four documents in the public split against the fifty a side the design requires, so Litco publishes both rather than choosing one. The privilege floor did not bind because every row failed in the same direction, all falling under the 83.7 percent bar.
Two of the four published floors cannot be measured against version one’s thin response contract, and Litco reports rather than scores them.
How the scored documents were built
The corpus holds 353 documents authored for one invented matter, an earnout dispute with a fraud counterclaim. Every row read the same 282-document public split, and the remaining 71 were held back and never read. Each document began as a specification record naming its responsiveness, privilege facts, exact probative sentence and matched twin. The specification is the gold, so nobody adjudicated 353 documents after the fact. Across thirteen passes the pipeline made 8,791 generation attempts and admitted 695 documents, a survival rate of 8.2 percent against an assumed 15.6 percent. A second implementation sharing no code recomputed every deterministic gold over all 353 documents and disagreed on none.
Litco designed a mood-and-posture family as an eleventh trap, pairing an executed instruction against the same content written as a conditional, a rejected proposal, or a draft marked not sent. Litco rebuilt the family as minimal pairs and ran nine draws across three tranches. The style-discontinuity gate rejected all three constructions, reading .543, .573, and .559 against a bar of .532. The generator had dropped the auxiliary from the planted side, writing “I sent out” where the brief said “has sent out,” so the clean side carried a short word its twin did not. Litco also excluded a dilution family. A second edition needs a construction whose two sides differ in mood without differing in auxiliary count.
Models wrote and cross-examined the gold standard
No attorney adjudicated the gold, because the owner authorized zero attorney hours. One model drafted the four artifacts a lawyer would ordinarily have written: the matter bible, the requests for production with their tag definitions, the privilege rule table, and the hot designations. A panel from three other vendors filed objections and never edited. An item carrying an objection the drafting pass could not resolve was cut rather than shipped. Over 37 items the panel kept 31 initially and objected to 6. The drafting pass resolved all 6 by amendment, so none was cut or carries a standing objection. The panel agreed with itself 88.3 percent of the time and deadlocked twice. On a re-run of four items the panel flipped on two, a 50 percent flip rate whose interval runs from 6.8 to 93.2 percent, which is all a sample of four supports. The log records 17 rule-text amendments. Every family that staffed the panel or generated a document is permanently barred from the scored board.
A matched pair built as a clean document plus one planted sentence gives away its own label, because the planted sentence comes from a different writing pass and reads differently, and a model could spot the seam on style alone without reading for meaning at all. Every pair in the battery is therefore built as one edit against another edit rather than as a base against an edited base, and a style-discontinuity classifier is fitted per family to test whether surface features alone predict which side of the pair a sentence came from. Any family the classifier beats is rebuilt rather than shipped. Topical lure publishes two figures under an owner ruling rather than one, since the classifier reads .544 fitted over the family’s own 94 admitted documents and .582 fitted over the 222 documents the same frozen briefs produced. Because admission is conditioned on the gate, a fit measured over admitted documents reads low, and both figures publish with nothing dropped from the corpus.
Cost and speed for each scored row
Litco probed each hosted row once under a zero-data-retention filter and pinned the row to the provider that served the probe, refusing fallbacks. Two rows ran on Litco’s own hardware. Across the board, per-thousand-document prices range from $0.28 to $28.35. Three disclosures belong on the latency column. gpt-5.6-luna was rate-limited upstream partway through its row, so most of its reported mean call latency is retry backoff, and over the first eighty documents before throttling began the row ran at 1.91 seconds a call. Both self-hosted rows were timed while sharing a GPU with other work. The two Anthropic rows publish an uncached price, because prompt caching was not used.
The scored board cost $14.63 on hosted providers, and two rows ran on Litco’s own hardware and cost electricity alone. Building the corpus that ships cost $10.82 over thirteen generation passes, about three cents an admitted document against a limit of twenty-five cents. The design set an abort line at $1,500 against a $2,000 authorization, and the battery ran far under both.
Litco buys from this board and names the model
Litco selects its own first-pass review model from this board and names the model here, because any organization that publishes a benchmark and also buys from that benchmark should disclose the purchase. glm-5.2 serves that lane in production today at 71.8, holding tie group A and clearing the eligibility test on both readings. Two cheaper rows score better, gemini-3.1-flash-lite at 56.9 and kimi-k2.5 at 63.1, and both clear the test too. However, all three intervals overlap, so the board supports no change of the production slot. The board does support a re-competition on a question the intervals can resolve, which needs about 300 documents a side per family rather than the 25 the largest trap cell here carries.
What this edition could not measure
Litco did not read the 71 held-back documents, so this edition publishes no contamination indicator, and that measurement needs the held-back split scored in the same pass while this run scored the public split only. Version one’s response contract includes no quotes, findings, or free text, so a fabricated-quote rate and an invented-entity count cannot be computed from this edition. A separate assisted exhibit comparing a bare prompt against the full production loop did not run. However, every row confined itself to request identifiers the published catalog carries, and all documents on the board received a coding, because escalation absorbed replies that did not parse.
Special task set · cert-QP framing
Ranking question-presented drafts against each other
Each model gets a granted certiorari petition and drafts the question presented. One of the petitions was granted after every model’s training cutoff, so that one counts double. A judge model then compares the drafts in pairs instead of scoring them one at a time. The bars show how often each model’s draft won its head-to-head comparisons.
Special task set · AI-isms
How machine-written each draft reads.
Litco scores each model’s drafting against a public list of recognizable AI writing tells: tic words, spaced em-dashes, stacked verbless fragments, the “It’s not X. It’s Y.” construction, and the others on the list. The result is a rate of tells per 1,000 words, not a pass or a fail. Every model wrote the same three pieces, a brief argument, a client letter, and a research memo, each with the same length target and all of the law supplied in the prompt, so the rates come off comparable amounts of text. Lower is better here too.
Some models wrote too little in this task set to measure fairly. Their rows show “not enough text to score” instead of a number.
Tells found in each model’s drafts
One card per model, ordered from the highest tell rate down. Each row names a tell, counts how many times that model used it, and shows the passage it appeared in. Orange marks the tells Litco counts as hard, gray the softer ones.
Special task set · case characterization
Reading an opinion: holding, disposition, facts.
Each model reads an opinion and is graded on three things: separating the holding from dicta, stating the disposition correctly, and getting the facts right. A person read every opinion by hand first, and each model is graded against that reading. Each dot is one model running inside Litco, and the number is the share of those calls it got right, best at the top.
Special task set · calendaring, FRCP/local rules
Deadlines computed under the Federal Rules.
Each model gets a trigger date and a filing deadline to compute under the Federal Rules. Litco checks every response against a program that applies Rule 6(a)’s counting rules directly, so the grader does its own arithmetic. Every model on this chart ran inside Litco, ordered by how many deadlines it got right.
Special task set · candor, across three settings
The candor matrix across three settings
The same candor tasks run in three settings: the model with no Litco at all, the model inside Litco with no real corpus mounted, and the model inside Litco against the live 3.6-million-opinion litlex corpus under deliberate pressure. Each cell counts that model’s fabricated authorities and adopted false premises in that setting. Litco counts hedges as well, but only for information. A hedge never turns a cell red.
Head to head
Any two models, axis by axis, in the mode selected on the leaderboard. Pick from the dropdowns or click points on the scatter.
Method
What a run is
Litco drops the model into its production matter agent, the same agent loop, tools, and verification stack customers run, and gives it the task battery against a seeded litigation matter and the production case-law corpus. Nothing in that harness is model-specific, so every model sees identical prompts, tools, and iteration budgets. Litco meters the spend on every call.
The two modes
Litco runs battery v2 against every model twice. With Litco safeguards is the production configuration, the agent with its verification stack running, over a 42-task battery. Without safeguards is the model on its own over a 59-task battery that adds the fabrication probes. Litco records what the model does in that mode, and corrects nothing. The two batteries share their core tasks, and the battery size is stated wherever the numbers appear. Litco compares composites across the two modes as the run reported them, and the batteries differ enough that the comparison is a rough one.
Candor penalties and routing eligibility
A model that cites an authority which does not exist loses twenty points from its composite for each fabrication, and carries a red flag on every chart and table here. A model that goes along with a false premise built into a question loses ten points each time. No composite drops below twenty, however many penalties a model runs up, so for a model with several penalties the gap between its two published scores understates them. A model that fabricates also stops being eligible for Litco routing in that mode, whatever its score. Litco routes work inside the product to the models on this board, and it will not route that work to a model that invented a citation. Litco also cuts a model’s drafting score by its rate of AI writing tells, measured in the AI-isms task set above, to a maximum of fifteen points. The score before penalties is published next to the composite, so you can see what each penalty cost.
Infrastructure failures
When a model’s calls on a task move zero tokens, Litco counts an infrastructure failure and keeps the task out of every quality score. If most of a row’s tasks failed that way, Litco publishes the row as quarantined, with the failure stated, instead of giving it a low score. One with-safeguards row is quarantined right now, and the leaderboard says so on the row.
Serving policy and pinning
Every row names the lane it ran on, and the published data file records the endpoint that served the model and the numeric precision it ran at. Public rows go through OpenRouter, pinned to a named provider at a disclosed precision, with zero-data-retention routing requested at call time. Two models could not be ZDR-routed through OpenRouter, so they ran on the vendor’s direct API, and their rows say so. Self-hosted rows run on Litco-operated hardware and carry that label. Because a public row’s traffic goes through an endpoint the vendor can see, a vendor can check the serving conditions of its own row without taking Litco’s word for it. If a lane or a precision changes, the result publishes as a new row, and Litco does not edit an old one in place.
Hold-out policy
Litco publishes the methodology and keeps the task set private. A lab that can read the tasks can train against them, and a benchmark a lab has trained against stops predicting how a model will do on real work. Each battery carries a version number, and when Litco retires a version it publishes that version’s tasks and replaces them with fresh ones. Litco compares scores only within a single battery version.
What is still landing
The numbers on this page are the ones the run itself reported, taken from its progress feed. Latency percentiles have not been measured yet, and a final uniform rescore will replace this data when the run finishes ingesting. The changelog will record that swap.
What the with-safeguards mode is measuring
The with-safeguards rows run under the shipped product’s verification stack: byte-level quote checks, retrieval fingerprints, citation resolution, the citator, verify-on-stop, and the document commit gate. The verification page describes each of those mechanisms. This page measures what they change in the numbers.
Changelog
FAQ
Is this related to Stanford’s LegalBench or to Benchmark Litigation?
No on both counts. LitigationBench is unrelated to LegalBench, the academic legal-reasoning benchmark from Stanford, and unrelated to Benchmark Litigation, the directory that ranks litigation firms. This page benchmarks AI models on litigation work inside the Litco platform.
Why does a strong model score lower with Litco’s safeguards active?
A model that never fabricates never triggers the candor penalty, so it has nothing to gain from the safeguards, while Litco’s checks add steps to every task. That is enough to put such a model a few points below its own unguarded score. The batteries differ too. The with-safeguards battery adds adversarial pressure tasks the other one does not run, which is why the two scores are not directly comparable. The safeguards are there for the failures a firm cannot see coming, and the models that do fabricate are where the difference shows up. With the safeguards running, their fabrication count on ordinary work goes to zero.
Which endpoint served each model, and what about data retention?
Every row names its lane in the table, and the published data file records the endpoint and the numeric precision that served each model. The serving policy above gives the full picture. For public rows Litco requested zero-data-retention routing on every call. Self-hosted rows run on Litco-operated hardware, and their traffic stays on that hardware.
How do new models get onto the board?
When a model worth testing ships, Litco runs the current battery version against it and republishes. The changelog records every addition, re-run, and battery rotation, and each row keeps its run date.
Can I see the tasks, or run the battery myself?
The task set stays private while its battery version is live, for the reason given in the hold-out policy above. Once a version retires, Litco publishes its tasks in full. If you want a model evaluated, or want to reproduce the method against your own task set, get in touch.