The legal-AI hallucination problem (and the layer that catches it)

In 2023, a lawyer citing a case that didn’t exist was a curiosity — one viral story about ChatGPT inventing precedent in Mata v. Avianca. In 2026, it’s a documented epidemic. A research database maintained at HEC Paris has cataloged more than 1,200 cases worldwide in which court filings contained AI-generated fabrications. A Bloomberg Law analysis found the number of such filings surged roughly sevenfold in 2025 alone.

And courts have stopped being patient. In June 2026 the Ninth Circuit sanctioned two attorneys $2,500 each and suspended them for six months over briefs citing opinions that didn’t exist and quotes that were never written. Months earlier the Sixth Circuit hit attorneys with $30,000 in sanctions for a brief containing more than two dozen fake citations. A federal prosecutor resigned over it. The reputational and financial stakes are now fully real.

If you’re building AI for the legal market, this is your defining product risk. Here’s the anatomy of it — and an honest account of what a safety layer can and can’t do about it.

Why legal is uniquely exposed

Most domains can absorb the occasional confident-but-wrong answer. Legal can’t, for three compounding reasons.

First, legal work is authority-based: an argument is only as good as the cases, statutes, and quotes it rests on. A fabricated authority doesn’t just weaken a brief — it can invalidate it and sanction its author.

Second, hallucinations in legal text are camouflaged by format. A made-up citation follows the exact syntactic shape of a real one — a plausible reporter, a plausible court, a plausible year. The model isn’t lying clumsily; it’s pattern-matching the form of legal authority while inventing the substance. To anyone who doesn’t independently verify, a fabrication is indistinguishable from genuine law.

Third — and this is the part most teams underestimate — even purpose-built legal AI hallucinates. A study cited by the Ninth Circuit found that legal research tools from Westlaw and Lexis hallucinated 17% and 33% of answers respectively on a representative set of questions. Domain training narrows the problem; it does not close it.

The most dangerous hallucination isn’t the one you think

When people picture legal AI hallucinations, they picture invented cases — Smith v. Jones, a citation to a case that was never decided. Those are bad, but they’re also the easiest to catch, because the case simply doesn’t exist when someone looks.

The far more dangerous failure is subtler. The Sixth Circuit, in United States v. Farris, made a point that should worry every legal-AI builder: citing a real case incorrectly can be just as serious as inventing one. A brief that quotes a real reporter citation — but attaches a fabricated quote, or cites it for a proposition it doesn’t support — sails past casual review precisely because the citation checks out. The reader sees a familiar, valid-looking authority and moves on. The fabrication is in the words attributed to it, not in the citation itself.

Hold onto that distinction, because it’s exactly where the honesty about any safety layer lives.

What the courts actually demand

Across every one of these rulings, the courts converge on one principle, and it’s the foundation any product in this space has to be built on: the lawyer is accountable, and verification cannot be outsourced. The ABA’s Formal Opinion 512 established that using generative AI doesn’t dilute a lawyer’s duties of competence and candor one bit. The Seventh Circuit put it bluntly — citation checking is easier now than ever, and failing to do it is inexcusable. Intent isn’t even required for sanctions; negligent reliance is enough.

So the goal of any safety layer is not to replace that verification. It’s to make sure nothing slips to a human’s desk — or a court’s — unflagged.

Where a gateway-level layer fits — honestly

Here’s the honest version, because legal professionals will see through anything less.

A layer that sits in front of your model — like Metriqual — scans every response, before it reaches a user, for the patterns that signal hallucination risk: citation structures, quotation blocks, references to statutes and authorities, suspiciously precise figures. It scores that risk and flags it, so a response carrying legal authorities is surfaced for verification rather than passing silently through.

In a citation-dense domain, that means it flags a lot — and that’s the point, not a flaw. You don’t want a layer that selectively decides which citations are worth checking. You want one that guarantees no quote or authority reaches a filing without being surfaced for the review the court already requires. It turns “we hope someone checked” into “everything that needed checking was marked.”

Now the boundary, stated plainly: this catches the signal, not the truth. Pattern-matching can tell you a response contains a citation and a quote that warrant verification. It cannot, on its own, tell you that the quote is fabricated or that the real case doesn’t support the point — that’s exactly the Farris problem, where the citation is genuine and only the words are invented. Determining that still requires checking against a real legal corpus, or a human who knows the law. A gateway flag is a smoke detector: it tells you where to look. It does not put out the fire, and it does not certify that there’s no fire where it stayed silent.

That’s not a weakness to hide; it’s the correct division of labor. The layer catches the risk signal so the accountable human — or a downstream verification step against authoritative sources — can catch the actual fabrication. Defense in depth, with the lawyer’s duty firmly intact at the end of it.

The honest pitch

If you’re building legal AI, the move that fails is promising your users that a tool “catches hallucinations” and lets them file with confidence. The courts have been explicit that nothing removes the lawyer’s duty to verify, and a product that implies otherwise isn’t just overclaiming — it’s setting its users up for the exact sanction they’re trying to avoid.

The move that works is honest defense in depth: every response carrying legal authority gets flagged before it ships, so verification is systematic instead of hoped-for, and the human stays exactly where the law puts them — accountable, with a net underneath them that makes sure nothing reaches the brief unexamined.

That’s the layer worth having. Not one that claims to make hallucinations disappear, but one that makes sure they can’t slip by unseen on the way to a filing. (For how differently models behave when they hit the edge of what they know — and why that behavior, not raw capability, is the real variable — see We asked three models about a book that doesn’t exist.)


Metriqual is the control layer for AI in production — one OpenAI-compatible key across every major provider, with per-key controls, observability, and response-level risk flagging. Start free or book a demo.

Scroll to Top