Evaluation

Right number, wrong clause: measuring citation faithfulness

Field accuracy is the metric everyone reports. It misses the quadrant that fails an audit: the value is correct and the clause it points to is not. That quadrant was 4.6 percent of our fields before the validator and is 1.1 today, inside a 97.4 percent first pass, and it is invisible to a spot check.

1,840 fields, scored on both halves of the answer value accuracy alone would report 98.3%. Citation-exact reports 97.4%. The difference is one quadrant. clause cited is the operative one yes no valuecorrect yes no 97.4% right, and provably so 1.1% passes every spot check, fails the audit was 4.6% before the validator 0.9% misread, but you can see where 0.6% wrong twice; caught by the validator production first pass, September 2026. The second read on the routed 15% of fields takes the top-left cell to 99.7%.
The top-right cell is the one to worry about. The number matches the answer key, so a reviewer checking values signs it off, and the clause it points to says something else.

Ask an extraction vendor for their accuracy and you will get one number, and it will be value accuracy: the share of fields where the extracted value matched the answer key. It is the natural metric, and it is the one we reported ourselves until we started scoring citations separately. Then a quadrant appeared that value accuracy cannot see, and it turned out to be the one our clients' auditors care about most. Our first pass now scores 97.4 percent on both halves of the answer at once, 99.7 after the second read. This note is about why we count it that way, and what the count showed before we fixed it.

Two halves of an answer

Every field in our register is two things: a value, and the clause the value came from. "Maximum Total Net Leverage: 4.50x, Section 8.01(a), page 142." The value drives the test. The clause is what a credit officer opens when the test is near the line, what an auditor asks for, and what a lawyer reads when the borrower disputes it. A field with the right value and the wrong clause is a test that passes every month and a finding waiting to happen.

So we score four outcomes rather than two: value right or wrong, clause right or wrong. The clause counts as right only if it is the operative provision a reviewer would cite. Not a recital, not a schedule that repeats the figure, not a definition that happens to contain the same number.

What the matrix shows today

On our 1,840-field set, the production first pass splits as the figure shows. Value accuracy alone: 98.3 percent, a number anyone would put on a slide. Citation-exact: 97.4. The 0.9 points between them are twenty fields where the value is right and the clause is not, and they are the fields the confidence routing sends to a person first. After that second read the set stands at 99.7.

What it showed before

That top-right cell was not always small. Before the span validator went into the pipeline, with a frontier model reading the document directly, it was 4.6 percent: 85 fields that would have passed any value check and failed the first audit. Value accuracy at the time was 95.9. Citation-exact was 91.3. We read all 85, and they fell into three groups.

  • The number appears twice. A threshold stated in the covenant and restated in a recital, a term sheet schedule or a certificate template. The model cited the first place it saw it. Forty-one fields.
  • The clause moved. An amendment restated Section 8.01 and the model cited the original, which still carried the old and coincidentally identical figure for the current period. Twenty-six fields.
  • Right section, wrong subsection. The value sits in 8.01(a)(ii) and the citation says 8.01. Close enough for a human, not for a link that has to open the right paragraph. Eighteen fields.

Why a spot check misses it

The standard acceptance test for extraction is a sample: pull fifty fields, check the values, sign off if forty-nine match. Every field in the dangerous quadrant passes that test, because its value is right. The check that would catch it, opening the cited clause and reading it, is the expensive one nobody runs on fifty fields. So a pipeline with 4.6 percent of its citations pointing at the wrong place ships with a 96 percent sticker on it, and the first person to find out is the one who needed the citation to be right.

Value accuracy measures whether the register is right today. Citation accuracy measures whether anyone will believe it in a year.

What changed

Two things, both described in the architecture note. The model now chooses from candidate spans found by the parser rather than searching the document, which removes most of the duplicate-number problem because the candidates for a covenant field come from the covenant section. And the validator checks that the cited span contains the value, which turns a wrong-value answer into a retry rather than an error. Together they took the top-right cell from 4.6 percent to 1.1 and the first pass from 91.3 to 97.4 on the same set. The twenty fields that remain are mostly amendments, a document-set problem more than a reading one, and they are exactly what the second read is for.

Ask for the matrix

If you are evaluating extraction, ask for accuracy as a two-by-two rather than a single figure, and ask how the clause half was scored. A vendor that has measured it will have the matrix to hand. A vendor that has not will have a very high number.

See the pipeline on your own agreements

Bring three agreements, one of them a scan. We run the extraction in front of you, field by field, with the clause each number came from and the confidence it carried.