Ask an extraction vendor for their accuracy and you will get one number, and it will be value accuracy: the share of fields where the extracted value matched the answer key. It is the natural metric, and it is the one we reported ourselves until we started scoring citations separately. Then a quadrant appeared that value accuracy cannot see, and it turned out to be the one our clients' auditors care about most. Our first pass now scores 97.4 percent on both halves of the answer at once, 99.7 after the second read. This note is about why we count it that way, and what the count showed before we fixed it.
Two halves of an answer
Every field in our register is two things: a value, and the clause the value came from. "Maximum Total Net Leverage: 4.50x, Section 8.01(a), page 142." The value drives the test. The clause is what a credit officer opens when the test is near the line, what an auditor asks for, and what a lawyer reads when the borrower disputes it. A field with the right value and the wrong clause is a test that passes every month and a finding waiting to happen.
So we score four outcomes rather than two: value right or wrong, clause right or wrong. The clause counts as right only if it is the operative provision a reviewer would cite. Not a recital, not a schedule that repeats the figure, not a definition that happens to contain the same number.
What the matrix shows today
On our 1,840-field set, the production first pass splits as the figure shows. Value accuracy alone: 98.3 percent, a number anyone would put on a slide. Citation-exact: 97.4. The 0.9 points between them are twenty fields where the value is right and the clause is not, and they are the fields the confidence routing sends to a person first. After that second read the set stands at 99.7.
What it showed before
That top-right cell was not always small. Before the span validator went into the pipeline, with a frontier model reading the document directly, it was 4.6 percent: 85 fields that would have passed any value check and failed the first audit. Value accuracy at the time was 95.9. Citation-exact was 91.3. We read all 85, and they fell into three groups.
- The number appears twice. A threshold stated in the covenant and restated in a recital, a term sheet schedule or a certificate template. The model cited the first place it saw it. Forty-one fields.
- The clause moved. An amendment restated Section 8.01 and the model cited the original, which still carried the old and coincidentally identical figure for the current period. Twenty-six fields.
- Right section, wrong subsection. The value sits in 8.01(a)(ii) and the citation says 8.01. Close enough for a human, not for a link that has to open the right paragraph. Eighteen fields.
Why a spot check misses it
The standard acceptance test for extraction is a sample: pull fifty fields, check the values, sign off if forty-nine match. Every field in the dangerous quadrant passes that test, because its value is right. The check that would catch it, opening the cited clause and reading it, is the expensive one nobody runs on fifty fields. So a pipeline with 4.6 percent of its citations pointing at the wrong place ships with a 96 percent sticker on it, and the first person to find out is the one who needed the citation to be right.
Value accuracy measures whether the register is right today. Citation accuracy measures whether anyone will believe it in a year.
What changed
Two things, both described in the architecture note. The model now chooses from candidate spans found by the parser rather than searching the document, which removes most of the duplicate-number problem because the candidates for a covenant field come from the covenant section. And the validator checks that the cited span contains the value, which turns a wrong-value answer into a retry rather than an error. Together they took the top-right cell from 4.6 percent to 1.1 and the first pass from 91.3 to 97.4 on the same set. The twenty fields that remain are mostly amendments, a document-set problem more than a reading one, and they are exactly what the second read is for.
Ask for the matrix
If you are evaluating extraction, ask for accuracy as a two-by-two rather than a single figure, and ask how the clause half was scored. A vendor that has measured it will have the matrix to hand. A vendor that has not will have a very high number.