Extraction

Extraction is where monitoring starts, not where it ends

LLM extraction from prospectuses and ISDA schedules works. What it does not give you is a register that stays live: a fresh NAV arriving, a test re-running, a threshold moving, a breach routed to the right desk.

EXTRACTION: a snapshot ProspectusISDA Schedule Amendmentside letter Terms, cited yours or ours seeds LIVE REGISTER: a loop that runs every day Register calendar, queue, audit NAV arrives tests re-run breach routed threshold moves silence: stale chase logged a photograph of a moving object on the left; the thing that keeps moving on the right
Extraction produces a cited snapshot of the terms. Monitoring is the loop that runs against that snapshot for years. ExactCov consumes the snapshot, whoever produced it.

Large language models are now good at reading fund documents. Point one at a prospectus, an ISDA Schedule or a prime brokerage agreement, ask for the NAV decline triggers, the thresholds, the windows and the basis, and you will get a structured answer with page references, most of the time correctly. We know because we do it, and because several of the teams we talk to have built the same thing in-house over a few weeks.

If you have built that, keep it. This post is not an argument that extraction is hard. It is an argument that extraction is the start of the job, and that the rest of the job is a different kind of system.

What extraction gives you

Extraction turns a document into a snapshot: on the day you ran it, this agreement contained these terms. That snapshot is valuable. It is what a register is seeded from, and it is the thing that lets you say "we read every agreement in the book" rather than "we read the ones we remembered to."

A snapshot has no clock. It does not know that the fund reports monthly and that the next statement is due on the fifth business day. It does not know that an amendment was signed in June, or that a sub-fund was added to the umbrella in August. It does not know that yesterday's NAV notice was indicative, not official. It is a very good photograph of a moving object.

What a live register needs

Monitoring is what happens to the snapshot over the following years. In our experience it comes down to a short list of things a register has to do on its own, without anyone remembering to run a script.

  • A fresh NAV arriving. By email, from the administrator's portal, from a data vendor, as a PDF or a spreadsheet. It has to be recognised, attributed to the right fund and period, validated against the previous figure, and stored with its source page.
  • A test re-running. The moment the observation is validated, every test whose window covers it runs again, against the basis the clause specifies, and writes a result with headroom and the value used.
  • A threshold moving. The audited accounts land and the year-end floor re-bases. An amendment changes the one-month trigger from 10% to 15%. The change is recorded against the document that caused it and the tests pick it up from that date.
  • A breach routed. A failed test is not a row in a table. It is an escalation with an owner, a due date for acknowledgement, and the notice and the clause attached, sent to the desk that holds the limit.
  • Silence noticed. When nothing arrives by the due date, the figure is marked stale, the chase is logged, and the fund moves up the morning list.

Each of those is an event that happens at an unpredictable time. A register that stays live is an event-driven system with a calendar, a queue, an exceptions list and an audit trail. That is not what an extraction pipeline is, and it is not what an extraction pipeline should try to become.

Where ExactCov sits

We built ExactCov to consume extraction output, not to compete with it. Our own extractor produces a register with citations, and it is the fastest way to get started. But the register format is open. If your team has already extracted the book, you can load those terms directly, keep your extractor as the source of truth for the documents, and let the register take over the calendar, the tests, the limits and the routing.

Extraction answers "what does the agreement say?". Monitoring answers "what is true this morning, and what changed?". They are different questions and they deserve different systems.

If you are the team that built the extractor, that should read as validation. The hard part of what you did was getting the documents read accurately at scale, and that part stands. The part you have not built yet is the part that runs at 6am every day for a decade. That is the part we built.

See it against your own book

Bring a handful of agreements and the NAV notices you already receive. We show what the register looks like, what is stale, and what would have fired.