Financial Spreading With Citations: Why Every Figure Needs a Document and Page Reference
Because a figure nobody can trace is a figure nobody can defend. A citation ties each spread line to a document, a page and a location on that page, so a reviewer verifies by clicking rather than searching. It does not make the figure correct. It makes it checkable — a weaker claim, and a far more useful one.
Key facts
- YuSight spreads with 100% of figures cited and one-click source verification. A sample credit assessment memo produced in 28 minutes carried 142 citations — one per figure, not one per section.
- Supervisors already require the tie-back. BCBS 239 Principle 3 states that "risk data should be reconciled with bank's sources, including accounting data where appropriate, to ensure that the risk data is accurate," and that data should be "aggregated on a largely automated basis so as to minimise the probability of errors" (BIS, January 2013).
- Self-reported numbers and replayed numbers diverge, measurably. In the FinBalance accounting-reconciliation benchmark, four of six evaluated models showed a 26–41 percentage-point gap between the balance sheet they reported and the balance sheet obtained by replaying their own journal entries through a deterministic ledger. The highest exact-balance-sheet score across all models was 46% (Tumpati et al., arXiv:2606.15949, June 2026).
- Model governance moved, and did not get looser about evidence. SR 26-2, *Revised Guidance on Model Risk Management*, 17 April 2026 supersedes SR 11-7 and states it is "most relevant to banking organizations with over $30 billion in total assets." Whatever your supervisor's current scope, the thing they ask for in a file review is unchanged: show me where this came from.
- A citation is a pointer, not a proof. Four distinct ways a perfectly cited figure is still wrong are set out below, with what to do about each.
Why is an uncited figure unusable in a regulated credit file?
Not because it is likely wrong. Because it cannot be adjudicated.
Put yourself at the end of the chain. A credit committee member reads "FY25 DSCR 1.42x, comfortably above the 1.35x covenant floor" and wants to know whether the 1.42x used total finance costs or term-loan interest only. In an uncited file the answer takes one of three forms: the analyst remembers, the analyst re-derives it, or the analyst opens a 96-page PDF and hunts. All three cost time and none of them is evidence.
Now scale that. A working file for a mid-market borrower holds three years of audited accounts, a provisional set, bank statements, a bureau report and tax filings — call it 400 pages. A spread carries 90 to 150 figures. Every one of those figures is a claim about a specific location in that stack. Without a stored pointer, the location is reconstructed by memory each time somebody asks, and memory is not an audit trail.
Three consequences follow, in order of how much they cost:
- Review becomes a search problem. The reviewer's scarce attention goes on finding the number rather than judging it.
- Disagreements become unresolvable. Two analysts quoting different EBITDA have no shared object to point at, so the argument runs on seniority.
- The file cannot be re-examined later. Eighteen months on, in a post-mortem or an examination, nobody can reconstruct what was actually read off the page in 2025 versus what was assumed.
This is a different problem from model hallucination, and it needs a different control. Stopping an AI tool from inventing numbers is about generation. Citations are about retrieval — whether an accepted figure can be sent back to where it came from.
What must a citation contain to be worth anything?
"Source: audited financials" is not a citation. Neither is "page 14" on its own. A citation that survives contact with a reviewer carries seven elements.
Element | What it must contain | What breaks without it |
|---|---|---|
Document identity | Immutable document ID or content hash, document type, entity name, period, and which upload — files get re-uploaded | Two versions of the same statement in the file; the citation points at the wrong one |
Page | Page number of the stored rendering, not of the printed statement's own numbering | Statements number their own pages from 1; the PDF does not |
Location on page | Bounding box coordinates, or table cell as (table ID, row, column) | The reviewer lands on page 14 and still has to scan it |
Source token | The literal characters read, before normalisation — | Scale, sign and separator errors become invisible |
Transform | Sign convention applied, scale factor, currency, and the chart-of-accounts mapping used | The number in the memo differs from the number on the page and nobody knows why |
Confidence and checks | Per-field extraction confidence, plus which deterministic checks the figure passed (footing, cross-statement tie) | You cannot triage. Everything looks equally certain |
Edit provenance | If a human changed it: who, when, previous value, stated reason, and the citation's new status | The most-likely-to-be-wrong figures are exactly the ones a person overrode |
The two that get skipped most often are source token and transform, and they are the two that catch the errors that foot perfectly. A statement printed in ₹ lakh read as ₹ million produces a spread that balances, ties across statements and is wrong by 10x. Only a stored raw token and a stored scale factor let a reviewer see that in one glance.
Derived figures need a different kind of citation
DSCR does not appear on any page. Neither does TOL/TNW, working capital cycle, or interest coverage. A page citation on a computed ratio is a category error, and a system that emits one is guessing.
A derived figure needs a formula citation: the definition used, in full, plus the operand list, and each operand carries its own page citation. That is what makes covenant arguments tractable, because covenant disputes are almost never about arithmetic — they are about which line items went into which side. Covenant testing for DSCR and leverage covers the definitional variants; the citation layer is what makes the variant you actually used inspectable.
Worked walkthrough: tracing one number from the memo back to the page
A mid-market manufacturer. The memo says: FY25 DSCR of 1.42x against a 1.35x covenant floor. Here is the whole chain, exactly as a reviewer would click through it.
Step 1 — the memo figure carries a formula citation.
DSCR (term debt basis) = (Profit after tax + Depreciation and amortisation + Interest on term loans) ÷ (Interest on term loans + Term loan principal repayable within 12 months)
Step 2 — four operands, four page citations.
Operand | Value ($'000) | Document | Page | Location | Source token | Confidence |
|---|---|---|---|---|---|---|
Profit after tax, FY25 | 486.20 | Audited FS FY25 (doc #A-4471) | p.12 | Table 3, row "Profit for the year", col "FY25" |
| 0.99 |
Depreciation and amortisation, FY25 | 341.20 | Audited FS FY25 (doc #A-4471) | p.12 | Table 3, row "Depreciation and amortisation expense", col "FY25" |
| 0.99 |
Interest on term loans, FY25 | 188.40 | Audited FS FY25 (doc #A-4471) | p.31 | Note 24 "Finance costs", row "Interest on term loans" |
| 0.97 |
Term loan principal due within 12 months | 527.00 | Audited FS FY25 (doc #A-4471) | p.26 | Note 14 "Borrowings", row "Current maturities of long-term borrowings" |
| 0.96 |
Step 3 — the arithmetic, shown.
- Numerator: 486.20 + 341.20 + 188.40 = 1,015.80
- Denominator: 188.40 + 527.00 = 715.40
- DSCR: 1,015.80 ÷ 715.40 = 1.4199 → 1.42x
Step 4 — the check the citation makes possible. Note 24 on page 31 also shows interest on working capital borrowings of 71.30 and other borrowing costs of 9.00. Total finance costs on the face of the P&L are 188.40 + 71.30 + 9.00 = 268.70. So there is a second, equally defensible DSCR:
- Numerator: 486.20 + 341.20 + 268.70 = 1,096.10
- Denominator: 268.70 + 527.00 = 795.70
- DSCR: 1,096.10 ÷ 795.70 = 1.3776 → 1.38x
1.42x passes the 1.35x covenant. 1.38x also passes — but by 0.03x of headroom instead of 0.07x. And a third variant that adds lease interest, or nets cash, moves it again.
That is the entire value of the citation layer in one example. Nobody's arithmetic was wrong. The two figures differ because one operand was drawn from a note and the other from the face of the statement, and only a citation makes that visible in under ten seconds. Without it, the committee debates a number. With it, the committee debates a definition, which is the debate worth having.
How do citations change the review workflow?
They convert reviewing from search to verification, and that changes what the reviewer's time is spent on.
Without citations, an analyst reviewing a spread performs the extraction a second time. Open the PDF, find the statement, find the line, compare. On a 90-field spread this is not a review; it is a re-spread, and it is why "checking the spread" quietly costs as much as producing it. See how much analyst time automated spreading saves per file for what that costs in hours.
With citations, the reviewer's loop is: click the figure, the source page opens with the cell highlighted, eye confirms, move on. Roughly two to five seconds per field rather than sixty to ninety. More importantly, the reviewer can choose where to spend attention, because the spread now sorts by confidence and by check status. Fields that failed a footing test, fields the extractor was unsure of, and fields a human overrode go to the top. Everything else is spot-checked.
That triage is the actual productivity gain, and it is why the citation layer belongs to the human-in-the-loop review design rather than to the extraction engine.
What happens when the analyst edits a figure
This is where most citation implementations fall apart. An analyst decides the extracted "Other income" is misclassified, types a new number, and the citation either silently persists — now pointing at a page that says something different — or silently disappears.
Neither is acceptable. An edited figure should keep its original citation, gain an override record, and change status:
State | Citation shows | Reviewer reads it as |
|---|---|---|
Extracted, checks passed | Document, page, box, token | Verifiable against source |
Extracted, low confidence | Same, flagged | Verify this one first |
Human-edited, source-consistent | Original citation + who/when/previous value/reason | Reclassified, not re-read — the page still says what it said |
Human-edited, source-divergent | Original citation + override record + explicit divergence flag | Analyst judgement. The page does not support this number |
Manually entered, no source | No citation | Assumption. Must carry a stated basis |
The fourth row is the one examiners care about, and the one that should never be silently indistinguishable from the first. A spread where every figure looks equally sourced, but 8% of them were typed in, is worse than an honest spread with 8% flagged as judgement.
What do citations get you in an audit or examination?
Three things, and it is worth being precise about which.
A reconstruction, not a recollection. BCBS 239's reconciliation expectation is satisfied by a stored pointer, not by an analyst's recall. When an examiner asks how a risk-weighted figure was derived, the file answers.
A bounded review. An examiner sampling ten files can verify twenty figures in each in minutes rather than a morning. This changes the character of the sample — reviewers who can check cheaply check more, which is uncomfortable and correct.
A defensible position on AI-assisted work. The question in examiner review of an AI-drafted memo is rarely "did a model touch this". It is "can you show, for this specific number in this specific file, where it came from and who accepted it". A citation with an edit history answers that; a confidence score does not.
Being honest: citations do not make a figure correct
This matters more than anything above, and the industry is sloppy about it.
A citation asserts this number appears at this location in this document. It asserts nothing about whether the number belongs where you put it, whether the document is genuine, or whether the statement is true. Four failure modes survive a perfect citation:
- Right number, wrong period. The extractor read the correct cell in the wrong column. The citation is accurate. The spread has FY24 revenue labelled FY25. Only a period-attribution check — does the column header parse to the stated year — catches this.
- Right number, wrong entity. In a group structure, the subsidiary's accounts spread into the parent's file. The citation points faithfully at a real page in a real document belonging to the wrong borrower. This is a document-mapping problem, addressed before extraction: see multi-entity document mapping.
- Right number, wrong mapping. A director's loan cited correctly from page 27 and mapped to "Other long-term liabilities" instead of "Related party debt". Extraction was perfect; the spread is misleading. Only a chart-of-accounts review catches it, and the citation is what makes that review possible.
- Right number, forged document. A citation to page 3 of a fabricated bank statement is a perfectly valid citation. Authenticity is a separate control entirely — see how to detect a fake or tampered bank statement.
So the honest claim is narrow: citations make a spread checkable at low cost. They do not make it right, and any vendor — us included — who lets "fully cited" slide into "fully accurate" is selling the wrong thing. What makes a spread right is the combination of citations, deterministic checks in code, and an analyst who spends the time citations free up on the judgement fields. Accuracy claims themselves are a separate and messier subject, covered in how accurate AI actually is at extracting financial statement data and in extraction accuracy vs straight-through rate.
How do you test a vendor's citations in 20 minutes?
Take one of your own worst documents — a scanned set with a continuation schedule — and run this list.
- Click ten figures at random. Does the source page open with the cell highlighted, or does it open the PDF at page 1?
- Click a computed ratio. Does it show the formula and every operand with its own citation, or does it show a page number? A page number here is a red flag.
- Find a figure that continues across a page break. Where does the citation point? Both pages, or one arbitrarily?
- Edit a figure and re-open it. Is the original citation still there? Is the override recorded with a reason field? Is it visually distinct?
- Ask for a manually entered figure. Can the system represent "no source" honestly, or does it fabricate a citation?
- Re-upload a corrected version of the same statement. Do citations rebind to the new document, or silently keep pointing at the superseded one?
- Export the file. Do the citations survive export to PDF, or is the audit trail only visible inside the vendor's UI?
Question 7 separates products more than any of the others. Financial spreading software: what it does and what it doesn't goes further into what to test.
FAQ
Why does spreading need citations?
Because a credit file gets read by people who were not there when it was built — a committee member, an auditor, a supervisor, or you in eighteen months. A citation is the only thing that lets any of them check a figure without redoing the work.
How do you verify a spread figure against the source document?
You click it. In a properly cited spread the figure opens the source page with the exact cell highlighted, and shows you the raw characters that were read before any sign or scale conversion. If you find yourself scrolling a PDF, the citation was not doing its job.
How do we keep an auditable trail when AI drafts our credit memos?
Store the pointer, not just the answer. Every figure keeps its document ID, page, location on the page and raw token; every human edit keeps the previous value, the person, the timestamp and a reason. The trail is the pair — the extraction record and the override record — and it has to survive export.
What is the difference between a citation and a confidence score?
A confidence score tells you how sure the system was. A citation tells you where to look. You need both, and they do different jobs: confidence decides what gets reviewed first, the citation makes reviewing it fast.
Can a language model just tell me the page number?
It can tell you a page number, and it will sound convincing. But a page number generated by the same process that generated the figure is not independent evidence. A citation only counts when the extraction layer produced it from coordinates, not when a model wrote it into a sentence.
Do citations slow down spreading?
No, but they change where the time goes. Producing the citation costs nothing extra if the extraction layer already knows the coordinates. What changes is that review gets much faster and much more targeted, because attention concentrates on low-confidence and overridden fields instead of being spread evenly.
How should an edited figure be shown in the file?
Visibly differently from an extracted one. Keep the original citation, add who changed it, when, from what, and why, and flag whether the new value still agrees with the source page. An analyst reclassifying a line item and an analyst overriding what the page says are two different acts and should not look the same.
Does a citation prove the figure is correct?
No — and this is the important limitation. A citation proves the number appears at that location in that document. It says nothing about whether the period, the entity, the mapping or the document itself is right. Those need their own checks.
What should a citation on a ratio look like?
The full formula, the operand list, and a separate page citation on each operand. Ratios do not exist on any page, so a page reference on a DSCR means the system is guessing — and it is exactly the place where two defensible definitions produce two different covenant answers.
Conclusion
Three things to take away:
- A citation must carry seven things, not one: document identity, page, location on the page, the raw token, the transform applied, confidence and check status, and edit provenance. Anything less and a reviewer is still searching.
- Derived figures need formula citations. The covenant arguments you will actually have are about which operands went in, and only an operand-level trail settles them — as the 1.42x versus 1.38x walkthrough above shows.
- Citations make a figure checkable, not correct. Right number, wrong period; wrong entity; wrong mapping; forged document — all four survive a perfect citation, and all four need their own control.
YuSight's Financial Spreading extracts and standardises financials, computes DSCR, leverage, liquidity and profitability ratios in code from traceable inputs, and keeps 100% of figures cited with one-click source verification — every figure traced to source document and page, spreads analyst-editable, with full version history of who changed what. A sample CAM produced in 28 minutes carried 142 citations. If you want to see whether that holds on a document set of yours rather than a clean demo file, bring the worst scan you have.
Watch YuSight spread a real balance sheet — book a live demo.