How Should a Commercial Lender Evaluate Spreading Software? A 22-Point Scorecard
Score every candidate against 22 criteria in six groups, weighted to 100, with a hard gate on two of them: mapping to your own chart of accounts, and figure-level citation. Seven criteria are table stakes worth 21 points combined; the other fifteen carry 79 points and are where vendors actually separate. YuSight extracts at 95.2% accuracy against a manual benchmark, with every figure traceable to its source document and page.
The scorecard exists because demos do not discriminate. Every product in this category reads a clean audited PDF and produces a tidy table in eight minutes. The differences appear on the fourth borrower, in the notes, and in whoever has to file a change request to redefine a ratio.
Key facts
- Company-prepared statements are the normal input, not the exception. 87% of banks evaluate non-audited financial statements for most or all $250,000 loans, and 94% do so at $1 million and $3 million (FDIC, 2024 Small Business Lending Survey, Section 3). Any accuracy claim measured on audited filings is measured on the wrong document.
- Automation is still the minority position. Only 9% of small banks use a credit-scoring model compared with 46% of large banks, and auto-approval for large loans sits at 0% and 1% respectively (same source). The spreading engine is feeding a human, and should be built that way.
- Examiners test the file, not the tool. The OCC's Commercial Loans booklet directs examiners to "assess the quality of credit file documentation" and to test for "loans not supported by current and complete financial information" (Comptroller's Handbook, Section 206). A spread nobody can trace to a page is weak documentation however it was produced.
- Model governance moved in 2026. SR 26-2, Revised Guidance on Model Risk Management (17 April 2026), supersedes SR 11-7 and SR 21-8 and sets a risk-based approach scaled to a firm's model risk profile (Federal Reserve). Ask the vendor which version of the guidance their documentation was written against.
- YuSight extracts at 95.2% accuracy against a manual benchmark, cites 100% of figures to document and page, and produces a 28-minute sample CAM carrying 142 citations.
If you have not yet drawn the line between a spreading engine and an extraction tool with a spreadsheet export, start with what spreading software does and doesn't do and run its twelve demo tests. This page is the scored instrument you take into procurement afterwards.
How does the scoring work?
Score each criterion 0 to 5. Weighted points = weight × score ÷ 5. The maximum is 100.
- 5 — demonstrated on your files, in your environment, without vendor intervention
- 3 — demonstrated on vendor files, or on yours with vendor hands on the keyboard
- 1 — asserted, roadmap, or "we can configure that"
- 0 — cannot do it
Two hard gates. A score below 3 on criterion 5 (chart-of-accounts mapping) or criterion 9 (figure-level citation) disqualifies regardless of total. These are not weighted more heavily than everything else because weighting is a poor way to express a veto.
Verdict bands: 75 and above, proceed. 60–74, proceed with named conditions in the contract. Below 60, decline.
Can it read the documents you actually receive?
1. Company-prepared and accountant-prepared statements — weight 6, differentiator. Strong: handles a compilation with no notes, an inconsistent prior-year column and a page numbered by hand, and flags what it cannot see rather than inferring it. Weak: accuracy figures quoted only for audited filings, with company-prepared statements described as "supported."
2. Degraded scans — weight 3, table stakes. Strong: skewed, stamped, 200 dpi and photographed-on-a-phone pages processed, with extraction confidence collapsing on illegible fields. Weak: silent degradation into plausible wrong numbers, which is worse than a failure.
3. Notes, schedules and annexures — weight 5, differentiator. Strong: current maturities of long-term debt disclosed only in a note are carved out of the aggregate line automatically; the composition of "other current liabilities" is pulled from the schedule. Weak: the three primary statements only. This is the single most common gap and the most consequential, because the note is where the leverage is.
4. Tax returns as first-class documents — weight 4, differentiator. Strong: the return and its schedules are spread as a source in their own right, with a stated rule preventing owner income being counted twice when the return and the statement both appear. Weak: the main form is read and the schedules are attachments.
Does it produce your numbers or the borrower's?
5. Mapping to your standardised chart of accounts — weight 9, differentiator, hard gate. Strong: output arrives in your row names, your ordering, your policy treatment, configured once and applied to every borrower. Weak: a table in the borrower's own labels, with mapping described as a professional-services engagement. This is the criterion that separates a spreading engine from an extractor, and it carries the largest weight for that reason.
6. Label independence across borrowers — weight 7, differentiator. Strong: two borrowers whose accountants write "Trade payables" and "Sundry creditors" for the same economic item land in the same standardised row without anyone touching a rule. Weak: per-borrower templates, which is manual mapping with a configuration screen in front of it.
7. Normalisation rules you control — weight 6, differentiator. Strong: your credit team can change whether operating leases sit in debt, whether related-party advances are deducted from net worth, or how a 15-month period is annualised — and can see which spreads the change affects before applying it. Weak: every rule change is a vendor ticket with a release cycle attached.
8. Multi-period consistency and restatement flagging — weight 4, differentiator. Strong: FY2025 and FY2026 map identically, and where the comparative column has been restated the spread says so on its face. Weak: three independent single-period extractions stacked side by side, with column drift undetected.
Can an analyst trace, challenge and override every figure?
9. Figure-level citation — weight 7, differentiator, hard gate. Strong: click any number in the ratio output and the source page opens with the line highlighted. Weak: a document-level reference, or a citation to "the FY2026 financials." A reviewer cannot verify what they cannot find in ten seconds, and unverified is functionally uncited.
10. Confidence scoring that discriminates — weight 4, differentiator. Strong: confidence varies meaningfully across fields on the same document, and low-confidence fields are the ones an analyst actually corrects. Weak: every field returns 0.97. A score that never varies conveys nothing.
11. Override workflow and the machine-versus-analyst diff — weight 6, differentiator. Strong: the analyst edits in place, the original machine value is retained beside the override, a reason is captured, and the override rate by field is reportable. Weak: editable output with no record of what was changed. Override design is a topic in its own right — see designing the analyst override workflow.
12. Ratio definition control — weight 5, differentiator. Strong: DSCR can be defined pre-tax or post-tax, with or without existing debt service, by facility type, by your team, and the covenant test uses the identical definition. Weak: a fixed ratio library. The definitions that cause disputes are catalogued in covenant testing for DSCR and leverage.
Does it handle the structures in your portfolio?
13. Multi-entity and group handling — weight 6, differentiator. Strong: submit an operating company, a holding company and two guarantors unlabelled, and each document lands against the right entity, with intra-group items identified. Weak: every PDF processed independently and assigned by filename. The failure mode is described in multi-entity document mapping.
14. Currency, units and period length — weight 3, table stakes. Strong: thousands loaded into a millions template are caught; brackets are read as negatives; a 15-month first period is annualised and flagged. Weak: silent unit errors, which are the most expensive small mistake in this category.
15. Framework and sector coverage — weight 3, table stakes. Strong: IFRS, local GAAP and your jurisdiction's filing formats, plus templates for the sectors that dominate your book. Weak: one framework with the others on a roadmap.
Will it survive audit, examination and model governance?
16. Audit trail and version history — weight 4, table stakes. Strong: who changed which figure, when, from what to what, and which version fed the approved memo — reconstructable months later. Weak: a processing log. What an examiner expects to see is set out in examiner review of an AI-drafted memo.
17. Model documentation and change control — weight 4, differentiator. Strong: written documentation of what the model does, its limitations, its validation results, and a notification process when the underlying model changes. Weak: "we update the model continuously," offered as a benefit. Silent model change under a static credit policy is a governance problem, not a feature.
18. Data residency, retention and security posture — weight 3, table stakes. Strong: named hosting regions, a stated retention period, current third-party attestation, and a clear answer on whether your documents train anyone's model. Weak: an evasive answer to the training question.
What does it cost to run, and what happens when it is wrong?
19. Accuracy evidence you can reproduce — weight 3, differentiator. Strong: the document mix, the sample size, the definition of a field, and how partial correctness was scored — plus willingness to be measured on your files. Weak: a single percentage with no denominator. The distinction that matters most is covered in extraction accuracy vs straight-through rate.
20. Straight-through rate on your document mix — weight 3, differentiator. Strong: a measured percentage of files needing no correction, quoted specifically for scanned company-prepared statements. Weak: a straight-through figure drawn from clean audited PDFs.
21. Integration with the LOS and core — weight 3, table stakes. Strong: a documented API, a named integration with your LOS, and a clear statement of which system owns which record. Weak: CSV export described as an integration. The boundary question is addressed in does credit memo automation replace your LOS.
22. Implementation and configuration ownership — weight 2, table stakes. Strong: a named timeline, a defined configuration workshop, and your team owning the chart of accounts afterwards. Weak: configuration that only the vendor can change, which is a renewal-leverage problem disguised as a support model.
The weighted scorecard
# | Criterion | Weight | Type |
|---|---|---|---|
1 | Company/accountant-prepared statements | 6 | Differentiator |
2 | Degraded scans | 3 | Table stakes |
3 | Notes, schedules and annexures | 5 | Differentiator |
4 | Tax returns as first-class documents | 4 | Differentiator |
5 | Chart-of-accounts mapping (gate) | 9 | Differentiator |
6 | Label independence across borrowers | 7 | Differentiator |
7 | Normalisation rules you control | 6 | Differentiator |
8 | Multi-period consistency, restatement flags | 4 | Differentiator |
9 | Figure-level citation (gate) | 7 | Differentiator |
10 | Discriminating confidence scores | 4 | Differentiator |
11 | Override workflow and diff | 6 | Differentiator |
12 | Ratio definition control | 5 | Differentiator |
13 | Multi-entity and group handling | 6 | Differentiator |
14 | Currency, units, period length | 3 | Table stakes |
15 | Framework and sector coverage | 3 | Table stakes |
16 | Audit trail and version history | 4 | Table stakes |
17 | Model documentation and change control | 4 | Differentiator |
18 | Data residency, retention, security | 3 | Table stakes |
19 | Reproducible accuracy evidence | 3 | Differentiator |
20 | Straight-through rate on your mix | 3 | Differentiator |
21 | LOS and core integration | 3 | Table stakes |
22 | Implementation and configuration ownership | 2 | Table stakes |
| Total | 100 | 21 table stakes / 79 differentiator |
Be honest about that last row. Seven criteria worth 21 points are things any credible vendor should clear — if a product cannot read a skewed scan or produce an audit trail in 2026, the evaluation ends early and cheaply. The 79 points that remain sit on capabilities a large share of the market does not have, and three of them — mapping, label independence and normalisation control, worth 22 points between them — are the same capability seen from three angles. A vendor that scores well on those three is a spreading engine. One that does not is an extraction tool, whatever the brochure says.
Worked example: scoring a hypothetical vendor (illustrative)
Vendor: "Meridian Spread" — a strong document-extraction platform extending into credit. Scores below are from a four-week evaluation on 40 of the lender's own files. Weighted points = weight × score ÷ 5.
# | Criterion | Weight | Score (0–5) | Points |
|---|---|---|---|---|
1 | Company/accountant-prepared statements | 6 | 3 | 3.6 |
2 | Degraded scans | 3 | 4 | 2.4 |
3 | Notes, schedules and annexures | 5 | 2 | 2.0 |
4 | Tax returns as first-class documents | 4 | 3 | 2.4 |
5 | Chart-of-accounts mapping (gate) | 9 | 2 | 3.6 |
6 | Label independence across borrowers | 7 | 2 | 2.8 |
7 | Normalisation rules you control | 6 | 1 | 1.2 |
8 | Multi-period consistency | 4 | 3 | 2.4 |
9 | Figure-level citation (gate) | 7 | 4 | 5.6 |
10 | Discriminating confidence scores | 4 | 2 | 1.6 |
11 | Override workflow and diff | 6 | 3 | 3.6 |
12 | Ratio definition control | 5 | 2 | 2.0 |
13 | Multi-entity and group handling | 6 | 2 | 2.4 |
14 | Currency, units, period length | 3 | 4 | 2.4 |
15 | Framework and sector coverage | 3 | 4 | 2.4 |
16 | Audit trail and version history | 4 | 4 | 3.2 |
17 | Model documentation and change control | 4 | 3 | 2.4 |
18 | Data residency, retention, security | 3 | 5 | 3.0 |
19 | Reproducible accuracy evidence | 3 | 2 | 1.2 |
20 | Straight-through rate on your mix | 3 | 1 | 0.6 |
21 | LOS and core integration | 3 | 4 | 2.4 |
22 | Implementation and configuration ownership | 2 | 3 | 1.2 |
| Total | 100 |
| 54.4 |
Group subtotals:
Group | Weight available | Points scored | % of available |
|---|---|---|---|
Reading the documents | 18 | 10.4 | 57.8% |
Producing your numbers | 26 | 10.0 | 38.5% |
Traceability and challenge | 22 | 12.8 | 58.2% |
Portfolio structures | 12 | 7.2 | 60.0% |
Governance | 11 | 8.6 | 78.2% |
Economics and evidence | 11 | 5.4 | 49.1% |
Total | 100 | 54.4 | 54.4% |
Verdict: decline, on two independent grounds.
First, the total of 54.4 falls below the 60-point band. Second, and decisive on its own, criterion 5 scored 2 — mapping to the lender's chart of accounts required a vendor consultant for each new borrower template — which trips the hard gate.
The shape of the result is the useful part. Meridian scored 78.2% of available points on governance and 60.0% on portfolio structures. It is a well-run, secure, auditable product. It scored 38.5% on the group that determines whether the output is a spread or a data dump — criteria 5, 6 and 7 returned 7.6 of 22 available points. That is the profile of an extraction platform with a credit skin, and no amount of strength elsewhere compensates, because the standardised, comparable spread is the entire reason for buying.
Also note criterion 20 at score 1. Meridian quoted a straight-through rate of 71% but could not reproduce it on the lender's own scanned company-prepared files, where the measured rate was 34%. Given that 87% to 94% of banks evaluate non-audited statements at the loan sizes in question, that is the population that matters.
How do you run the evaluation without wasting a quarter?
Four practical constraints, learned expensively.
Use your own files, at least 40 of them, weighted to your real mix. A demo on the vendor's sample set tells you about the sample set. Include the ugly ones deliberately: the 15-month first period, the statement that does not balance, the group file with four unlabelled entities.
Score in the room, not afterwards. A scorecard completed a week later records impressions. One completed during the session records evidence.
Separate the person who wants it from the person who scores it. The sponsor who has been talking to the vendor for six months should not be holding the pen.
Time-box to four weeks. Beyond that the evaluation becomes an implementation, and the sunk cost starts scoring for you.
For the adjacent question of what to ask a vendor conversationally rather than score, questions to ask a credit memo automation vendor at a demo is the companion piece; nine platforms compared is the market map.
FAQ
How should a commercial lender evaluate spreading software?
Score candidates against 22 weighted criteria across six areas — document handling, standardisation to your chart of accounts, traceability, portfolio structures, governance, and economics — using your own files rather than the vendor's. Apply hard gates on chart-of-accounts mapping and figure-level citation, because a failure on either makes the rest of the score irrelevant.
What is the best financial spreading software for commercial lenders?
There is no single answer, because the weights differ by lender. A community bank underwriting company-prepared statements should weight criteria 1, 3 and 9 heavily; a lender with complex group borrowers should push criterion 13 up. Score the shortlist on your own weights and your own files, and the ranking usually settles itself.
Can financial spreading software read accountant-prepared statements?
The better products do, and it is worth testing specifically, because compilations and reviews have no notes and inconsistent labels. Ask for accuracy and straight-through figures on non-audited statements alone — the FDIC survey shows those are the statements most banks are actually working from at $250,000 and above.
Which criteria are genuinely table stakes?
Seven, worth 21 points combined: degraded scans, currency and units, framework coverage, audit trail, security posture, integration, and implementation model. Any credible vendor clears these. Use them as an early filter, not as a way to distinguish finalists.
Why is chart-of-accounts mapping weighted so heavily?
Because it is the capability that makes a spread comparable across borrowers and across years, and it is the one most often absent. Without it you have the borrower's numbers faster, which is a different and much less valuable product than your numbers, consistently.
How many test files does a fair evaluation need?
At least 40, weighted to your actual mix rather than chosen for cleanliness. Fewer than that and one difficult document swings the result; many more and the evaluation stops being time-boxed.
Should the scorecard weights be the same for every lender?
No. The 100 points here reflect a mid-market commercial lender with mixed document quality. Move weight toward criterion 13 if your book is group-heavy, toward criteria 15 and 18 if you operate across jurisdictions, and toward criteria 19 and 20 if your business case rests on throughput.
What happens if a vendor scores 70?
That is the conditions band. Proceed, but write the gaps into the contract as dated deliverables with a remedy, not as roadmap discussion. A capability that is not contracted is a capability you are hoping for.
Key takeaways
- Twenty-two criteria, six groups, weighted to 100, scored 0–5 on your own files.
- Two hard gates — chart-of-accounts mapping and figure-level citation. A score below 3 on either ends the evaluation.
- Twenty-one points are table stakes and 79 are differentiators. Criteria 5, 6 and 7 alone carry 22 points and decide whether you are buying a spreading engine or an extractor.
- Verdict bands: 75+ proceed, 60–74 proceed with contracted conditions, below 60 decline.
- Score in the room, use 40 files or more, and keep the sponsor away from the pen.
The conceptual grounding sits in what financial spreading is and the mechanics in the spreading process step by step. For the throughput side of the business case, how much analyst time automated spreading saves per file has the arithmetic.
Watch YuSight spread a real balance sheet — book a live demo.