How Do You Evaluate AI Underwriting Software for Examiner Readiness?
Evaluate the vendor the way an examiner will evaluate you. Demand a due-diligence file that satisfies the 2023 Interagency Third-Party Guidance, evidence of measured extraction accuracy against a stated benchmark, a live test on your own worst file, and four contract terms — audit rights, data retention, model change notification and exit. Marketing language is not evidence.
Key facts
- YuSight cites 100% of figures with one-click source verification. In an evaluation, that is not a feature claim — it is a testable assertion. Open a memo, click a number, see whether it lands on the source page.
- The model risk rulebook changed in April 2026. SR 26-2 (17 April 2026) superseded SR 11-7 (2011) and SR 21-8 (2021); the OCC issued it as Bulletin 2026-13, rescinding the "Model Risk Management" Comptroller's Handbook booklet and OCC Bulletins 1997-24, 2011-12 and 2021-19.
- Generative and agentic AI are expressly out of scope. The guidance states: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The agencies signalled a separate request for information on banks' AI use. Confirm whether that RFI has been issued before relying on its timing.
- The scope threshold is $30 billion. The guidance "is expected to be most relevant to banking organizations with over $30 billion in total assets," with smaller institutions generally outside it subject to a carve-back for significant model exposure.
- Third-party guidance still bites regardless. OCC Bulletin 2023-17 (6 June 2023) applies "commensurate with the bank's risk profile and complexity as well as the criticality of the activity supported by the third party." Buying an AI underwriting tool is a third-party relationship whether or not the tool is a model.
- Adoption is mainstream. 49% of banks and 59% of credit unions had already deployed generative AI, across 416 senior executives surveyed (Cornerstone Advisors, 29 January 2026).
This page is about evaluating the vendor. If you need the governance side — which rules apply, whether the tool is a model, what an examiner asks in the room — read what examiner review requires from an AI-drafted credit memo first and come back.
What changed in April 2026, and why does it make vendor evaluation harder?
Counter-intuitively, taking generative AI out of the model risk framework raised the bar on you, not lowered it.
Under SR 11-7 there was a well-worn playbook: call it a model, run it through validation, produce a validation report, done. SR 26-2 narrowed the definition — a model is now "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates," excluding "simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes." Then it removed generative AI by name, and added: "a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document."
That sentence hands the design of the control set back to you. There is no validation template to fall back on. So the evidence you extract from the vendor at diligence stage becomes the substance of your control set — and if the vendor cannot supply it, you cannot build one.
Two practical consequences for evaluation:
- You cannot outsource the classification decision to the vendor. A vendor telling you "we're not a model, so you don't need validation" is answering a question that belongs to your model risk committee. Ask what evidence they will supply either way.
- Third-party diligence is now the primary control, not the secondary one. Bulletin 2023-17 does not care about the model question at all.
What should the vendor due-diligence file contain?
Ask for these as documents, before contract, with a named owner at the vendor. If the answer to any of them is a sales deck, note it.
Company and viability
- Audited or reviewed financial statements, or a written explanation of why not.
- Ownership, funding stage, runway, and any change-of-control events in the last 24 months.
- Client count in US commercial lending specifically, with reference customers at your asset size.
- Business continuity and disaster recovery plan, with the last test date and result.
Security and data
- SOC 2 Type II report (the full report, not the certificate), current, with the bridge letter if the period has lapsed.
- Penetration test summary from the last 12 months, with remediation status.
- Data flow diagram: where borrower documents are stored, where they are processed, in which regions, and by which subprocessors.
- Subprocessor and subcontractor list — including the foundation model providers behind any generative component, and whether inference happens on the vendor's infrastructure or a third party's.
- Written statement on training: is your data used to train or fine-tune any model, ever, including "aggregated and anonymised"? Get it in the contract, not the FAQ.
- Encryption at rest and in transit; key management; who at the vendor can read a borrower's tax return, under what process.
Model and accuracy
- Measured extraction accuracy, the benchmark it was measured against, the document mix in that benchmark, and the date.
- Error taxonomy: what kinds of errors occur, at what rate, and which are silent versus flagged.
- Confidence scoring — does the tool tell the analyst which fields it is unsure about, or present everything with equal confidence?
- Model change management: how often do underlying models change, what regression testing runs before a change ships, and how are customers notified.
- Whether the tool scores, rates, ranks, prices or recommends approval or denial in any respect. Written answer, signed.
Controls and audit
- Sample audit trail export for a single memo — the actual file format an examiner would receive.
- Immutability design: can a prior version be altered or deleted, and by whom.
- Retention: how long versions, citations and source documents remain resolvable, and what happens on termination.
- Role-based access control model, and the admin action log.
- Fair lending posture: written scope statement, and whether the vendor has had any ECOA/Reg B-related finding or inquiry disclosed to it by a customer.
Which tests should you actually run?
Documents tell you what a vendor believes. Tests tell you what the software does. Run these on your own files, in a sandbox, with your own analysts — not on the vendor's demo set.
# | Test | What you feed it | Pass condition |
|---|---|---|---|
1 | Hard file | Your worst real file: 3+ entities, mixed year-ends, scanned and stamped statements, one company-prepared set | Correct entity mapping with no manual assignment; ≤2 analyst corrections per entity |
2 | Citation resolution | Any generated memo | Every numeric assertion clicks through to the correct document, correct page, correct extracted region |
3 | Adversarial document | A statement with a deliberately altered figure, and a page inserted from a different borrower | Tool flags the mismatch or fails loudly; silent acceptance is a fail |
4 | Missing data | A file with a page removed mid-schedule | Tool reports the gap; does not interpolate or infer the missing figure |
5 | Reproducibility | Same file, run twice, 48 hours apart | Identical figures; any difference is explained and logged |
6 | Version diff | Analyst edits three figures, re-generates | Diff shows exactly what changed, who changed it, when, and the downstream ratio impact |
7 | Termination drill | Ask for a full export mid-pilot | You receive memos, citations and source documents in a readable format within the contractual window |
Test 3 and test 4 are the ones vendors dislike, and they are the ones examiners care about, because they distinguish a tool that knows what it does not know from one that fills gaps confidently. See how to stop an AI credit memo tool from hallucinating numbers for the mechanism behind that distinction.
What does a worked scoring evaluation look like?
Score each dimension 0-5, apply the weight, and set a floor on the non-negotiables. Below is a completed example for a hypothetical vendor. Illustrative scores — the framework is the deliverable, not the numbers.
Dimension | Weight | Score (0-5) | Weighted | Floor |
|---|---|---|---|---|
Citation coverage and resolution | 20% | 5 | 1.00 | ≥4 |
Extraction accuracy, measured and benchmarked | 15% | 4 | 0.60 | ≥3 |
Audit trail and immutable version history | 15% | 4 | 0.60 | ≥4 |
Security posture (SOC 2 Type II, pen test, subprocessors) | 12% | 4 | 0.48 | ≥4 |
Contract terms (audit, retention, change notice, exit) | 12% | 3 | 0.36 | ≥3 |
Multi-entity and messy-document handling | 10% | 4 | 0.40 | — |
Model change management and notification | 8% | 3 | 0.24 | ≥3 |
Fair lending scope clarity (drafts vs scores) | 5% | 5 | 0.25 | ≥4 |
Vendor viability and reference depth | 3% | 3 | 0.09 | — |
Weighted total, line by line: 1.00 + 0.60 + 0.60 + 0.48 + 0.36 + 0.40 + 0.24 + 0.25 + 0.09 = 4.02 out of 5.00
As a percentage: 4.02 ÷ 5.00 = 80.4%.
Weights sum check: 20 + 15 + 15 + 12 + 12 + 10 + 8 + 5 + 3 = 100%. ✅
Two rules make this useful rather than decorative. First, floors override the total — a vendor scoring 4.6 overall but 2 on audit trail fails, because the total is a purchasing convenience and the floor is a control. Second, a score of 5 requires evidence in the file, not a demo. Anything demonstrated but not documented caps at 3.
Weights should be set by your own model risk and credit policy committees before you see a single vendor. Setting them afterwards is how a scorecard becomes a justification.
Which contract terms matter most?
Four clauses do the heavy lifting. Everything else is commercial.
1. Access to records and audit rights. Bulletin 2023-17 frames third-party risk management around the relationship life cycle, and your examiner will ask what rights you actually hold. Specify: the right to receive the SOC 2 Type II annually without asking; the right to audit or to commission an audit on notice; the obligation to notify you of any security incident affecting your data within a defined number of hours; and the right to receive subprocessor changes in advance, with a right to object.
2. Data ownership, use and retention. State plainly that borrower documents and derived data are yours, that the vendor obtains no licence to train on them, and that on termination the data is returned in a specified format and then deleted with written certification. Set retention while the contract runs to at least your credit-file retention period — Regulation B's floors are short (25 months consumer, 12 months business, 60 days for business applicants above $1 million in revenue, extended to 12 months on written request under 12 CFR 1002.12), but credit file expectations run life-of-loan and that is what governs in practice.
3. Model change notification. This is the clause almost nobody negotiates and the one that causes the most trouble. Require written notice before any change to an underlying model, extraction pipeline or prompt architecture that could alter output; a stated regression test standard; and a right to test in sandbox before the change reaches production. Without it, your accuracy evidence goes stale silently and your "measured error rate" answer to an examiner becomes untrue on a date you cannot identify.
4. Exit and continuity. Define the export format, the window, the assistance obligation, and — critically — whether citations still resolve after termination. A memo whose 142 citations point at a decommissioned service is a document, not an audit trail. Ask for a self-contained export where source documents travel with the memo.
Clause | Weak version | What to insist on |
|---|---|---|
Audit rights | "Vendor will provide reports upon reasonable request" | Annual SOC 2 Type II delivered automatically; audit on 30 days' notice; incident notice within a stated number of hours |
Data use | "Vendor may use aggregated data to improve services" | No training or fine-tuning on customer data, full stop; written certification of deletion on exit |
Model change | Silent | Prior written notice; regression test standard; sandbox testing right; rollback path |
Exit | "Data available for 30 days post-termination" | Defined format, defined window, migration assistance, citations resolvable offline |
Subprocessors | "Vendor may engage subcontractors" | Named list, advance notice of change, right to object, flow-down of security terms |
What separates a tool that drafts from a tool that decisions?
Get this answer in writing before anything else, because it changes which regulatory conversation you are having. A tool that produces standardised financials, ratios and narrative leaves the credit decision with the analyst and the committee. A tool that emits an approve/decline, a score, a rating or a price has taken part of the decision inside itself, which pulls in fair lending testing, adverse action reason accuracy and — depending on your size and the technique — a live argument about model status.
Ask the question in the vendor's own words and keep the answer: "Does your product score, rate, rank, price, or recommend approval or denial of a credit application in any respect?" A hedged answer is a finding.
YuSight drafts and spreads. It does not score, rate or decision. Whatever you buy, put the equivalent sentence in the contract, because that is the sentence your compliance team will be asked to defend. The distinction runs through the whole commercial underwriting process in US banks, and it is worth mapping against your own credit policy before you shortlist.
What does a complete evaluation timeline look like?
Eight weeks is realistic for a bank between $2 billion and $30 billion in assets. Compress it and you will be assembling the diligence file during the exam instead.
- Weeks 1-2 — Scoping. Model risk, credit policy, compliance and IT agree the weights and floors. Write the requirement, then look at vendors.
- Weeks 3-4 — Documents. Issue the 20-item due-diligence request. Score what comes back before any demo.
- Weeks 5-6 — Testing. Sandbox, your own files, the seven tests, your own analysts. Record every correction.
- Week 7 — Contract. The four clauses. Legal and the business in the same room, because "model change notification" only survives if credit explains why it matters.
- Week 8 — Committee. The classification memo, the scorecard, the test results, the contract summary, the residual risks and the compensating controls, all in one file with a named approver.
That final file is the artifact. When an examiner asks how you selected this vendor, you hand it over rather than reconstructing it. Pair it with the ongoing artifacts described in the credit memo checklist and the bureau report analysis controls, and the file answers most of the exam before it is asked.
Frequently asked questions
How do I evaluate AI underwriting software for examiner readiness?
Run three streams in parallel: a documented due-diligence file that satisfies the 2023 interagency third-party guidance, hands-on tests on your own difficult files, and negotiation of four contract terms — audit rights, data use and retention, model change notification, and exit. Score all three against weights and floors your committees set before you saw a vendor.
What documentation should an AI underwriting vendor supply?
At minimum: a full SOC 2 Type II report, a recent penetration test summary, a data flow and subprocessor list including any foundation model providers, measured extraction accuracy with the benchmark behind it, a model change management policy, and a sample audit trail export in the format an examiner would receive.
Which controls do examiners test first?
Provenance and human review. They pick one loan, point at one number in the memo, and ask where it came from and who checked it. If that takes two clicks you are in good shape; if it takes two days, everything else in the file is discounted.
Is AI underwriting software a model under SR 26-2?
SR 26-2 narrowed the model definition and put generative and agentic AI expressly outside its scope, so for many tools the answer is no. That does not remove the governance obligation — the guidance redirects it to your own risk management practices, so document the classification, who made it and what controls you applied instead.
Does SR 26-2 mean banks under $30 billion can skip this?
No. It means the interagency model risk guidance is generally not addressed to you and that your own practices, sized to your risk profile, govern. Third-party risk guidance, credit administration expectations and fair lending rules apply regardless of asset size.
What extraction accuracy should we require?
Require the number, the benchmark and the date rather than a threshold. A 95% figure measured on clean digital PDFs is weaker than a 90% figure measured on scanned, stamped, multi-entity statements — so ask what was in the test set and re-run it on yours.
Should we require the vendor to be SOC 2 certified?
Ask for the SOC 2 Type II report itself, not a certificate or a logo. Read the exceptions section and the complementary user entity controls, because that section tells you which controls the vendor is assuming you operate.
How do we handle model change notification if the vendor uses a third-party foundation model?
Push the obligation through anyway. The vendor may not control the upstream model, but they control which version they call and when they switch, so require notice of that decision plus regression test evidence against your document types.
What happens to our memos and citations if we terminate?
Whatever the contract says, which in most standard forms is not enough. Insist on a self-contained export where source documents travel with the memo so citations resolve offline, and test that export during the pilot rather than discovering the gap at exit.
Who should own the vendor evaluation internally?
Credit or lending should own the requirement and the testing; model risk or risk management should own the classification decision; IT and information security should own the security review; compliance should own the fair lending scope statement. One named person should own the resulting file.
Key takeaways
- SR 26-2 removed generative AI from the model risk framework and handed control design back to the institution. That makes vendor evidence the foundation of your control set, not a supplement to it.
- Third-party risk guidance applies whether or not the tool is a model. Bulletin 2023-17 is the guidance that bites on a purchased AI underwriting platform.
- Demand documents before demos, and cap any capability that is demonstrated but not documented at a middling score.
- Run the adversarial and missing-data tests. A tool that fills gaps confidently is the one that generates exam findings.
- Negotiate four clauses: audit rights, data use and retention, model change notification, and an exit where citations still resolve.
- Set weights and floors before you meet vendors. A scorecard built afterwards is a justification, not an evaluation.
YuSight's Workflow & Audit Trail keeps a complete, single-platform record of who reviewed what and when, with 100% of figures cited and one-click source verification, full memo version history and a diff between the machine draft and the approved document — the evidence a due-diligence file needs, produced as the work happens rather than reconstructed for the exam.
See the audit trail an examiner would see — book a live demo.
This article is general information for credit and risk professionals, not legal or regulatory advice. Supervisory guidance changes; verify every citation against the issuing agency before relying on it.