Underwriting Automation in Commercial Lending: The Complete 2026 Playbook
Underwriting automation means machine-handling the extraction half of a commercial credit file — classification, spreading, ratios, covenant tests, exception flags, first-draft narrative — and leaving the judgement half to a credit professional. Done to the whole chain it supports 5x throughput at the same headcount. Done to extraction only, it caps at roughly 1.8x, and this playbook shows the arithmetic.
Key facts
- Roughly 45% of the analyst minutes on a mid-market commercial file are mechanical. On our stage-by-stage estimate below, 265 of 595 minutes carry no judgement. That share sets a hard ceiling on any extraction-only tool:
1 ÷ (1 − 0.45) = 1.8x. - Straight-through processing does not scale up the ticket. The FDIC found that 28% of large banks would auto-approve a small loan, 11% a medium loan and 1% a large loan (FDIC, *Small Business Lending Survey 2024*, Section 3). The curve does not bend at $1m; it collapses.
- Model risk guidance changed on 17 April 2026. SR 11-7 and SR 21-8 are superseded by SR 26-2 / OCC Bulletin 2026-13, which narrows a model to "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates" and states that "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance" (Federal Reserve, SR 26-2; OCC Bulletin 2026-13).
- The guidance is aimed at the large end. SR 26-2 "is expected to be most relevant to banking organizations with over $30 billion in total assets." Below that threshold it is a reference point, not a checklist.
- Buying does not move the accountability. Third-party arrangements sit under OCC Bulletin 2023-17 / interagency third-party risk guidance (6 June 2023), which expects risk management "commensurate with the bank's risk profile and complexity as well as the criticality of the activity supported by the third party."
Where do the minutes actually go in a commercial credit file?
Take one unit of work: a $2m–$10m secured request, three years of financial statements plus matching returns, one operating entity, one holding entity, one guarantor, twelve months of bank statements. Not a renewal, not a $250k equipment note.
# | Stage | Nature | Minutes |
|---|---|---|---|
1 | Intake, collation, building the missing-items list | Mechanical | 30 |
2 | Document classification and entity mapping | Mechanical | 25 |
3 | Extraction and keying — 3 years × 2 entities + guarantor | Mechanical | 70 |
4 | Tie-out and cross-check (footing, opening-to-closing, return-to-statement) | Mechanical | 35 |
5 | Ratio computation — DSCR, leverage, liquidity, coverage | Mechanical | 20 |
6 | Covenant and policy testing against the grid | Mechanical | 20 |
7 | Exception flagging, stale-data and completeness checks | Mechanical | 15 |
8 | First-draft financial narrative (what moved, by how much) | Mechanical | 50 |
| Extraction half subtotal |
| 265 |
9 | Recasting and normalisation decisions (add-backs, shareholder loans) | Judgement | 30 |
10 | Purpose, structure and repayment-source analysis | Judgement | 40 |
11 | Industry and market view | Judgement | 30 |
12 | Risk factors and mitigants | Judgement | 45 |
13 | Collateral and guarantor assessment | Judgement | 30 |
14 | Risk rating and rationale | Judgement | 20 |
15 | Policy deviation call and recommendation | Judgement | 35 |
16 | Credit officer review | Judgement | 60 |
17 | Rework and re-issue after review | Judgement | 40 |
| Judgement half subtotal |
| 330 |
| Total |
| 595 |
595 minutes is 9 hours 55 minutes of analyst-and-officer touch time. It is not elapsed time — a file that takes ten hours of work routinely takes four weeks of calendar, because most of the calendar is queue. The stage-by-stage elapsed picture is in commercial loan underwriting in US banks; the spreading slice on its own is timed in how much analyst time automated spreading saves.
Why does automating extraction alone cap your speed-up?
This is the commonest failure in underwriting automation, and it is arithmetic, not opinion.
Let p be the fraction of the work a tool touches and s the speed-up it achieves on that fraction. Total speed-up is:
Total = 1 ÷ ( (1 − p) + p/s )
As s goes to infinity — a perfect tool, instant, free — the ceiling is 1 ÷ (1 − p). That is Amdahl's law, and it is indifferent to how good your vendor is.
With our file: p = 265 ÷ 595 = 0.445.
- A tool that is 10x faster on extraction:
1 ÷ (0.555 + 0.0445) = 1 ÷ 0.5995 = 1.67x - A tool that is infinitely fast on extraction:
1 ÷ 0.555 = 1.80x
So the gap between a good extraction tool and a magical one is 1.67x versus 1.80x — about eight minutes on a ten-hour file. If your business case promised 4x from a document extraction purchase, the business case was wrong before the contract was signed. The brief version, worth memorising: fix 40% of the work and you cannot beat 1.67x.
Getting past the ceiling means attacking the judgement half — not by automating judgement, which you must not do, but by removing the friction around it. Review and rework are 100 of the 330 judgement minutes, and neither is judgement. A reviewer spends most of that hour re-performing tie-outs and hunting for the page a number came from. Cite every figure to its source document and page, and that hour becomes twenty minutes without a single credit decision changing hands.
A fully worked example
Same file. Four changes, each with an honest reduction.
Line | Before (min) | After (min) | Why |
|---|---|---|---|
Extraction half (stages 1–8) | 265 | 30 | Classification, extraction, ratios, covenant tests and first-draft narrative machine-produced; analyst reviews exceptions only |
Judgement, stages 9–15 | 230 | 230 | Unchanged. This is the work you are paying an analyst for |
Credit officer review (16) | 60 | 20 | Every figure one click from its source page; no re-performance of tie-outs |
Rework (17) | 40 | 15 | Fewer arithmetic and consistency defects reach review |
Total | 595 | 295 |
|
Speed-up = 595 ÷ 295 = 2.02x
Minutes saved = 595 − 295 = 300 min = 5.0 hours per file
Now convert to throughput. Assume an analyst has 7 productive hours a day, 20 days a month, and spends 60% of it on files rather than meetings, chasing and portfolio work:
Capacity = 7 × 60 × 20 × 0.60 = 5,040 minutes per month
Files before = 5,040 ÷ 595 = 8.5 files per analyst per month
Files after = 5,040 ÷ 295 = 17.1 files per analyst per month
Throughput gain = 17.1 ÷ 8.5 = 2.0x
Be honest about the gap to 5x. Two times is what the arithmetic gives on a bespoke mid-market file where the judgement half is genuinely thick. To reach 5x you would need 119 minutes per file, which implies judgement collapsing to about 90 minutes — not credible for a $5m structured facility. It is credible for renewals with no material change, programme lending, and small-ticket secured requests where stages 10–13 are a paragraph each. The 5x figure describes the whole chain across a realistic portfolio mix, not a single complex deal. Ask any vendor quoting a multiple which segment it came from.
What is genuinely automatable, and what is judgement?
The line is not "hard versus easy". It is whether the output is a fact or a position.
Task | Verdict | Why |
|---|---|---|
Document classification | Automatable | A tax return is a tax return. Deterministic answer, verifiable against the artefact |
Entity mapping in multi-entity structures | Automatable with review | Machine proposes, analyst confirms. Wrong mapping poisons everything downstream — see multi-entity document mapping |
Line-item extraction and spreading | Automatable | Fact recovery from a document — the mechanics are in what financial spreading is. Every figure must carry a page citation |
Ratio computation | Automatable | Arithmetic on agreed definitions |
Covenant testing | Automatable | A defined test against an extracted number, on a schedule |
Exception and completeness flagging | Automatable | Rule evaluation. Flagging is not deciding |
First-draft financial narrative | Automatable | "Revenue fell 11.4% while gross margin held at 31%" is a description, not a view |
Recasting and add-back decisions | Judgement | Whether the owner's $180k salary is compensation or distribution is a position with a written basis |
Purpose and repayment source | Judgement | "Working capital" is not a purpose. Requires knowing the business |
Structure and pricing | Judgement | Tenor, amortisation, covenants, relationship economics |
Risk factors and mitigants | Judgement | Identifying what could go wrong, and whether the mitigant is real |
Risk rating | Judgement | A scorecard can propose. A credit officer owns it |
Policy deviation | Judgement | By definition the case for departing from the rule |
The recommendation | Judgement | Someone's name goes on it |
The test to apply to any vendor claim: if the output is wrong, is it wrong in a way anyone can check against a document? Extraction fails that way — you open page 14 and see it. Judgement does not. Which is why the honest deployment posture is: automate the checkable, and make the uncheckable easier to defend. What that looks like in memo form is set out in what a credit assessment memo is and, for Indian formats, what a credit appraisal memorandum is.
Does straight-through processing work in commercial lending?
In consumer and small-ticket business lending, yes. Above that, no — and it is worth being precise about why, because "we'll do what the consumer team did" is a board-level mistake that gets made every year.
Segment | STP realistic? | Binding constraint |
|---|---|---|
Unsecured consumer | Yes | Bureau score plus income verification is the whole file |
Small-ticket business, under ~$250k | Partly | Bureau plus bank statements; standard product; thin structure |
Renewal, no material change | Partly | Prior credit already exists; only the deltas need judgement |
Mid-market C&I, $1m–$25m | No | Everything below |
Structured, sponsor-backed, CRE construction | No | Everything below, squared |
Four constraints, in order of how often they break the STP business case:
Heterogeneous documents. A consumer file has three document types. A mid-market file has audited statements, reviewed statements, compilations, tax returns with K-1s, management accounts, AR and AP agings, inventory listings, debt schedules, rent rolls, and a personal financial statement in whatever format the guarantor's accountant favours. Extraction technology handles this — see OCR vs IDP vs LLM extraction — but "handled" means classified and extracted with citations, not decided.
Entity complexity. The borrower is rarely the whole obligor group. Real estate in a sister LLC, an ESOP holding, a trust, a foreign parent, related-party rent that flatters one entity and burdens another. There is no rule set that reliably infers a consolidation perimeter from the documents alone.
Relationship pricing. The decision is not about this facility. It is about the deposit balance, the treasury mandate, the two existing loans and the referral flow. That information is not in the credit file and often not in any system.
Non-standard structures. Covenants with a step-down, an earnout, a subordination, a springing guarantee. The moment terms are negotiated rather than selected, the decision leaves the reach of a rules engine.
The FDIC numbers quoted above are the empirical version of this argument: 1% auto-approval at $1m+. That is not technology lagging. That is the market pricing the risk of automating a position.
How do you design human-in-the-loop underwriting that actually saves time?
Human-in-the-loop is often implemented as "the analyst checks everything the machine did", which reproduces the manual cost and adds a licence fee. Three design decisions separate a working implementation from that.
1. Confidence thresholds that route, not thresholds that reassure. Every extracted field carries a confidence score. Set a threshold per field class, not one global number:
- High-consequence, low-ambiguity (total revenue, total debt, interest expense, cash): route to a human below ~0.95
- High-consequence, high-ambiguity (line classification into your chart of accounts, entity assignment): route below ~0.90 and always route where two entities have similar names
- Low-consequence (address lines, document dates): route below ~0.80
2. Route exceptions, not documents. The unit of human work should be a flagged field with its context, not a 200-page PDF with a note saying "please check". A reviewer handed a queue of 11 flagged fields clears them in minutes. A reviewer handed the file re-does the file.
3. Put everything needed for the decision on one screen. The reviewer needs, without navigating away: the extracted value, the source page image with the figure highlighted, the prior-year value for comparison, the rule that fired, and an edit box that writes an audit entry. If resolving one exception takes three clicks and a document search, you have moved the bottleneck, not removed it. The 60-to-20-minute review reduction in the worked example above depends entirely on this and on nothing else.
What of credit policy can be written as code?
More than most credit teams expect, and less than most vendors imply.
Encodable, cleanly:
- Eligibility gates — industry codes, entity age, geography, legal form
- Hard limits — maximum exposure, LTV caps, tenor limits, house and concentration limits
- Ratio floors and ceilings — minimum DSCR, maximum leverage, minimum current ratio, with the exact formula and the exact input definitions
- Document requirements by product, amount and entity type, including recency rules
- Approval authority matrices — who can approve what, at what amount, at what rating
- Covenant test schedules and cure periods
Resistant to encoding:
- "Satisfactory management" and "acceptable industry outlook"
- Whether a mitigant offsets a weakness, and by how much
- The weighting between a strong guarantor and a weak operating company
- Whether an exception is a one-off or the start of a pattern
- Anything requiring knowledge of the borrower held only in the relationship manager's head
There is a second-order benefit that credit heads tend to value more than the automation itself: encoding forces the policy to be unambiguous. Most credit policies contain a DSCR definition that three analysts compute three ways. You cannot code that. Resolving it is a policy improvement that survives the technology. The ratio definitions worth pinning down first are catalogued in credit analysis ratios, and covenant mechanics in covenant monitoring in commercial lending.
Should underwriting automation replace the LOS or sit alongside it?
Alongside, in almost every case. The loan origination system owns the application record, the workflow states, the approval routing, the document repository and the boarding handoff to core. Replacing it is a multi-year programme with a business case that rarely survives contact with the integration inventory.
What sits alongside it is the analysis layer: document intelligence, spreading, analyzers, memo generation, audit trail. The integration contract is narrow — documents in, structured financials, ratios, exceptions and a completed memo out, written back to the LOS record. The full comparison is in does credit memo automation replace your LOS.
Buy versus build, with honest numbers. The build case looks attractive because the first 70% is genuinely tractable — a competent team gets a classification and extraction pipeline working on your document set in a quarter. The cost sits in the remaining 30% and, more, in the years after.
Cost line | Build | Buy |
|---|---|---|
Initial delivery | 4–8 engineers, 9–18 months | 8–16 weeks configuration |
Extraction accuracy tail | Continuous — new formats break it | Vendor's problem, contractually |
Template and format drift | Ongoing engineering | Included |
Audit trail and citation infrastructure | Built once, maintained forever | Product feature |
Regulatory change | Your policy team plus your engineers | Vendor tracks, you validate |
Key-person risk | High. The two people who understand it leave | Contractual continuity |
Exit cost | Sunk | Migration, contained |
Build if extraction is genuinely core to your differentiation and you already run a platform engineering function. Buy if it is a cost line. The tie-breaker question: who fixes it at 6pm on a Friday when a lender changes its statement format and every file in the queue stalls? If you cannot name the person, buy.
How do you get credit teams to adopt it?
The stated objection is accuracy. It is almost never the real one. Analysts who have spent six months checking a tool that is right 95% of the time stop worrying about accuracy and start worrying about two other things, and neither is addressed by a demo.
Accountability. "If the machine spread it wrong and I signed the memo, whose finding is it?" This is a fair question with a clear answer, and the answer must be operational, not reassuring: the analyst owns the file, the platform makes ownership defensible by citing every figure to a source page so that verification takes seconds, and the audit trail records who accepted what and when. Under the 2026 guidance this matters more, not less — a generative drafting tool sits outside SR 26-2's model perimeter, which means it falls to general safety-and-soundness, internal control and third-party expectations instead. Examiners will not ask you to fit it in a model inventory; they will ask who validated the output. What they actually ask is in what examiner review requires from an AI-drafted credit memo and what questions examiners ask about AI underwriting.
Job security. Do not answer this with "it frees you for higher-value work" — credit teams have heard it and correctly read it as evasion. Answer it with the throughput arithmetic above, in the open: the plan is 17 files per analyst instead of 8.5, at the same headcount, because the pipeline is constrained and the bank wants the volume. If the plan is actually headcount reduction, say so; analysts find out either way, and the ones who find out second stop cooperating with the rollout.
Three practical moves that work: let analysts override any machine output without asking permission and log the override; publish the override rate back to the team as a quality metric on the tool rather than on the analyst; and pick your two most sceptical senior analysts as the pilot cohort, not your two most enthusiastic.
How do you measure underwriting automation — and what are the traps?
Metric | Definition | The trap |
|---|---|---|
Turnaround time | Complete application to credit decision | Includes queue and borrower delay. Improves when documents arrive faster, which has nothing to do with your tool. Measure analyst-touch minutes separately |
Files per analyst per month | Completed files ÷ FTE | Mix shift. Ten renewals is not ten new-money structures. Band by complexity or the number is meaningless |
Rework rate | Files returned by review ÷ files submitted | Goes up first. Reviewers scrutinise machine output harder for the first quarter. Expect a spike and do not panic |
Exception rate | Fields routed to a human ÷ fields extracted | Tunable to whatever you want. A low rate can mean good extraction or thresholds set too loose. Always report it alongside post-review error rate |
Time to first draft | Document intake to complete draft memo | The cleanest single measure of the extraction half, and the one most resistant to gaming |
Post-review error rate | Errors found after approval ÷ files | The metric nobody wants to publish and the only one that proves the thing is safe |
Baseline before you start. Twenty recent files, timed by stage, medians by complexity band. Two weeks of logging gives you a number you can defend at credit committee — which is more than any vendor benchmark will.
What does a realistic implementation sequence look like?
Phase | Focus | What "done" means |
|---|---|---|
Days 0–90 | One segment, one product, extraction and spreading only. Baseline measured. Ratio and covenant definitions pinned down in writing | Machine-produced spread with full citations for every file in the pilot segment; analyst override logged; time-to-first-draft measured against baseline |
Days 91–180 | Exception queue and confidence thresholds calibrated on the pilot's labelled data. First-draft memo narrative. Review screen with one-click source verification | Review minutes down measurably; exception rate stable; rework spike resolved |
Days 181–365 | Second and third segments. Encoded policy gates. LOS write-back. Covenant testing on the existing book | Throughput gain visible at portfolio level; audit trail complete enough to hand an examiner unassisted |
Do extraction and spreading first, always — not because it is the biggest prize (it is capped at 1.8x, as shown) but because it is the only stage where correctness is objectively checkable, which is how the credit team builds the trust that the later phases need.
What actually fails, and why?
- Automating extraction and declaring victory. The ceiling argument, ignored. Turnaround barely moves, the business case misses, and the programme loses its sponsor in month nine.
- Piloting on the easy segment. Clean single-entity files prove nothing. The pilot must include the multi-entity mess or the rollout discovers it at scale.
- No baseline. Without measured pre-state, every post-state number is contestable, and someone will contest it.
- Thresholds set once. Confidence cut-offs chosen at go-live and never recalibrated drift into either an unusable exception queue or silent errors.
- Buying accuracy instead of traceability. A 97% tool without citations is worse in practice than a 94% tool with them, because the reviewer cannot verify the 97% cheaply. Extraction accuracy claims and how to test them are unpacked in how to evaluate AI underwriting software for examiner readiness.
- Letting the machine write the recommendation. The fastest way to lose your credit committee, and the one failure that is also a governance finding.
- Treating adoption as training. It is an accountability question. Training does not answer it.
FAQ
What is underwriting automation?
It is machine-handling the parts of a credit file that produce facts — document classification, extraction, spreading, ratio computation, covenant testing, exception flagging and a first-draft narrative — while a credit professional keeps the parts that produce positions. It is not automated decisioning.
Which parts of underwriting can actually be automated?
Anything whose output can be checked against a document. Extraction, spreading, arithmetic, rule evaluation and description all qualify. Purpose, structure, mitigants, deviations and the recommendation do not, because being wrong about those is not visible on a page.
How do lenders keep judgement in an automated process?
By routing exceptions rather than documents, by letting an analyst override anything without asking, and by logging every override. If the analyst cannot change a machine output in one click, judgement has already left the process regardless of what the policy says.
Does straight-through processing work for commercial loans?
For small-ticket and renewals with no material change, partly. For mid-market and above, no. The FDIC found only 1% of large banks would auto-approve a large loan. Heterogeneous documents, entity complexity, relationship pricing and negotiated structures each defeat it independently.
Does SR 11-7 still apply?
No. SR 11-7 and SR 21-8 were superseded on 17 April 2026 by SR 26-2 and OCC Bulletin 2026-13. The definition of a model narrowed, and generative and agentic AI were placed expressly outside scope. The guidance is aimed mainly at organisations above $30 billion in assets.
If generative AI is out of scope for SR 26-2, is it unregulated?
No — it moves, it does not disappear. Out of the model perimeter means into general safety and soundness, internal control, third-party risk and consumer protection expectations. An examiner will still ask who validated the output and whose name is on the decision.
How much faster can underwriting realistically get?
On a bespoke mid-market file, about 2x on analyst-touch time — 595 minutes to 295 on our worked example. Higher multiples are real in segments where the judgement half is thin: renewals, programme lending, small-ticket secured. Ask any vendor which segment their multiple came from.
Should we replace our LOS?
Almost certainly not. The LOS owns the application record, workflow and boarding. Put the analysis layer alongside it and write the results back. Replacement programmes are multi-year and rarely survive the integration inventory.
Why does our rework rate go up after go-live?
Because reviewers check machine output harder than they check a colleague's for the first quarter. It is a trust curve, not a quality problem, and it usually resolves by month four. Watch post-review error rate instead — that is the one that matters.
What should we automate first?
Document classification, extraction and spreading, on one segment, with a measured baseline. Not because it is the biggest prize — it is capped at roughly 1.8x — but because it is the only stage where correctness is objectively checkable, and that is what earns the credit team's trust for everything after.
Key takeaways
- The extraction half of a mid-market commercial file is roughly 45% of the minutes. Automating it perfectly caps total speed-up at 1.80x; a realistic tool gives 1.67x.
- Getting past the ceiling means attacking review and rework — not by automating judgement, but by citing every figure so verification takes seconds.
- STP works in consumer and small-ticket. It does not transfer to mid-market commercial, and the FDIC's 1% auto-approval figure at $1m+ is the evidence.
- Encode eligibility, limits, ratio floors, document requirements and authority matrices. Do not try to encode mitigants, weightings or the recommendation.
- Sit alongside the LOS. Buy unless extraction is core to your differentiation.
- The adoption objection is accountability and job security. Answer both in the open, with the arithmetic.
- Baseline first, twenty files, timed by stage. Everything downstream depends on it.
See the audit trail an examiner would see — [book a live demo](https://yuverse.ai/yusight).