What Questions Should You Ask a Credit Memo Automation Vendor During a Demo?
Ask questions with denominators. Not "how accurate are you" but "on which fields, across what document mix, against whose ground truth, and what is the number on your worst 10% of files?" Below: 29 such questions, a scoring table, and a plain statement of where YuSight — 95.2% extraction accuracy against a manual benchmark — is not clearly ahead of anyone.
Key facts
- 29 questions across seven areas: accuracy, provenance, coverage, human control, workflow fit, auditability, commercials. Three of them we do not win — named below.
- The 2023 Interagency Guidance on Third-Party Relationships (88 FR 37920, effective 6 June 2023) directs US banks to assess a third party's business experience, financial condition, information security, operational resilience and subcontractor reliance before contracting — not after the pilot.
- Under the EU AI Act, Annex III point 5(b), AI used "to evaluate the creditworthiness of natural persons or establish their credit score" is high-risk — relevant if any of your book is EU consumer or sole-trader lending.
- YuSight reports 95.2% extraction accuracy, validated against a manual benchmark. Ask us for the denominator too.
What does your accuracy number actually mean?
- Is that figure field-level or document-level? Wildly different numbers.
- What was the sample size and document mix — how many audited financials, scanned bank statements, bureau reports?
- Who created the ground truth — your annotators or the customer's credit team?
- Was the benchmark drawn from live customer files or curated by you?
- What is the accuracy on the worst-scanning decile of that sample?
- How is a partially correct field scored — right value, wrong period; right number, wrong entity?
- When the system is wrong, does it know it is wrong — is there a confidence score, and is it calibrated?
Work the arithmetic in the room. Take 10 files, 120 extracted fields each: 1,200 fields. A vendor quoting 96% field-level accuracy is telling you 1,200 × 0.04 = 48 wrong fields. Ask how they cluster. If those 48 fall across 6 of the 10 documents, then documents with zero errors = 4, and document-level accuracy is 40% — from the same "96%." Your analyst reviews documents, not fields. Insist on the second number.
Can every figure be traced to a source page?
- Click any number in the draft memo. Does the source document open at the correct page, in one click, without leaving the memo?
- What share of figures carry a citation — and is a computed ratio cited back to its inputs?
- When an analyst edits a figure, does the citation break, persist, or get flagged as human-overridden?
- Can you export the memo with citations intact for a credit committee pack?
Which documents can you actually handle?
- Name the document types supported, separating "extracts reliably" from "accepts the upload."
- What happens with scanned, photographed, rotated or password-protected files?
- Which languages and scripts, and what is the accuracy delta versus English?
- On a multi-entity group, does it map each document to the right obligor or dump them in one pile?
- Multi-period files with restated comparatives: which year wins, and is the restatement flagged?
- Behaviour on a document type it has never seen — silent guess, or explicit "unsupported"?
What can my analyst override, and does the system record it?
- Which fields are editable, and which are locked?
- Is every override written to an immutable log with user, timestamp, prior value and new value?
- What routes to human review automatically, on what threshold, and can we set that threshold?
- Can we block approval until every low-confidence field has been touched by a human?
Does this replace my LOS or sit alongside it?
- Is this a system of record or a system of analysis? Wanting to be the record is a much bigger project.
- Which LOS and core systems are you integrated with in production — named customers, not a logo wall.
- What does integration actually mean: API, SFTP drop, or manual export?
What would an examiner see?
- Show the version history of a memo that changed three times: who changed what, when, why.
- Can we export the full audit trail for one file, unaided, in a format we control?
- What are the data retention and deletion rules, and are they configurable per jurisdiction?
What am I actually signing?
- Is pricing per seat, document, memo or decision — and what happens when volume doubles?
- On termination, do we get our data and memos back, in what format, and how long do you retain copies?
Where is YuSight not clearly ahead?
Three honest answers:
The hardest 10% of files (Q5). Nobody has solved poor scans, handwritten annotations and photographed statements. Our 95.2% is a blended number; on genuinely bad inputs every engine in this category degrades, ours included. Push us hardest here, and disbelieve anyone who will not concede the same.
LOS integration depth (Q22–24). We integrate by API and sit alongside the LOS rather than replacing it. So do most credible competitors. Unless a vendor has a pre-built, in-production connector to your specific LOS version, treat every integration claim — ours included — as a project with a timeline, not a feature.
Commercials and implementation time (Q28–29). No vendor has a technical moat in pricing, and implementation duration depends more on your data hygiene and IT calendar than on the software. A shorter quoted timeline is not evidence of a better product.
We are stronger on provenance, multi-entity document mapping and India/UAE document sets (GST and VAT returns, AECB and CBRB bureau data, trade licences). A vendor purpose-built on US tax forms may beat us on Form 1120S line-item nuance. Test it; do not take our word.
Scoring table: take this into the demo
Question area | A strong answer sounds like | A weak answer sounds like |
|---|---|---|
Accuracy | "94.1% field-level on 3,200 fields across 40 files, 30% scanned; 81% on the worst decile" | "Over 99% accurate" |
Ground truth | "Annotated by the customer's credit team, not us" | "Industry-standard benchmarks" |
Provenance | Clicks a number; source page opens, on your file | Shows a screenshot of a citation |
Edits | "Override logged with user, timestamp, old and new value — here it is" | "Yes, it's fully editable" |
Coverage | "These types work; these do not. Here is what we reject" | "We handle anything you throw at us" |
Routing | "Threshold is configurable; below 0.85 it queues for review" | "The AI knows when it's unsure" |
Integration | Names three production customers on your LOS | "We have an open API" |
Audit | Exports the full trail live, unaided | "We can arrange that for you" |
Exit | "Full export in CSV and PDF within 30 days" | "That's in the MSA" |
Red flags
- An accuracy percentage with no denominator. "We're 99% accurate" is marketing until you hear the sample.
- A demo run only on the vendor's sample files. Those have been rehearsed on for two years.
- Refusal to run your documents live. "Send them over and we'll come back next week" means the demo was not the product.
- A citation feature shown but never clicked. Make them click it, on your file, on a number you choose.
- "Our model is proprietary" used to deflect every question about error behaviour.
- No named production customer on your LOS, paired with a confident integration timeline.
- A comparison chart with no stated criteria — including ours. Ask for the methodology; if there is none, discard the chart.
Bring ten of your own files — including the two worst
Pick eight ordinary files and two you would not wish on anyone: the photographed bank statement, the multi-entity group with restated comparatives, the audited financials scanned at an angle. Send nothing in advance.
The eight good files tell you nothing — every vendor here handles clean, machine-readable PDFs. Your cost sits in the 10–15% that stall for two days while an analyst rekeys by hand. Only the worst files discriminate between vendors, because they are the only ones where behaviour differs: one system fails loudly and routes to review, another silently produces a plausible wrong number. Only a live run shows which you are buying.
On governance: SR 26-2, Revised Guidance on Model Risk Management (17 April 2026), superseded SR 11-7 and narrowed expected applicability to organisations above roughly $30 billion in assets — but lighter model-risk expectations do not lighten your evidence burden at credit committee.
Related reading
FAQ
What questions should I ask a credit memo automation vendor during a demo?
Ask questions that force a denominator: field-level versus document-level accuracy, the sample the number came from, and performance on the worst decile. Then make them click a citation on your own file.
How do I evaluate AI underwriting software for examiner readiness and audit trail?
Ask to export a complete audit trail for one file yourself, live, without vendor help. If you cannot produce version history, overrides and source citations on demand, you are not examiner-ready.
Does credit memo software replace underwriter judgment?
No. It removes the extraction and formatting work — spreading, ratio computation, drafting, citation. The credit decision, the risk narrative and the structure remain human, and any vendor implying otherwise is selling you a governance problem.
How long should a credit memo software evaluation take?
Four to eight weeks for a mid-sized lender: two weeks of demos, three to four weeks of POC on your own files, security and legal review in parallel. Past twelve weeks usually means nobody owns the decision.
Should we run a paid pilot or a free trial?
Prefer a short paid pilot. A free trial gets a shared solutions engineer and the vendor's own files; a paid pilot gets a named contact and the right to insist your document types be handled.
How many documents should we test in a POC?
At least 50 from your actual mix, with 20% drawn deliberately from the difficult tail. Below 30 files, the accuracy figure you compute has no confidence interval worth quoting.
What is a fair accuracy benchmark to ask for?
Field-level accuracy in the low-to-mid 90s on a representative mix, with the worst-decile figure disclosed separately. A vendor who volunteers the worst-decile number unprompted is telling you something about their honesty.
Key takeaways
- Every accuracy claim needs a denominator, a document mix and a worst-decile figure. Without all three, the percentage is decoration.
- Provenance is testable in ten seconds: click a number on your own file and see whether the source page opens.
- Bring ten of your files, two of them terrible, and send none in advance.
- Ask every vendor — us included — where they are not ahead. The ones who can answer are worth shortlisting.
See your first CAM in 30 minutes — book a live demo. Bring your ten files. Bring the two worst ones.