I led a buy-side QoE in the first quarter of 2026 on a $65M-revenue consumer brand, and that engagement was the first time we used AI across the full diligence workflow rather than a single workstream. The deliverable was still hand-built by the QoE team. The path to the deliverable was meaningfully accelerated by agent work, and the engagement also surfaced the specific places where AI cannot be trusted to land the answer on its own. The McKinsey numbers — 20% average cost reduction, 30-50% faster deal cycles for the teams using gen AI in M&A — match what we measured against a comparable engagement six months earlier, but the saving lands in different places than the vendor decks suggest. The most material gains were in the first 72 hours of the engagement and in the reserve-adequacy first-pass; the back half of the QoE looks almost identical to the all-human version. Here is what actually happened, in the order it happened, and where the discipline matters most.
01 Data-room ingestion: the first 48 hours
The seller opened the data room on a Wednesday afternoon. By Friday morning the agent had ingested every document — about 3,400 files across 18 top-level folders — classified each into the QoE workstream taxonomy (Revenue, COGS, Opex, Working Capital, Legal, HR, Tax, IT) and produced a structured catalogue mapped against our standard 187-line diligence request list. Output: a labelled document universe with confidence scores, a deal index reorganised by workstream rather than by the seller's folder structure, and a Gap Matrix flagging every request-list item as Satisfied, Partially satisfied, or Not found. The QoE lead read the catalogue Friday afternoon. Two years ago that triage took a junior analyst the better part of a week.
The Gap Matrix mattered more than the indexing. Of 187 request-list items, the agent flagged 41 missing and 28 partially-satisfied (right document type, wrong period coverage, or missing dimension). By Monday morning we had filed a prioritised Q&A list back to the seller covering only those 69 items, sequenced by analytical impact: missing AR aging for FY-2 blocks DSO trend analysis, missing payroll for two subsidiaries blocks wage-inflation assessment, missing trade-promotion calendar by retailer blocks the deduction-reserve work. The shape of the engagement was set in three working days instead of two weeks.
The agent did not do the QoE. It pulled forward the moment we could see what the QoE needed to address — and the seller could see what they needed to produce.
Two operating disciplines matter here. First, the agent ran inside a VDR-integrated, firm-controlled environment with logging and access controls — not a consumer-grade LLM with confidential deal data pasted in. The Datasite survey finding that 36% of dealmakers cite data security as the top AI concern in M&A is the right concern; the answer is environment governance, not abstinence. Second, the QoE manager spot-checked 15% of classifications before trusting them downstream. Two of the 18 top-level folders had been mis-classified at the document level in ways that would have biased our financial workpapers; the spot-check caught both.
02 Contract metadata extraction and concentration analytics
The brand had 14 retail-customer contracts, 9 large vendor contracts, and 31 secondary commercial agreements (warehousing, fulfilment, tech, marketing) in the data room. The agent extracted structured metadata from each: counterparty, term, renewal mechanics, termination clauses (for convenience vs for cause, notice period), pricing model (fixed, tiered, indexed, MFN), change-of-control language, exclusivity, SLAs, penalties. Output: a flat table — 54 rows, 22 columns — built in twenty minutes automatically plus an hour to spot-check high-impact rows. Building the same table on a comparable 2023 deal was a two-analyst, ten-day exercise.
The concentration view was where the table earned its keep. The top three retail customers represented 47% of fiscal-2025 revenue. Two of the three contained change-of-control clauses requiring written consent within thirty days of the transaction; the third had a unilateral termination-for-convenience clause with a sixty-day notice period. None was hidden, but the agent surfaced the concentration-plus-termination overlay as a structured deliverable rather than something noticed in week three of contract review. The buy-side principal had the customer-concentration risk profile in hand for the Friday IC preview, one week into the engagement.
The honest qualifier: we manually re-read each of the top-five customer contracts and the top-three vendor contracts before the QoE report went out. The agent's extraction was 94% accurate on a hand-checked sample of 200 fields — roughly 12 fields needed correction. Some mattered (a misread on an MFN clause that would have changed a pricing-step-down assumption); most were minor. Treating the table as a starting point rather than a finished product was the discipline; the time saving versus building from scratch was still north of 90%.
03 Reserve-adequacy screening: returns, deductions, inventory
The reserve-adequacy work is the part of consumer-brand QoE that drives the most adjustment dollars and that I was most curious to test AI inside. Across the eighteen consumer-brand QoEs we have advised on since 2022, the median EBITDA adjustment from returns-reserve restatement runs 14%, from trade-deduction restatement 7%, and from inventory standard-cost restatement 4% — each one a number with a defensible calculation behind it. The question we set going in was whether the agent could run a credible first-pass on each of the three lines against the source data and produce numbers that the QoE lead could refine rather than rebuild.
Returns reserve
The agent pulled trailing-24-month returns data at order/SKU/channel level (about 410,000 rows against 2.1 million order lines), built return curves by SKU-family and channel time-bucketed at 30/60/90/180 days, applied curves to the unexpired sales-cohort pool, and produced base-case and 90th-percentile-stress implied reserves. Base case ran $1.84M against booked $1.61M — a $230K gap, roughly 11% of headline EBITDA. The QoE lead flagged that the agent had used a uniform curve across DTC and wholesale where wholesale return rights are contractually different, refined the model to split the channels properly, and re-ran. The refined number landed at $1.92M and a $310K adjustment — 15% of EBITDA. Three rounds of human-agent iteration converged in four working days of QoE-lead time, versus the eight-to-ten days the all-human version had taken on a similar deal in mid-2025.
Trade-promotion deduction reserve
The deduction work was harder for the agent and ultimately more valuable. The agent ingested the brand's promotion calendar (47 active retailer promotions during the trailing-12), the trade-promotion accrual subledger, and the actual deduction-claim ledger by customer. It matched roughly 78% of deduction claims to a specific promotion through retailer code, dollar-amount banding, and timing. The remaining 22% — mostly compliance fines, slotting true-ups, and post-audit chargebacks — were flagged unmatched and routed to manual review. First-pass implied accrual ran $640K against booked $480K — an under-accrual of $160K, about 8% of EBITDA, in line with our cross-deal median of 7%.
The agent was useful on the matching — the 78% match rate would have taken a senior associate three to four days manually. The agent could not be trusted on the unmatched 22%: each one needed a human read against retailer-specific history and open-claim status, because two of the larger unmatched claims were on appeal and the seller had a credible argument for partial recovery. That nuance was not in the data room; it came out of a management interview in week two. The agent surfaced the question; the human answered it.
Inventory standard cost
The inventory standard-cost screen is the line that kills deals when the gap is wide. The agent reconstructed actual landed cost (base + freight + duty + brokerage + handling) at the SKU-month level across the trailing-12, compared to the standard cost in the item master, and flagged SKUs where actual cost had drifted from standard by more than 5% for two or more consecutive quarters. Out of 312 active SKUs, 41 showed material drift. The aggregate inventory revaluation adjustment came out at $120K — 6% of EBITDA, slightly above our cross-deal median of 4%, and important enough to surface to the buyer's integration team because the standard-cost update process was clearly broken, not just stale.
04 Concentration, quality-of-revenue, and cohort work
The customer-cohort and quality-of-revenue work sits in the middle of the engagement and is where the agent contributed most to analytical depth rather than just speed. The brand had three years of order-line data across DTC and wholesale — roughly 480,000 distinct DTC end customers and 11 active wholesale accounts. The agent built monthly DTC customer cohorts, computed retention curves and contribution margin by acquisition channel and acquisition month, and produced a layered LTV view tied back to gross-margin profile by cohort.
The headline finding was useful: the brand's blended LTV had compressed by roughly 18% over the trailing 24 months, but the compression was concentrated in two specific acquisition cohorts (Meta-acquired customers in late-2024 and early-2025) where CAC inflation had outpaced the retention curve. The 2023 cohorts and the 2025 organic and SEO cohorts held LTV almost intact. That nuance — which a senior associate would have caught with two weeks of cohort modelling — was visible in the agent's output by the end of week two, and the buy-side principal used it directly in the underwriting model rather than waiting for the formal QoE report.
The same discipline applied. We re-ran the cohort math independently in a separate workbook, reconciled to the agent's output line-by-line, and resolved a 2% discrepancy to a definition mismatch on how returns were netted against gross revenue. Both numbers were defensible; the QoE report disclosed both and footnoted the methodology gap. The cross-check was load-bearing.
05 What the AI did not do
Three categories of work in this engagement looked almost identical to the all-human version of a comparable QoE eighteen months earlier. Naming them is the honest part of this writeup.
The management interviews were entirely human. Five sessions with the CFO and two with the head of operations covered revenue-recognition policy, the history of the standard-cost update process, the seller's view on contested trade deductions, the integration assumptions baked into projected synergies, and the soft-signal diligence on key-person and culture-fit risk. The agent surfaced the right questions; the answers came from human-to-human conversation. We tested an experiment where the agent summarised recorded interview transcripts and flagged inconsistencies against the data room — that surfaced one minor issue worth following up on, but it was not a substitute for the conversation.
The materiality and judgement work — what is an add-back, what is a normalisation, where to draw the line on non-recurring versus recurring — was entirely human. The agent produced first-pass schedules of candidate add-backs from the GL with reasoning text attached, but every line was reviewed against the seller's representation, contractual basis, historical pattern, and cross-deal experience. The agent's first-pass schedule was about 70% of what we eventually used; the missing 30% and the corrections to the included 70% were where the QoE judgement happened.
Report drafting and IC-memo work was human-led with agent-drafted starting points. The agent produced first-pass narrative for each section keyed to the workpapers. The QoE lead rewrote substantially every section — partly because the agent's prose tended toward generic consultant-speak, partly because the narrative judgements (what to emphasise, what to caveat, what to surface to the IC) were judgement calls. The drafting time saving was real but smaller than the data-handling saving — probably 20-30% rather than 50-70%.
06 Hallucination discipline and verification protocol
The verification protocol is the thing I would urge anyone running this stack to copy first. Four pieces, each non-negotiable.
- 01 Source traceability is mandatory and machine-enforced. Every factual assertion the agent produces carries a citation back to a specific document, page, and section (or workbook, sheet, and cell). The output schema requires a source_reference field on every fact; the agent is configured to return NOT_FOUND rather than guess when no source meets the relevance threshold. The QoE manager spot-checks 15% of citations against source; the engagement audit trail logs every check.
- 02 Deterministic vs generative tasks are partitioned at the workflow level. Numbers come from deterministic calculations in Excel, Alteryx, or Python against structured data. The agent annotates, summarises, and explains; the agent does not compute the headline numbers in the QoE report. Returns curves, implied reserve calculations, cohort LTV math — all of it lives in a transparent, replicable model that the agent reads from. The LLM is the interface to the analytics, not the analytics itself.
- 03 Dual extraction with reconciliation on high-risk fields. For categories that move the deal most — change-of-control clauses, termination rights, contested accounts, key revenue concentrations — we ran two independent extractions (rules-based pass and LLM pass) and reconciled discrepancies before trusting the field. The dual-pass added two hours of runtime and surfaced four discrepancies on this engagement; one was a genuine misread that would have biased the customer-concentration narrative.
- 04 Human sign-off on every section, partner sign-off on the report. No section leaves the engagement without manager-level review of underlying workpapers and citations. The partner signs the report knowing what the agent did, what the human did, and where each finding sits on the deterministic-vs-generative spectrum. The accountability sits where it has always sat.
The protocol is a direct adaptation of ABA Formal Opinion 512 and the California State Bar GenAI guidance — independent verification of any AI output that affects the work product, supervisory responsibility under Model Rules 5.1 and 5.3, confidentiality discipline under Model Rule 1.6. The same pattern applies to QoE work for the same reason: the professional accountability sits with the human signing the report, regardless of which tools produced the underlying draft. Treating the agent as a research analyst with imperfect output, not as a substitute decision-maker, is the operating model that lets the rest scale.
The agent surfaces; the human decides. The protocol is what keeps the surfacing useful instead of corrosive.
07 What this changed for the engagement economics
The engagement-level numbers matter because the vendor narrative overstates the saving in one direction and the AI-skeptic narrative dismisses it in the other. Total QoE hours ran 38% lower than the comparable mid-2025 deal — 410 versus 660. The saving was concentrated in three places: the first 72 hours of data-room triage (roughly 60 hours saved), contract metadata extraction (roughly 50 hours saved), and reserve-adequacy first-pass (roughly 80 hours saved). The back half of the engagement — management interviews, judgement-heavy adjustment work, report drafting, IC presentation — was within 10% of baseline. Partner time was unchanged; manager time dropped roughly 20%; analyst-and-senior-associate hours absorbed most of the saving.
The cycle-time saving was bigger than the hours saving. Calendar time from data-room open to draft QoE delivery ran 4.5 weeks versus the 7-week baseline. Compression came from parallelism — we ran reserve-adequacy, contract analytics, and cohort analysis simultaneously in week two rather than sequentially across weeks two through four. The buy-side principal had preliminary findings in time to use them in the LOI negotiation, which had not been possible on the prior deal. McKinsey's 30-50% cycle-time finding for high-adopters lines up. The deliverable quality was at least as good as the all-human baseline and arguably better on the cohort and concentration analytics where the agent enabled depth we would not otherwise have produced inside the timeline. Our pricing held; the deliverable got better.
08 Where this goes next
Three things we are tightening. First, the management-interview transcription-and-cross-reference experiment is worth running properly — the agent reading transcripts against the data room could surface inconsistencies in close to real time, but we need to harden audit trail and confidentiality first. Second, contract metadata extraction needs to move from 94% to closer to 98% accuracy before we reduce the spot-check rate below 15%; better prompt templates and tighter dual-extraction on high-impact fields should get us there. Third, report-drafting is currently the weakest link — the right next step is a small library of in-house section templates fine-tuned against our prior QoE reports.
Companion reads: what a QoE actually finds in DTC and CPG at /blog/buy-side-qoe-dtc-cpg-what-it-finds/; the sell-side QoE-readiness checklist at /blog/qoe-checklist-template-dtc-cpg-acquisitions/; the CFO-facing AI evaluation framework at /blog/cfo-ai-evaluation-framework-2026/; the Q2 2026 DTC/CPG M&A multiples read at /blog/dtc-cpg-mid-market-ma-multiples-q2-2026/.
Frequently asked questions
Can AI run a QoE on its own in 2026?
What does AI actually do in the first 48 hours of a buy-side QoE?
How accurate is AI contract metadata extraction on M&A documents?
Can AI test reserve adequacy on returns, trade deductions, and inventory?
What is the hallucination risk on AI in M&A diligence and how do you control it?
How much faster is an AI-enabled QoE versus an all-human one?
Does this work for sell-side QoE-prep too?
Engagement: Q1 2026 buy-side QoE on a $65M-revenue DTC-and-wholesale consumer brand, conducted by Putra & Co transaction services. Comparison baseline: a Q3 2025 buy-side QoE on a similarly-sized brand under an all-human workflow. Hour and cycle-time figures are engagement-level actuals, not averages.
AI tooling: VDR-integrated ingestion agent on a firm-controlled, non-self-learning model environment with logging, access controls, and audit trail. RAG architecture with mandatory source citation, NOT_FOUND fallback, and dual-extraction reconciliation on high-impact fields. Deterministic analytics run in transparent Excel/Python models the LLM reads from, not computes.
Verification protocol drawn from ABA Formal Opinion 512 (2024) and the California State Bar Generative AI Practical Guidance, adapted for QoE. Source-traceability, deterministic-vs-generative partition, dual extraction, and human sign-off are non-negotiable engagement standards.
Cross-deal medians (returns 14%, trade-deduction 7%, inventory standard-cost 4% of EBITDA) from the Putra & Co sample of 18 consumer-brand QoEs, 2022-2025. Methodology at /blog/buy-side-qoe-dtc-cpg-what-it-finds/.
Full source list at content-pipeline/research/ai-ma-diligence-actual-workflow-qoe/sources.md. Primary references include Datasite, Intralinks, Imprima, Ansarada, McKinsey Gen AI in M&A, IBA Impact of AI on M&A, and ABA Formal Opinion 512.