Insights / Practice / Practice
Field note

The CFO's AI evaluation framework: buy, build, skip in 2026.

Half the agentic AI pitches arriving on the CFO's desk are real, half are dashboards with a new label, and the cost of getting the call wrong is six to twelve months of distraction. The three-question framework I use to triage every pitch into buy, build, or skip.

I have sat in a meaningful number of AI-for-finance pitches in the past eighteen months, and the gap between the pitch and the deliverable is wider than in any technology wave I have lived through. Gartner now forecasts that 60% of AI projects will be abandoned by end-2026 specifically because the data is not AI-ready, and that more than 40% of agentic AI projects will be cancelled by 2027. MIT's 2025 NANDA study of 300+ deployments found 95% of generative AI pilots fail to produce measurable financial impact. BCG's 2025 finance ROI survey reports a third of AI-using finance orgs running below 5% ROI, with only one in five clearing the 20% threshold most CFOs underwrote to. The pattern across the failures is the same three mistakes: the use case is misclassified, the data is not ready, and the failure cost is not recoverable inside the workflow. The framework I use to triage every pitch is three questions, asked in order, and it has held up across our practice.

01 Three questions, in order

The framework is three questions, and the order matters. One: is the use case a real-time recurring workflow, a periodic batch task, or a one-time analysis — and is the vendor pitching the right deployment model for that classification? Two: is the data the agent needs already structured, governed, and trustworthy enough to make autonomous decisions on, or does the agent need to be a co-pilot to a human owner because the underlying data is not yet AI-ready? Three: what is the failure cost if the agent gets it wrong, and is that failure cost recoverable inside the natural cadence of the workflow? The three answers together determine buy, build, or skip. Each question on its own will mislead you. The order traps the noise.

KPMG's 2025 paper on agentic finance describes a similar discipline under the label "TACO" — Task, Actionability, Control, Operational context — and reports that 67% of CFOs they surveyed prefer to buy agentic capability from a vendor rather than build it. Gartner's tiering of agentic use cases into safe-fail, material-but-recoverable, and direct-P&L-or-regulatory categories tracks the same logic. The three-question version I use is the one I have found mid-market operators actually apply without an offsite or a consultant. It is the same logic stripped to its operating skeleton.

Half the AI pitches that arrive on the CFO's desk are real. The framework is the discipline that separates the half that matters from the half that does not — and the order of the questions is where the discipline lives.
— From a CFO peer roundtable, March 2026

02 Recurring, batch, or one-time — classify before you value

Recurring real-time workflows — invoice categorisation against the chart of accounts, exception flagging in the close, AR collections prioritisation, deduction matching against the trade-promotion ledger, cash forecasting — are the workflows agentic AI can compound on. Each iteration trains the system; each iteration earns the cost of the integration; each iteration moves a measurable percentage of the work off a human queue. Vic.ai's published case studies show 70-80% of invoices fully automated after ramp-up; HighRadius and Tesorio report DSO reductions of 5-10 days on continuous AR-prioritisation deployments; AppZen runs 90%+ of expense and invoice audits autonomously after rules tune. None of this is glamorous. All of it compounds inside the first three quarters when the data is ready.

Periodic batch tasks — quarterly board pack assembly, monthly variance commentary, year-end audit prep — are AI candidates, but the ROI is harder to defend because the volume per cycle is lower and the per-cycle cost of getting it wrong is higher. Most of the FP&A vendors in the market today — Pigment, Cube, Mosaic, Datarails — sit in this band, and the honest reading of their case studies is that they deliver real value but on an 18-36 month payback, not 12-18, and the ROI depends heavily on whether the freed-up FP&A hours actually change the decisions the company makes. The agentic flavour in this band is mostly narrative generation, suggestive forecasting, and exception alerting — useful, but augmentation rather than autonomy.

One-time analyses — a strategic-options review, an M&A target screen, a one-off process design — are AI co-pilot tasks, not AI automation tasks. There is no integration to amortise, no training loop to compound on, no second iteration to earn back the setup cost. Anyone pitching a custom-built agent for a one-time analysis is selling you the wrong deployment model for the workload. Use a co-pilot. Pay by the seat. Move on. The framework places every candidate workload into one of these three buckets first, and a meaningful share of inbound pitches die at this question alone because the vendor is pitching automation infrastructure for what is structurally a co-pilot task.

03 Data readiness as the gating factor

The single most common reason an AI-for-finance project fails inside a $20M-$200M operating company is that the data is not ready for the agent to act on. Gartner reports that 63% of organisations do not have or are unsure they have the right data management practices for AI. Analytics8's 2024 mid-market study found that only 14% of mid-market organisations consider themselves fully data-ready, and 85% cite inconsistent data architecture, poor data hygiene, and siloed systems as the top three impediments. Tech-Azur's mid-market read is harsher: 71% of organisations are explicitly assessed as unready, and only 25% of the 91% running some form of generative AI have it integrated into core operations. The other 75% are in what they call "pilot purgatory." That is the realistic peer-group baseline. If you have not done the data foundation work, you are in the unready majority by default.

In our practice the specific failure modes are unglamorous and consistent. The chart of accounts has not been rationalised — most $20M-$200M companies are running 400-1,200 active GL accounts inherited from acquisitions and historical accident, with the same expense booked under inconsistent codes across subsidiaries. The vendor master has duplicates, missing or incorrect tax IDs, and non-mailable addresses, with the same legal entity represented under three or four different vendor records across ERP, AP, and procurement systems. The SKU master has competing primary keys across the ERP, the e-commerce platform, the warehouse system, and the 3PL extract — so the agent cannot reliably compute SKU-level margin without the human owner reconciling first. The sub-ledger does not tie to the GL on the AR or inventory side, with aged unreconciled items sitting in suspense accounts that no one has touched in eight months. Buying or building agentic AI on top of any of this produces high-confidence wrong answers, which is the worst category of output the finance function can produce. The discipline is to assess data readiness honestly before any AI capital is committed.

60%
Share of AI projects Gartner forecasts will be abandoned by end of 2026 specifically due to lack of AI-ready data (Gartner, 2023-2024).
14%
Share of mid-market organisations reporting full data readiness for AI in 2024 (Analytics8 mid-market study).
63%
Share of inbound AI-for-finance pitches we triage to "data readiness first" rather than "buy now," across 2025-2026 in our practice.

The realistic timeline to "AI-ready finance data" tracks the complexity of the operating estate. A relatively simple $20M-$100M company running one main cloud ERP — NetSuite, Sage Intacct, Microsoft Dynamics — and a basic data warehouse can typically get to a defensible foundation for a first wave of agentic deployments in three to six months of focused work. A more complex $100M-$200M company with multiple ERPs across acquired entities, multiple product systems, and unresolved master data should plan on nine to twelve months for a robust foundation and twelve to eighteen months for a leader-level state. CoA standardisation runs eight to twelve weeks of design and mapping plus three to six months of implementation. Vendor master cleanup runs four to eight weeks of intensive work on a 5,000-20,000 vendor file, then settles into a standing monthly process. SKU master harmonisation runs eight to sixteen weeks for a single business unit and six to twelve months across multiple systems. The work can be overlapped, which is how the totals compress, but it cannot be skipped without conceding the agent to the high-confidence-wrong-answer category.

04 Failure cost and recoverability inside the cadence

The third question is the one most often skipped and the one that produces the worst outcomes when it is skipped. What is the cost if the agent gets it wrong, and is the cost recoverable inside the natural cadence of the workflow? An invoice categorised to the wrong GL line — recoverable inside the month-end close during variance review, low cost, low signal. An AR-collections agent that mis-prioritises a customer by one day — recoverable inside the weekly cash review, no external commitment made. A cash forecast that mis-reads a seasonal spike — recoverable if the CFO still walks the model weekly, not recoverable if the forecast is acted on without review. These are the tier-one workflows where agentic deployment compounds with low downside.

The not-recoverable workflows are different in kind. A vendor payment authorised on a fraudulent invoice — not recoverable once cash leaves the account; the agent should not have payment-release authority without a human checkpoint, full stop. A tax filing submitted on incorrect jurisdictional rates — not recoverable inside the close cycle; the error sits as latent exposure for years until a state audit surfaces it, by which time interest and penalties have compounded. A customer-facing pricing decision or discount approval — binding under contract law in most jurisdictions; the Air Canada chatbot case (Moffatt v. Air Canada, 2024 BCCRT 149) settled this question publicly when the British Columbia Civil Resolution Tribunal held the airline liable for a refund policy its chatbot had invented and ordered it to honour the policy. The chatbot was removed from the website the following month. An external financial report — a bank covenant package, an investor update, statutory accounts — where the lender or shareholder relies on the figure: not recoverable from a reputation standpoint even if the math can be restated, because the signal of weak control is the damage.

The compounding-error math underneath this is what makes the recoverability question non-negotiable for multi-step workflows. Prodigal's reliability analysis is the cleanest framing: at 95% reliability per step over a 20-step process, the end-to-end success rate is 36%. At 99% per step over 20 steps, end-to-end is 82%. The OSWorld benchmark on multi-step GUI agents shows real-world end-to-end success rates as low as 14.9% for Claude 3.5 Sonnet and 22.7% for UI-TARS on long-chain computer workflows. If the workflow chain includes a not-recoverable step, the compounding-error math says the chain will fail catastrophically with predictable frequency. Agents with autonomy over not-recoverable steps are not a deployment strategy; they are a tail-risk position the CFO is unwittingly underwriting.

The decision rule

With the three answers in hand the decision is mechanical. Recurring + data-ready + failure-recoverable equals buy (or build only if the workflow is differentiated enough to justify the engineering cost — almost never the case at $20M-$200M scale). Recurring + data-not-ready equals a data project first, AI later — and the data project is sequenced honestly with a multi-quarter timeline, not disguised as an AI initiative. Recurring + failure-not-recoverable equals co-pilot mode only with a human checkpoint on every action that crosses the not-recoverable threshold; no autonomous execution authority. Batch or one-time equals co-pilot mode by default; do not invest in automation infrastructure for what is structurally co-pilot work. Anything else equals skip, which is most of what walks through the door.

05 Buy, build, or skip — what survives in mid-market

Run the three questions across the AI-for-finance vendor landscape and a clean pattern emerges. The AP-automation category — Tipalti, Vic.ai, AppZen, Stampli on the autonomous end; Ramp and Brex on the card-plus-policy end — is the most consistent buy at mid-market scale. The workflow is recurring and high-volume. The data is closest to ready because the invoice and GL master are the most-touched objects in finance. The failure cost is recoverable inside the close. Independent ROI work from HighRadius, Tipalti, Yooz, MadrasAccountancy, and KlearStack converges on a 70-85% reduction in per-invoice cost and a 6-18 month payback at mid-market volume. The Tipalti case study showing 195% year-one ROI and a ten-month payback is the cleanest published example. Buy this category if you process more than 10,000 invoices a year on manual or semi-manual workflows. Skip if you are below 3,000-4,000 invoices a year and already on a modern card platform.

The AR and collections category — HighRadius, Tesorio — is the second consistent buy when the AR balance is material. The agent monitors AR continuously, prioritises customers by risk and likelihood of paying, triggers outreach, and updates the cash forecast as new data arrives. The work is recurring; the data is in the order-to-cash sub-ledger and is usually cleaner than the vendor master; the failure cost on a mis-prioritised customer is one day of delay, fully recoverable in the weekly cash review. Vendor case studies regularly show DSO reductions of five to ten days, which on a $50M revenue base with a 50-day DSO unlocks several million in working capital. Buy if AR sits above $5-10M outstanding and DSO is above 45 days. Skip if AR is small or DSO is already below 30 days.

The close-automation category — BlackLine, FloQast, Numeric — passes the three questions but the ROI window is longer. The workflow is recurring; the data dependency on a clean CoA and reconciled sub-ledgers is higher; the failure cost is contained inside the close. Independent and vendor numbers converge on two to five days off the month-end close and a 20-40% reduction in manual reconciliation hours, with payback at twelve to twenty-four months conditional on the close being genuinely painful today. Buy if you are closing in more than ten business days, your controllership team is four or more people, and your audit-adjustment trail shows recurring control weaknesses. Skip if your close is already under five business days and your reconciliation discipline is intact — the marginal AI dollar is better spent in AP or AR.

The FP&A and planning category — Pigment, Cube, Mosaic, Anaplan, Planful, Datarails — is where the three questions get more interesting and the answer is genuinely company-specific. The workflow is mostly recurring but the per-cycle volume is lower; the agentic features are mostly narrative generation, suggestive forecasting, and exception alerting rather than autonomous decision-making. The honest ROI window for mid-market is eighteen to thirty-six months and the value realisation depends on whether the freed-up FP&A hours actually change pricing, hiring, and capital allocation decisions. BCG's research suggests organisations at higher analytics maturity spend 40% less time on variance reporting and 60% more on strategic planning — but that is a maturity-driven outcome, not a tool-driven one. Buy if Excel is visibly breaking under multi-entity, multi-scenario modelling. Hybrid-build is genuinely viable at the lower end of the band ($20M-$50M) with strong internal finance-engineering talent. Skip if analytics maturity is low and ERP data quality has not been fixed first.

The tax-compliance category — Avalara as the dominant mid-market name — is genuinely operational rather than dashboard-rebranded but the AI labelling is mostly incremental on top of automation that has existed for a decade. Buy if multi-state or multi-country sales-tax exposure is non-trivial. Skip if your footprint is simple and a local CPA can handle the filings. Build is economically irrational at any mid-market scale.

Below those categories, the dashboard-rebrand band is wide and the discipline of the three questions is the protection. A meaningful share of the FP&A and "AI CFO co-pilot" pitches in the market today are LLM narrative layers over existing analytics workflows with no recurring autonomous action. They are not bad products. They are not agentic deployments either, and they should be evaluated as co-pilot seat licences, not automation infrastructure investments. The deployment model determines the procurement model. The framework forces that distinction at the front of the conversation.

06 How to run the PoC without being sold a dashboard

Once a candidate vendor passes the three questions, the PoC is the test. The PoC discipline that works at mid-market scale has five components, and all five are non-negotiable. Pick one workflow with a binary outcome — "should this invoice be auto-approved and coded correctly: yes or no?" — not a vague "see what the AI can do" demo. Run a thirty-, sixty-, or ninety-day pilot on your data, not the vendor's reference data — one thousand to five thousand historical invoices or three to six months of close cycles is the minimum sample. Run shadow mode first; the agent makes a recommendation, the human acts, and the comparison is logged. Never put the agent into the production decision path during the PoC. Measure work removed, not just accuracy — the percentage of invoices touched zero times, the percentage of reconciliations auto-certified, the percentage of variance commentaries that pass review without rewrite. And run an exit drill at the end of the PoC — confirm you can export the configurations, the prompts, and the evaluation set and stand up an alternative or revert in a week. The exit drill is the lock-in test.

The TCO line that most CFOs underestimate is integration and config hours. Vendor pricing is the visible number; the engineering and finance hours required to ingest the data, tune the rules, and validate the outputs are the invisible number. Track them from day one. The vendors that survive the PoC test in our practice are the ones whose tuning time per workflow lands below 80 hours of finance-team effort. The ones that don't survive are the ones that require a full implementation partner and a six-month engagement to deliver the first dashboard.

07 Three questions for the next AI pitch

Carry these into the next vendor meeting. The order is the discipline.

  1. Is this a recurring real-time workflow, a periodic batch task, or a one-time analysis — and is the vendor pitching the right deployment model for that classification? If they are selling automation infrastructure for what is structurally co-pilot work, the deployment model is wrong before the demo starts.
  2. Is the data this agent needs cleaned up, governed, and trustworthy enough today — and if not, how long is the data project before the AI deployment can deliver? If the answer is "the agent will work around the data issues" the answer is wrong; agentic AI on dirty data is a confidence-multiplier on wrong answers.
  3. If the agent gets it wrong inside this workflow, is the failure recoverable inside the cadence of the workflow, or does it compound externally? If the failure crosses into payments, tax filings, customer-facing pricing, or external financial reporting, the agent should not have autonomous execution authority — co-pilot mode only, with a human checkpoint on every binding action.
The framework is three questions. Most pitches die at question one because the deployment model is wrong for the workload. Most of the rest die at question two because the data is not ready. The few that survive question three are worth the integration cost.
— From a working session on the 2026 AI-for-finance roadmap, April 2026

08 Sequencing the work — what to do this quarter

For most $20M-$200M operating companies the practical sequence in 2026 is unglamorous and consistent. Open a data-readiness assessment first — chart of accounts, vendor master, SKU master, sub-ledger reconciliation — and sequence the cleanup with a multi-quarter timeline. Pilot one Tier-1 agentic workflow on the cleanest data domain you have — typically AP categorisation or AR-collections prioritisation, sometimes close auto-certification — using the five-component PoC discipline. Treat every higher-risk workflow as co-pilot only with human checkpoints until the data foundation and the Tier-1 deployments have compounded for two to three quarters. Refuse the vendor pitches that are selling automation infrastructure for what is structurally a co-pilot task, and refuse the pitches that promise to "work around" the data readiness gap. The discipline is boring on the way in and decisive on the way out. The companies that will compound an AI-for-finance advantage in 2027 and 2028 are the ones that did the data work in 2026 and bought sparingly against the framework. The ones still in pilot purgatory at year-end 2026 will be the ones that skipped the order of the three questions.

For deeper reads on the sector-specific versions of this framework, the companion posts on the AI agent stack for a $50M CPG brand, on AI agents in inventory, supply chain and finance, on agentic AI in the financial close at $20M-$100M, and on AI in M&A diligence and quality-of-earnings work walk through what the three questions look like inside specific workflows. The framework is the same in every sector; the workflow inventory and the data readiness picture is what changes.

Frequently asked questions

What is the CFO AI evaluation framework for buy, build, or skip in 2026?
Three questions in order. First, classify the use case as recurring real-time workflow, periodic batch task, or one-time analysis, and check the vendor pitches the right deployment model. Second, assess data readiness — is the CoA, vendor master, SKU master, and sub-ledger reconciliation in shape for an agent to act on? Third, is the failure cost recoverable inside the workflow cadence? Recurring + data-ready + recoverable equals buy. Failing any one is skip, co-pilot only, or data project first.
Why do most AI-for-finance projects fail at mid-market scale?
MIT NANDA's 2025 study of 300+ deployments found 95% of generative AI pilots fail to produce measurable financial impact. Gartner forecasts 60% of AI projects abandoned by end-2026 due to data not being AI-ready, and 40%+ of agentic AI projects cancelled by 2027. In our practice the failures cluster into three patterns: misclassified use case, unready data foundation (CoA, vendor master, SKU master, sub-ledger reconciliation), and failure cost not recoverable inside the workflow.
How long does it take to get mid-market finance data AI-ready?
Simple $20M-$100M with one cloud ERP and a basic data warehouse: three to six months focused work. Complex $100M-$200M with multiple ERPs, entities, and product systems: nine to twelve months for a robust foundation, twelve to eighteen for leader-level. CoA standardisation runs 8-12 weeks design plus 3-6 months implementation. Vendor master cleanup runs 4-8 weeks on a 5,000-20,000 vendor file. SKU master harmonisation runs 8-16 weeks per business unit. Work overlaps, but it cannot be skipped.
Which AI-for-finance vendor categories deliver ROI inside 12-18 months at mid-market?
Three consistent buys. AP automation (Tipalti, Vic.ai, AppZen, Stampli) reports 70-85% per-invoice cost reduction and 6-18 month payback — Tipalti's Global Tech case shows 195% year-one ROI, ten-month payback. AR/collections (HighRadius, Tesorio) regularly shows DSO reductions of 5-10 days. Close automation (BlackLine, FloQast, Numeric) pays back over 12-24 months if the close is genuinely painful. FP&A platforms (Pigment, Cube, Mosaic, Anaplan, Planful, Datarails) typically realise value over 18-36 months.
How is an agentic AI deployment different from a dashboard rebrand?
An agentic deployment runs in a recurring workflow and takes autonomous action against measurable work — invoices categorised, reconciliations auto-certified, AR customers prioritised. The measure is work removed from a human queue per period. A dashboard rebrand is an LLM narrative layer over existing analytics with no autonomous action — the human still does the work. The PoC discipline that separates the two is measuring work removed (% invoices touched zero times, % reconciliations auto-certified) rather than accuracy or feature breadth on a vendor demo.
When should an AI agent NOT have autonomous execution authority?
Any workflow where the failure cost is not recoverable inside the natural cadence. Vendor payment authorisation. Tax filings and regulatory submissions. Customer-facing pricing or contract terms — the Air Canada chatbot case (Moffatt v. Air Canada, 2024 BCCRT 149) held the airline liable for a refund policy its chatbot invented. External financial reporting where a lender or shareholder relies on the figure — bank covenants, investor updates, statutory accounts. In each, the agent runs in co-pilot mode only with a human checkpoint on every binding action.
How do I run an AI vendor PoC without being sold a dashboard?
Five components. Pick one workflow with a binary outcome ("should this invoice be auto-approved and coded correctly: yes or no?"). Run 30-90 days on your data, not vendor reference data — 1,000-5,000 transactions or 3-6 close cycles. Shadow mode first; never put the agent in the production path during PoC. Measure work removed, not accuracy. Run an exit drill — confirm you can export configurations, prompts, and evaluation sets and revert in a week. Track integration and config hours; vendors needing more than 80 hours of finance-team tuning per workflow rarely survive.
Notes

Framework version 2026.2, applied across 14 advisory engagements 2025-2026. Three-question structure refined against Gartner's tiering (safe-fail / material-but-recoverable / direct-P&L-or-regulatory), KPMG's TACO framework (Task, Actionability, Control, Operational context), and BCG's ROI-threshold work.

Adoption and failure statistics: Gartner 2024 AI Hype Cycle and 2023-2024 forecasts (60% AI project abandonment by end-2026; 40%+ agentic AI cancellation by 2027); MIT NANDA study 2025 (95% of generative AI pilots fail to produce measurable financial impact, 300+ deployments); BCG 2025 finance AI ROI survey (one-third below 5% ROI; 1 in 5 above 20%); KPMG AI Pulse Survey 2024 (78% confidence; 67% prefer buy over build).

Data readiness research: Analytics8 Mid-Market Data Readiness Study 2024 (14% fully ready; 85% cite hygiene/architecture/silos); Huble AI Data Readiness Report 2024 (8.6% fully AI-ready); Tech-Azur 2024 (71% unready, 75% in pilot purgatory). Vendor master and CoA cleanup timelines from NetSuite and Sage Intacct partner guidance 2023-2024, GEP and TealBook vendor-master research.

Failure-cost case law: Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, 2024) — airline held liable for refund policy invented by AI chatbot; chatbot removed from website following ruling. Compounding-error math: Prodigal, "Why most AI agents fail in production," 2024; OSWorld benchmark on multi-step GUI agents 2024.

Full source list at content-pipeline/research/cfo-ai-evaluation-framework-2026/sources.md in the Putra & Co content pipeline.

About the author
Marcin Samiec
Partner · Practice

Marcin Samiec

Senior Partner, Tech & AI

Tech executive with 15+ years in transformation and IT strategy, now operating as a fractional CIO. Leads systems modernization, ERP and platform implementations, project rescue and turnaround, and the privacy-and-security build (GDPR, ISO 27001) inside operating businesses. Brings the AI and agentic-operations practice to where the operating stack actually runs — oil & gas, construction, real estate and professional services.