A finance-grade calculator for deciding whether one workflow should receive agentic investment
Altivate | Published 21 June 2026 | Updated 26 July 2026
1. Start with the queue the business already owns
Most AI agent cases begin in the wrong unit. They start with headcount, model capability or the number of tasks an agent might perform. Finance cannot govern any of those as a benefit.
Start with one queue:
- supplier invoices waiting for resolution;
- service cases waiting for evidence and routing;
- orders waiting for an exception decision;
- maintenance work orders waiting for triage;
- contracts waiting for first review;
- account records waiting for validation.
A queue has observable economics. Work arrives at a rate. Some cases follow a standard path, some need judgment and some must never leave accountable human control. Cases wait, age, get reopened, escalate and occasionally cause loss. People spend time on them, but time is only one part of the cost.
This is the first discipline of the investment case:
The unit should be narrow enough that a process owner can answer four questions from evidence:
- How many cases arrive?
- What happens to each case today?
- What outcome tells us the case is complete and correct?
- What is the consequence when the outcome is late or wrong?
Do not write “finance agent” or “customer-service agent.” Write “classify an inbound invoice exception, retrieve the relevant purchase order and policy, draft the proposed resolution, and route it to the authorised approver.” That description exposes the reads, actions, owner and residual human work that the calculator needs.
Use RAG, tools, workflow or agent to select the simplest architecture, then apply the Enterprise AI Agent Readiness Test to each action the system may take.
A positive headline before the final block is not a business case.
2. The calculator has five evidence blocks
The downloadable calculator is a worksheet, not a forecast. Its completeness gate refuses to show NPV or financial approval until each required input is explicit. The tables below define the evidence contract for those inputs; use an observed baseline, a contracted price or a bounded assumption with an owner and review date.
Workload and baseline
| Input | Evidence required | Evidence owner |
|---|---|---|
| Annual cases entering the queue | Full-period queue count with exclusions stated | Process owner |
| Average active handling minutes per case | Time study or system activity duration | Operations |
| Average elapsed cycle time | Open-to-close duration with median and tail cases | Operations |
| Rework or reopen rate | Reworked or reopened cases divided by completed cases | Quality owner |
| Current loss per failed case, where measurable | Verified unit loss, credit, penalty or recovery cost | Finance / Risk |
| Fully loaded hourly cost for affected work | Finance-approved labour and overhead basis | Finance / HR |
Do not multiply all handling time by labour cost and call it savings. First establish what happens to released capacity. If work disappears but headcount, overtime, contractor spend, backlog and service output do not change, no financial benefit has been realised.
Addressable scope
| Input | Required boundary | Evidence |
|---|---|---|
| Share of cases in the approved process scope | Enumerated inclusion and exclusion rules | Sampled queue classified against the rules |
| Share with sufficient data and tool access | Required sources, permissions and actions | Evaluation cases that pass and fail for missing context |
| Expected user adoption | Named user groups and changed behaviours | Observed pilot use by cohort |
| Share allowed to complete without per-case approval | Reversible actions and retained human consequences | Approved authority matrix and action-level tests |
Eligible cases
annual cases x in-scope share x data-ready share x adoption
The product of these inputs is often much smaller than the original volume. That is useful. It prevents a business case from taking credit for work the first release cannot reach.
Agent performance and residual work
| Input | How it is measured |
|---|---|
| Correct accepted rate | Outcome grader against final system state |
| Exception or escalation rate | Production-like evaluation set |
| Accepted-error rate | Downstream quality measure |
| Human review minutes on all accepted cases | Timed pilot observation |
| Human handling minutes on exceptions | Timed pilot observation |
| Rework minutes after an accepted error | Timed downstream observation |
Measure the outcome, not the agent’s claim. Anthropic’s evaluation guidance makes this distinction directly: a transcript may say that a booking was made, but the outcome is whether the reservation exists in the system. That is vendor guidance, and it is also sound measurement practice.
The correct-accepted, exception and accepted-error rates must be mutually exclusive and total 100% of eligible cases. This prevents an exception rate from being deducted once in volume and again inside a per-case time formula.
Benefit economics
Keep benefit classes separate:
| Benefit class | Calculation | Finance treatment |
|---|---|---|
| Released capacity | Baseline workload minus future workload, converted at loaded rate | Count only the share that changes an accepted cost or output |
| Avoided cost | Future contractor, overtime or hiring cost no longer required | Tie to an approved spending baseline |
| Quality | Avoided rework, credit, penalty or failure cost | Use observed incidence and accepted unit loss |
| Cycle time | Working-capital, service or throughput effect | Name the causal mechanism |
| Revenue | Additional completed or retained business | Apply probability, margin and attribution |
| Risk | Reduced expected loss | Probability x impact, reviewed by Risk and Finance |
Baseline annual workload
eligible cases x baseline handling minutes
Future annual workload
correct accepted cases x review minutes + exception cases x exception handling minutes + accepted-error cases x (review minutes + rework minutes)
Released hours
(baseline annual workload - future annual workload) / 60
If future workload is equal to or greater than baseline workload, released hours and capacity value are zero or negative. Do not hide that result.
Capacity value before realisation
released hours x loaded hourly cost
Recognised capacity benefit
capacity value before realisation x finance-approved realisation factor
The realisation factor is not a technology input. It expresses what the organisation will do with the capacity. It can be zero. A zero is more credible than pretending all time saved becomes cash.
Full cost and failure exposure
| Cost | Initial evidence | Annual-run evidence |
|---|---|---|
| Process discovery and redesign | Scoped analysis and redesign effort | Ongoing process ownership and improvement |
| Data preparation and access controls | Source remediation and permission setup | Data-quality controls and access reviews |
| Integration and tool execution layer | Build, test and cutover effort | Support, change and transaction cost |
| Model, platform and consumption | Setup, evaluation and environment cost | Contracted or metered usage |
| Evaluation and red-team work | Initial representative set and adversarial tests | Regression evaluation after material change |
| Identity, audit and security operations | Identity, policy and audit implementation | Monitoring, review and incident operations |
| Human oversight and exception handling | Operating-model and queue setup | Observed review and exception workload |
| Change, training and adoption | Role, process and training change | Reinforcement, onboarding and adoption support |
| Monitoring, incident response and improvement | Telemetry and response design | On-call, investigation and improvement work |
| Vendor management and internal product ownership | Selection, contracting and launch governance | Product ownership, assurance and supplier review |
Expected annual failure loss
annual agent actions x material failure probability x average loss per material failure
This is not permission to hide severe tail risk inside an average. A low-probability action with a catastrophic consequence needs a hard authority boundary, not a favourable expected-value cell.
3. Turn the worksheet into three scenarios
Do not debate one forecast. Compare a downside, base and upside case using the same formulas.
| Assumption | Downside evidence | Base evidence | Upside evidence |
|---|---|---|---|
| In-scope and data-ready share | Lower confidence bound from evaluated cases | Observed pilot share | Verified expansion with access already approved |
| Adoption | Lagging user cohort | Observed current adoption | Tested improvement after funded change actions |
| Correct completion | Acceptance floor | Multi-trial observed rate | Upper confidence case, not the best run |
| Exception rate | Adverse observed rate | Production-like average | Lower rate demonstrated across representative cases |
| Residual review minutes | High observed duration | Timed pilot median | Improvement proven in the tested workflow |
| Recognised capacity factor | Zero or Finance floor | Finance-approved capacity plan | Higher factor backed by a committed operating change |
| Run and control cost | High credible usage and oversight case | Contracted plus internal operating cost | Lower case supported by price and usage evidence |
| Expected failure loss | Adverse tested incidence and loss | Reviewed expected-loss case | Lower case only after controls are validated |
Then calculate:
Annual net benefit
recognised benefits - annual run cost - annual control cost - expected failure loss
Simple payback in months
initial investment / positive monthly net benefit
Payback is not defined when monthly net benefit is zero or negative. Report the result as “no payback in the modelled horizon,” not as an error or an infinite-looking number.
Net present value
sum of discounted annual net benefits - initial investment
Use the organisation’s approved discount rate and planning horizon. Do not extend the horizon simply to make the case positive. Model refresh, platform change, re-evaluation and process drift as real operating costs.
The downside case should be uncomfortable but plausible. It should assume slower adoption, more review, lower completion, higher exceptions and a longer stabilisation period. If the investment only works when every uncertain input moves in the favourable direction, it is not ready for approval.
Worked scenario
Assume 20,000 eligible cases at 18 baseline minutes and a loaded rate of AED 180 per hour.
| Scenario | Correct accepted / exception / accepted error | Review / exception / rework minutes | Released hours | Capacity value before realisation |
|---|---|---|---|---|
| Downside | 55% / 35% / 10% | 5 / 20 / 24 | 1,783 | AED 321,000 |
| Base | 70% / 25% / 5% | 4 / 16 / 20 | 3,333 | AED 600,000 |
| Upside | 82% / 15% / 3% | 3 / 12 / 16 | 4,390 | AED 790,200 |
These are arithmetic examples, not Altivate benchmarks. Apply the finance-approved realisation factor, annual run and control cost, and expected failure loss before calculating net benefit.
Download the polished AI agent business-case workbook with an executive scorecard and scenario charts, or use the open CSV model.
4. The business case is an operating contract
A good investment case assigns each benefit and cost to someone who can cause or verify it.
| Commitment | Accountable owner | Evidence at review |
|---|---|---|
| Queue volume and baseline | Process owner | Extract and sampling method |
| Model and integration cost | Product / Technology | Contract and consumption model |
| Adoption | Business change owner | Active use and process compliance |
| Correct outcomes | Product and process owner | Evaluation and production sample |
| Capacity conversion | Operating executive | Budget, backlog or output change |
| Financial recognition | Finance | Accepted benefit ledger |
| Authority and loss controls | Risk / Security / Process owner | Control test and incident record |
This table prevents the technology team from owning a benefit it cannot realise. Product teams can deliver a capable system. Only the business can change work allocation, retire cost, increase throughput or accept a different risk position.
OpenAI’s current use-case guidance recommends collecting specific work opportunities and comparing impact with effort. Anthropic recommends the simplest sufficient architecture because agentic complexity adds latency and cost. Both are vendor guidance. The finance implication is broader: every layer of complexity needs an incremental value claim and an owner.
If a deterministic workflow creates 80 percent of the expected value at half the run and control cost, the agent does not win because it is more advanced. It loses the business case.
5. Benefits that do not survive scrutiny
“Hours saved”
Ask what changes after the hours are released. Accepted answers include lower overtime, avoided hiring, increased completed volume, reduced backlog or reassigned capacity with a measurable output. “People will do higher-value work” is an intention until that work, owner and measure are named.
“Fewer errors”
Define the error, current incidence, detection point and unit consequence. Also measure new failure classes created by the agent, including plausible but unsupported decisions, unauthorized action and correlated failure at scale.
“Faster cycle time”
Identify the step the agent changes and the outcome affected by time. A two-hour reduction matters only if it changes cash, service, capacity, loss or a binding deadline.
“Better employee experience”
Measure task abandonment, rework, active use, override behaviour and time spent correcting output. A satisfaction survey can supplement operating evidence, not replace it.
“Strategic option value”
Option value can justify a bounded learning investment. It cannot be presented as realised ROI. Record the decision the pilot will enable, the evidence required and the expiry date of the option.
6. Put a stop threshold in the approval
The pilot should be authorised with a pre-agreed stop or redesign threshold.
Example structure:
- stop if material authority violations exceed zero;
- stop if the correct-outcome rate stays below the pre-approved floor after the minimum representative sample;
- redesign if exception handling exceeds the approved per-case workload ceiling;
- redesign if adoption remains below the approved floor after the agreed change period;
- do not scale if downside-case payback exceeds the investment committee’s approved horizon;
- do not expand tools or actions until the current action passes its evaluation and control gates;
- re-open the investment case if model, platform, price, process or policy changes materially.
This protects the organisation from “pilot momentum,” where continued spending becomes the reason to continue spending.
It also improves learning. A stopped pilot can still be a successful investment if it resolves an important uncertainty cheaply and prevents a larger loss.
7. The executive sign-off checklist
Before approval, the sponsor should be able to answer yes to each statement:
- ☐ One workflow and one accountable process owner are named.
- ☐ The current queue, handling, cycle, rework and failure baseline are evidenced.
- ☐ The eligible scope excludes cases the first release cannot safely reach.
- ☐ A simpler workflow or tool-based option was costed.
- ☐ Correct completion is measured from final state, not agent self-report.
- ☐ Residual review and exception work are included.
- ☐ Capacity has an accepted conversion mechanism.
- ☐ Each quality, cycle, revenue or risk benefit has a causal formula.
- ☐ Initial, run, control, change and product-ownership costs are included.
- ☐ Material failure exposure and non-negotiable authority boundaries are recorded.
- ☐ Downside, base and upside scenarios use the same formulas.
- ☐ The pilot has a stop threshold and a dated investment review.
If several boxes remain open, approve discovery, not deployment.
8. A better board question
The weak question is:
The better question is:
That question can produce “not yet,” “use a workflow,” or “pilot one action.” Those are investment decisions, not failures of ambition.
An agent earns scale when its evidence improves the business case while the operating controls remain intact. Until then, fund the next uncertainty, not the largest story.
Continue with How to Start with AI Agents Without Starting with an Agent, or return to White Papers.
Sources and verification status
Checked 26 July 2026. Vendor guidance is identified as such and is not used as an independent ROI benchmark.
- OpenAI, Identifying and scaling AI use cases – vendor guidance on workflow discovery and impact/effort prioritisation.
- OpenAI, A practical guide to building agents – vendor implementation guidance on agent components, guardrails and intervention.
- Anthropic, Building effective agents – vendor guidance on workflows, agents and the cost/latency trade-off of complexity.
- Anthropic, Demystifying evals for AI agents – vendor guidance on tasks, trials, traces, graders and outcomes.
- NIST, AI Agent Standards Initiative – current standards initiative covering security, identity, interoperability and evaluation.

