Start with the decision the pilot must support
Before building the workflow, finish this sentence: “We will expand this pilot if it helps this team complete this work, under these conditions, with this level of effort.”
Consider an illustrative supplier-onboarding workflow. An agent gathers documents, checks required fields, routes an approval, and creates a supplier record. The business needs a usable record with the right evidence attached, ready by the agreed deadline.
A technically successful run could leave a missing approval or duplicate record. A useful evaluation has to detect both, then show what it took to resolve the case.
Agree acceptance criteria, the minimum useful improvement, and conditions that would stop expansion with the people receiving the work, before results arrive.
Fix the population of cases before measuring success
Write admission rules that another person could apply consistently. For supplier onboarding, those might specify the business unit, supplier category, and document types in scope. Log excluded cases and their reasons so the reader can see how much of the wider queue the pilot covers.
Once an eligible case enters, keep it in the evaluation. An access failure, missing document, or manual handoff is part of what happened to that case. Reclassifying it after a difficult run makes the pilot look better by changing the workload being measured.
Keep a unique case identifier across attempts. Record arrival time, relevant complexity factors, and the workflow version used. Maintain separate counts for eligible arrivals, cases admitted under the pilot's allocation rule, and cases actually started by the agent.
NIST's AI Risk Management Framework calls for documented test sets and performance measurement in conditions resembling deployment. For this pilot, that means testing the permissions, record quality, and approval queues the operating team will encounter.
Include difficult cases deliberately in development tests. In the live evaluation, preserve the normal case mix or show the results by case type so deliberately oversampling exceptions does not obscure everyday performance.
Define completion and inspect the evidence
For each case, capture the expected end state and where it can be checked. In the supplier example, that includes the correct supplier identity, a single usable record, required evidence, and an authorised approval.
Anthropic's agent-evaluation guidance distinguishes an agent's recorded actions from the final state of its environment. Apply that distinction when grading: inspect the resulting record and approval evidence alongside the execution history.
Use deterministic checks for exact conditions, such as required fields and record identifiers. Ask a qualified reviewer to judge matters that need business interpretation. If a model helps grade those outputs, calibrate its judgments against practitioner review and investigate disagreements.
Keep these outcome categories separate:
| Case outcome | Evidence to retain | Treatment in the scorecard |
|---|---|---|
| Accepted through the planned process | Required end state and any scheduled approval | Completed; record normal review effort |
| Accepted after correction | Final acceptance plus the changes a person made | Completed with corrective intervention |
| Rejected or abandoned | Reason and resulting business disposition | No accepted completion; retain all effort |
| Still open at the reporting cutoff | Current state, age, and next dependency | Unresolved; retain in the admitted population |
A planned approval is part of the designed workflow. A person rebuilding the output is corrective intervention. Measuring both makes it possible to see whether the workflow is creating a sustainable review task.
Compare work under similar conditions
Collect a baseline with the same completion definition before launch. Measure actual handling and review time; estimates recalled at the end of the pilot are a weak basis for a business case.
A concurrent comparison group can help distinguish the pilot's effect from changes elsewhere. Where practical, allocate eligible cases between the existing process and the pilot before staff select which cases look easy. Random allocation can reduce selection bias; the design still needs enough cases and a suitable operational setting.
When that is impractical, compare similar cases from the existing process and state the differences. Supplier category, missing documents, business unit, and approval complexity can all matter. Show results within those groups before using an overall average.
The UK government's Magenta Book guidance on evaluation explains why attributing an improvement requires a credible view of what would have happened without the intervention. A before-and-after chart alone cannot establish that the AI caused the change.
Record staffing changes, policy updates, and unusual volume during the pilot. If the implementation team is quietly clearing difficult cases, include that support. A launch with daily expert attention tests a different operating arrangement from the one the business may inherit.
Keep unfinished work visible at the cutoff
Close the reporting period without erasing its unfinished cases. A case waiting on an external document may complete later; at the cutoff, its final outcome is still unknown.
Report the number and age of open cases, with reasons for delay. Show completion by an agreed case-age deadline, giving both groups the same time to reach it. Cases that have not yet reached that age should remain visibly pending assessment.
An average calculated only from completed cases describes the completed cases. A growing backlog of difficult work can sit outside that average. Put elapsed time for accepted cases beside the open-case count and ageing profile.
Keep admission cohorts identifiable when later outcomes arrive. Update the original cohort's results so that late completion can be reconciled with earlier reports.
Use a scorecard the expansion meeting can act on
The scorecard should make clear which result improved, who supplied the remaining effort, and how much work the evidence covers. Add baseline and pilot columns to these rows, with the observed case count beside every rate.
| Scorecard row | Definition and decision it supports |
|---|---|
| Scope and coverage | Eligible arrivals, admitted cases, and exclusions by reason. Establish how much of the operation was tested. |
| Accepted completion | Accepted cases divided by admitted cases at a stated cutoff, split by planned review and corrective intervention. Show pending cases alongside it. |
| Deadline performance | Accepted by the agreed deadline among cases old enough to assess. Identify whether the team receives the result in time. |
| Human effort | Handling, ordinary review, correction, and recovery time across all admitted cases. Include specialist delivery support. |
| Quality and consequences | Rejected outputs, reopened cases, and consequential mistakes, with severity. Review serious incidents individually. |
| Operating economics | Total operating cost, accepted completions, and unresolved workload. Show setup spending separately. |
Measure human effort in both groups using the same method. Sample ordinary work as well as escalations, and distinguish active handling from time waiting in a queue.
Include platform, model, tool, and runtime charges alongside labour. Avoid counting a charge twice when it is included in a platform fee. The cost per completed workflow guide explains how to allocate failed attempts and recovery work.
Released hours create capacity. A cash saving requires an actual change in expenditure; additional throughput requires evidence that the team handled more work at the required quality. Choose the benefit the business intends to realise and measure it directly.
Make a bounded expansion decision
Expand when the observed results meet the agreed bar across the intended case types and the ongoing support requirement is acceptable. Identify the next scope explicitly, such as another business unit with the same approval policy.
If completion improves but correction consumes the gain, target the recurring defect and rerun the relevant cases. If the comparison is too weak or too much work remains unresolved, extend observation with a specific question and end date.
A serious control failure may require pausing the affected action even when average performance improves. Keep that decision separate from the aggregate score.
Before expansion, confirm who will own the workflow after the pilot and how they will preserve the evidence. A new application, permission boundary, or case type needs evaluation of its own.
Where Monarch fits
Monarch supports the application work needed to reach and verify the agreed result. Its Product Graph discovers application behaviour available to the connected account and represents captured, verified operations for workflows to use.
For a multi-system pilot, scope those operations against the end states in the scorecard. The business supplies the authoritative-source rules, acceptance criteria, and approval policy. Evaluate the resulting workflow with the same completion and effort measures used for the existing process.
Frequently asked questions
What is the most useful success metric for an enterprise AI pilot?
Accepted completion across a defined population of business cases is a useful starting point. Read it alongside deadline performance, corrective human effort, and serious errors. The combined view shows whether the result is useful to the operation.
Does a human approval mean the pilot failed?
No. A required approval can be part of the intended design. Record its effort separately from corrections so the evaluation shows whether the planned human role is working.
How many cases should a pilot include?
Enough to cover the intended case types and support the decision being made. Volume alone cannot establish readiness. Rare consequential failures need targeted testing, while low-volume categories may require a longer observation period.
Can we claim ROI from hours saved?
Hours saved support a capacity estimate. Financial return depends on the full cost of delivery and the value actually realised, such as reduced expenditure or additional completed work. State assumptions separately from measured outcomes.
To scope a pilot around work your business can verify, talk to Monarch.