A failed agent run still costs money.

Enterprise AI agents fail when they cannot perform a required application action, select the wrong record, lose track of the workflow, or apply the wrong business rule. Model reasoning matters, and so does the system around it.

The bill includes every attempt. Your team may also have to check what changed, finish the work, and repair the automation before it runs again.

That is why I would start an agent business case with the cost of a completed job. Token prices become useful once you know how much work the system actually finishes.

Benchmarks show why you need to test complete workflows

TheAgentCompany built a simulated software company where agents used workplace tools to complete professional tasks. Its 2025 paper revision reported that the most competitive evaluated agent completed about 30% of tasks autonomously, with difficult, longer tasks remaining a problem.

Salesforce's CRMArena-Pro research also found lower success when its evaluated agents had to handle multi-turn business interactions. That was vendor-authored research in a constructed CRM environment.

These results describe the agents, tasks, and environments tested at the time. They give you a reason to test complete workflows; they do not establish the failure rate of your production system or identify the cause of a particular failed run.

For your own evaluation, define the end state before running the agent. If the job is to prepare a renewal for review, specify the evidence it must gather, the record it must update, and the approval it must obtain.

The required action may be missing

An agent can read an invoice and still have no usable way to correct it. The required operation might live in a custom API, a private web interface, or an admin screen that its tools do not expose.

A connector's application name tells you where to start checking. You still need to test the exact operation, fields, and permissions in the connected account. Custom integrations and authorised UI automation can extend what the agent reaches; our integration-platform guide explains those options.

When an operation is missing, repeated model calls can spend money searching for a route that has never been implemented. Record the blocked action so the team can decide whether to add it, change the workflow, or hand that case to a person.

Correct tool calls can act on the wrong record

Suppose a hypothetical renewal workflow selects the wrong account at its first lookup. Every later call could execute successfully while attaching billing evidence to the wrong opportunity.

Readback needs to establish that the correct record changed and the required condition now holds. An API success response alone does not answer either question.

The business also needs to define which system is authoritative when records disagree. Application discovery can reveal available fields and operations; people supply the meaning, ownership, and approval rules the workflow must follow.

That distinction changes how you investigate a failure. An unavailable field needs an access or integration fix. A disagreement about which value should win needs a business decision before the automation can safely proceed.

Long workflows need explicit state and a recovery plan

An agent can lose track of a prerequisite, repeat a step, or declare completion before checking the result. More capable models can help with those mistakes, but the workflow still needs a record of what has happened and what is allowed next.

Separate the steps that need model judgment from those whose sequence and rules are already known. Anthropic's guidance on building agents makes this distinction between model-directed agents and predefined workflows, and recommends starting with the simplest design that works.

A timeout creates a particularly expensive ambiguity. If a request to create a billing adjustment times out, the adjustment may already exist. Retrying without checking can create a duplicate.

Use a stable request identifier where the API supports it, check the resulting state, and define a recovery path for partial work. AWS describes this problem and the role of idempotent operations in making retries safe.

An agent should stop when it cannot establish whether a consequential action happened. Put enough evidence in the handoff for a person to resolve the case without repeating the entire investigation.

Repeated exploration adds to the bill

An agent that spends each run looking up the same schema and experimenting with the same operations repeatedly pays for that work. Retaining verified application knowledge and a tested execution path can reduce the exploration required on later runs.

The knowledge still needs maintenance. An application update, permission change, or new business rule can invalidate part of a working workflow.

Long conversations can also increase input consumption when earlier messages and tool results are included in later model requests. The cost depends on what is actually sent, cached, or compacted; it does not inevitably grow at one fixed rate with task length.

Anthropic's pricing documentation separates base input, output, cache writes, and cache reads. Use the usage records and rates for your actual model and deployment, including any separately billed tools or runtime.

Caching can lower the cost of repeated context. It cannot make an incorrect record selection acceptable, so measure the resulting work as well as the token savings.

Calculate cost per completed job

Use a fixed set of business cases over a defined period. Give each case one identity so retries do not become additional jobs in the denominator.

Cost per completed job = all operating costs for those cases ÷ cases completed to the agreed standard.

Include model and tool use across every attempt, allocated platform and runtime charges, and human review, recovery, and maintenance. If a platform fee already includes model use, count it once.

Here is a hypothetical monthly example in Australian dollars. These are illustrative costs and outcomes, not vendor prices or benchmark results.

Measure Initial workflow Revised workflow
Unique cases attempted 1,000 1,000
Cases completed without human intervention 500 800
Additional cases completed with human help 300 150
Cases unresolved at the end of the period 200 50
Model, tool, and runtime cost across all attempts A$2,000 A$3,000
Allocated platform cost A$1,000 A$1,000
Human review, recovery, and maintenance A$5,000 A$2,000
Total operating cost A$8,000 A$6,000
Completed cases, including human-assisted cases 800 950
Cost per completed case A$10.00 A$6.32

In this example, model, tool, and runtime spending rises while the cost of completing a case falls. The revised workflow leaves less work for people and fewer cases unresolved.

The separate completion rows matter. A system completing 800 cases, with 300 needing human help, has a different operating requirement from one completing 800 cases without intervention.

Keep setup costs visible too. If integration and deployment cost A$20,000, allocating that over 10,000 completed jobs adds A$2 per job.

Allocating it over 1,000 adds A$20. Both are planning assumptions until you know the volume the workflow can sustain.

Unresolved cases can have consequences beyond this calculation, such as a delayed renewal or an incorrect adjustment. Track those outcomes separately with the process owner so a cheap average does not hide expensive mistakes.

Diagnose the failed cases before changing the model

Take a representative group of failures and inspect the actual traces with the people who own the process.

What happened What to investigate next
The required action was unavailable Tool coverage, account permissions, and the authorised route into the application.
The agent used the wrong record or value Record selection, authoritative-source rules, and checks before writes.
The agent skipped a prerequisite Workflow state, reasoning performance, and which steps can use deterministic code.
A retry repeated a write Request identifiers, readback, and recovery after partial completion.
A person had to resolve an ambiguous case Whether the case needs a clearer policy, better evidence, or an intentional human decision.
The agent claimed success without the result Completion checks against the target system and the business acceptance standard.

Then change the relevant part and rerun the same cases. If you are testing a model change, hold the tools and workflow rules steady so you can see what the model contributed.

Include denied permissions and partial failures in that test. A deliberate, documented stop can be correct behaviour. Count it as a handoff until the business case actually reaches its required end state.

Make useful application knowledge reusable

Monarch addresses the application work behind these workflows. It is the orchestration layer between agents and enterprise systems, using a Product Graph to discover operations and support verified actions across connected applications.

Monarch can use the operations the connected account can access and discovery has captured and verified. AI can help define a workflow, with repeatable steps captured in deterministic code and consequential actions subject to the customer's controls.

For a team repeatedly exploring the same billing portal, the goal is a saved, tested operation that the next workflow can use. The application-map approach explains how that operational knowledge can be shared across runs.

The commercial test is straightforward: does this improve completion and reduce the total effort required to finish the job? Include discovery, review, and ongoing maintenance in the comparison.

Start with one workflow whose failures you can inspect and whose result you can verify. Put the full cost next to the completed work, then use that evidence to decide what to fix and where to expand.

Connect agents to real work

Monarch discovers the product depth a company needs so agents can act across existing systems without waiting on a perfect connector catalog.

We got your email. We'll reach out shortly to set up time.