BLOG
How much an AI agent costs when nobody’s looking
October 7 | Sharam Dadashnia
An AI agent can get started on a small budget. A trial account can be set up quickly, as can an initial workflow. The monthly bill seems manageable at first.
Then the agent goes live.
It processes emails, reads attachments, searches for information, calls APIs, creates drafts, triggers follow-up processes and asks for clarification when in doubt. A single model call turns into several steps. One agent becomes ten. A pilot project turns into a network of dependencies.
The real question regarding costs is therefore not: How much does a prompt cost?
But rather: What does a completed business transaction cost – from the initial input to a traceable decision?
This is precisely where many companies lose track.
Recent analyses indicate that token costs alone account for only part of the bill. The more autonomously a system operates, the greater the impact of factors such as tool calls, infrastructure, integration, monitoring, testing, human checks and governance. Computerwoche summarises this total cost perspective for Agentic AI.
Why low initial agent costs can be misleading
A traditional chatbot answers a question and ends the interaction. Its cost can be calculated relatively easily: input, response, done.
An AI agent works differently. It pursues a goal across several steps. It can consolidate information from multiple sources, communicate with an ERP system, generate a follow-up query, hand a task over to a human, and continue working once the human has given the go-ahead.
Each of these steps incurs a cost. An example from the quotation process:
- A customer enquiry arrives by email.
- The agent reads the message and any attachments.
- It extracts technical details.
- They cross-check the data against product information or ERP master data.
- Missing details lead to a follow-up enquiry.
- A member of staff checks the draft reply.
- Once approved, the CRM is updated and the process is forwarded.
The visible model response is only a small part of this process. The total costs arise throughout the entire process chain, and therefore if no one keeps an eye on things, this can quickly lead to five typical cost traps.
1. Token costs do not increase linearly with the number of enquiries
Tokens are the unit of measurement for many language models. However, for agents, consumption depends not only on how many enquiries are received.
The following factors are also relevant:
- the length of the inputs and responses
- the volume of accompanying documents and contextual data
- additional research or retrieval steps
- tool calls and their results
- repeats in the event of errors or uncertainty
- verification and reflection loops
- the number of agents involved
A single process can thus trigger several model calls. It becomes particularly costly when agents work with context windows that are too large, process the same content multiple times, or reload the entire process dossier with every follow-up query.
The challenge is fundamental in nature: an agent does not operate in a fully deterministic manner. Two similar processes can generate different processing paths, different response lengths and, consequently, different costs. ‘etailment’ describes this dynamic, which is difficult to plan for, in token-based agent-driven services.
That is why it is not enough to report model costs monthly under a single cost centre. Companies need to see which agent incurs which costs – and for which process step.
2. The most expensive calls are often not the model calls
An agent without access to corporate systems remains merely an assistant. It is only through integrations that it becomes an operational part of the process.
It reads data from the CRM, checks orders in the ERP, opens documents in the archive, creates tickets, updates records or initiates follow-up processes. Every integration must be developed, tested, secured, documented and operated.
This incurs costs that are often overlooked in a quick pilot calculation:
- API management and interface operation
- identities, roles and permissions
- data preparation and data quality
- encryption and access controls
- error handling and recovery procedures
- version management for interfaces and agent logic
- monitoring and alerting
An agent that analyses a quotation request may be cost-effective. An agent that autonomously modifies business data as a result requires a robust technical environment. This is not an optional extra, but a prerequisite for secure operation.
3. An error costs more than a response
The costs of a faulty agent cannot be reduced to tokens.
If an agent misclassifies a query, it may simply result in rework. However, if it sets the wrong customer status in the CRM, closes a ticket prematurely or includes unauthorised information in a draft response, the cost can rise significantly.
The question is therefore not just: How much does the agent cost per month?
But also:
- How many results need to be checked?
- How many cases are escalated?
- How quickly are errors detected?
- Can an action be undone?
- Is it possible to trace the basis on which the agent acted?
- Who is responsible for the outcome?
In the case of business-critical processes, a human checkpoint is not a sign of a lack of automation. It is a deliberately implemented quality and risk control mechanism.
A well-planned process therefore does not make a blanket distinction between ‘fully automated’ and ‘manual’. It defines precisely which steps are automated, in which exceptional cases a human is involved, and which decisions must never be made without approval.
4. Operations only begin after go-live
A productive agent is not a one-off implementation. Models change. Interfaces change. Data structures change. Business rules change.
What works reliably today may produce incorrect results in three months’ time because a form, a product structure or an internal policy has changed.
Operations therefore involve the following on an ongoing basis:
- Quality measurement and technical spot checks,
- Tests prior to changes to the model, prompt or process logic,
- Monitoring of error rates and execution times,
- Adjustment of permissions,
- Maintenance of knowledge sources,
- Handling of exceptions,
- regular review of benefits.
That is the difference between a demo and an enterprise system.
A prototype demonstrates that an agent can, in principle, carry out a task. An operational process demonstrates that it carries out this task in a traceable, secure and cost-effective manner over a period of months.
5. Without allocation, there is no cost management
In many organisations, expenditure on AI models, cloud resources, integration platforms and external APIs is managed under separate budgets. The business unit sees its licence costs. IT sees infrastructure costs. Operations sees support costs. Nobody sees the total costs of a single process.
This means there is no basis for making a cost-effective decision.
A robust cost model allocates expenditure across at least four levels:
Level | Key question |
Business process | What is the cost of processing a complete transaction? |
Agent | Which agent causes which usage, errors and follow-up work? |
Process step | At which stage do a particularly high number of calls, waiting times or escalations occur? |
Business outcome | What measurable improvement is there in relation to the total costs? |
The key metric is not ‘cost per agent’. Nor is it ‘cost per token’.
A more meaningful metric is, for example:
What does it cost to process a complete and correct quote enquiry – compared with the previous process?
Or:
What are the costs per correctly classified service case, taking into account rework, checks and escalations?
It is only this perspective that links technology expenditure to operational impact.
FinOps for agents: Three rules to get you started
Cost control doesn’t have to start with a major controlling project. Three rules quickly create transparency.
1. Define a budget per agent and per task
Every productive agent needs a usage budget. Not just per month, but also per task or case type.
This protects against endless loops, excessively long contexts and uncontrolled retries. If an agent reaches its limit, the process can pause, switch to a more cost-effective model or be escalated to a human.
2. Use the right tool for the right step
Not every task requires an autonomous agent.
- Fixed thresholds belong in rules or decision tables.
- Recurring data transfers belong in integrations and workflows.
- Simple extraction or summarisation can be carried out with a limited number of model calls.
- Agents are useful where context, language, documents, judgement and the dynamic use of tools are actually required.
This distinction reduces costs and increases reliability. A deterministic process step is generally cheaper, faster and easier to verify than an agent-based decision.
3. Monitoring costs, quality and risk together
An agent that operates cost-effectively but generates many errors is not economical. Nor is a high-quality agent that becomes disproportionately expensive due to unnecessary loops.
That is why three metrics go hand in hand:
- Costs: model usage, infrastructure, tool calls, and operations.
- Quality: hit rate, rework, escalations, and human corrections.
- Risk: permissions, policy breaches, sensitive data, and decisions that cannot be traced.
It is only when viewed in context that it becomes clear whether an agent is helping the process or merely generating activity.
From a cost centre to controllable process performance
This is a key function of an Agent-based Process Orchestration Platform.
Scheer PAS integrates processes, data, APIs and AI agents into a single execution environment. This means that an agent can not only be technically invoked, but also assigned to a specific process step. Its role, its permissions, its data handovers, and its results become part of a controllable overall process.
This creates transparency in areas where isolated agent projects often remain obscure:
- Which agent was used in which process?
- Which systems and data sources did it use?
- How many loops and tool calls were required?
- Where was human approval obtained?
- What were the costs incurred in relation to the outcome?
- Which process version was active at the time of processing?
Cost management thus does not become a retrospective monthly review. It becomes an integral part of the ongoing process.
The right question is not: ‘Is the agent cheap?’
An AI agent does not have to be cheap, but rather it must make economic sense.
An expensive agent can be worthwhile if it reliably speeds up complex cases, reduces errors and relieves specialist staff of time-consuming research or document checks. A cheap agent, on the other hand, can end up being expensive if it operates un ly, initiates incorrect follow-up processes and generates a constant need for corrections.
Anyone who views agents merely as cheap digital labour is underestimating their actual architecture. They are software components with variable usage costs, permissions, data access, integrations and operational responsibilities.
The good news is that these costs can be managed.
The prerequisite is that agents do not run as isolated experiments alongside business processes. They must be embedded within processes, with clear rules, measurable results, technical control points and full traceability.
Then the question ‘How much does an AI agent cost?’ becomes a better one:
What measurable contribution does this agent make
– and what does each successfully completed task cost us?