Paper
What it costs to keep an agent running
A. Silva, J. Ferreira
·
·
19 pages
Twelve months of operating cost data from four production agent deployments, broken down by inference, retrieval, monitoring, and human review. Human review was the largest line item in three of four cases, and grew as a share of cost in all four. Inference, the line most teams optimise first, was never the largest. We report the numbers, the drivers behind them, and where reduction effort actually paid off.
The question every client asks before deployment is what the model will cost to run. The honest answer is that the model is rarely the expensive part.
Key findings:
Human review was the largest cost line in three of four deployments, and second in the fourth
Review cost grew as a share of total in all four systems as usage scaled
Inference cost fell in absolute terms across the year in every deployment, driven by model price drops the teams did nothing to earn
The only intervention that durably cut review cost was narrowing what the agent attempts, not improving the model
Method: monthly cost data from four deployments we operate or co-operate, normalised per resolved task. Review time was measured from ticketing logs, not estimates.
What it changes in practice: we now scope review load in the proposal, alongside inference. A system that saves model cost by escalating uncertainty to humans has not saved anything.
CITATION
Silva, A., & Ferreira, J. (2025). What it costs to keep an agent running. Forja Studio.
CONFRONTO