Skip to content

Menu

Case study / Audit and fix

Cutting the cost
of an AI agent.

An AI agent reads before it works, and you pay for every word. We redesigned its tools so it reads far less. Then we measured it.

Before

400 pages

172,078 tokens read per task

After

160 pages

67,932 tokens read per task

Each jar holds what the agent reads to do one task. One page is about 320 words, or 430 tokens.

Time per task
32 s18 s
Tool calls per task
52
Tasks completed
87.5%88.9%

What that means for a bill

Reading cost each month

$1,032 $408

Tokens read each month
1.03 billion 408 million
Hours spent waiting
53 30

Arithmetic on our median results, at a price you set. We measured tokens and time, not dollars. Cached tokens cost less, so real bills are lower.

fewer tokens read per task
61 to 73%
less time per task
27 to 44%
controlled runs, three models
595

The problem

An AI agent used many small tools to do one job. Each step sent the full conversation back to the model.

The median task read more than 130,000 tokens. Tokens are the unit that AI providers bill.

What we changed

We redesigned the agent's infrastructure: the tools it calls to do its work. Now one call does the work of several. Code now handles the IDs, the versions and the retries. The model does not.

Then we ran the old design and the new design on the same 24 tasks. The wording was the same, and the database was restored before each run.

What came out

Across three models, the median task read 61 to 73% fewer tokens. It finished in 18 to 22 seconds, where it took 30 to 34 seconds before.

Task success went up for all three models in total: 87.5 to 88.9%, 87.5 to 98.6%, and 96.6 to 100%.

The impact

Most companies turn on an AI agent and never check what it costs to run. We checked.

The same agent does the same work with 61 to 73% less reading, so the cost stays low as the company grows.

An AI agent and its tools. First-party benchmark, not an independent audit. 589 of 595 runs were scored, over 24 tasks. Each model is compared with itself.

Start with a
thirty minute call.