Skip to content

What Does Your AI Agent Really Cost? (Including the Runs That Fail)

Your dashboard shows what you spent, not what a finished task costs, how much went to failed runs, or what one stuck loop would cost overnight. costcheck measures all three with three lines of Python. Free, open source, and nothing leaves your computer.

Emmanuel Naweji

5 min read

What Does Your AI Agent Really Cost? (Including the Runs That Fail)

Your AI provider's dashboard tells you one thing: how much you spent.

It doesn't tell you the three numbers that actually matter:

  1. What one finished task costs.
  2. How much of your bill went to runs that produced nothing.
  3. What one stuck run would cost if it looped all night.

I built costcheck to answer those three questions. It's a small, free, open-source Python tool. No dependencies, no account, and nothing leaves your computer.

What Does Your AI Agent Really Cost? (Including the Runs That Fail)

Prefer to watch? The video guide walks through everything in this post step by step.


The problem: you're measuring the miles, not the arrivals

Think of it like a taxi. Your dashboard shows what you paid per mile. But what you care about is what it costs to arrive.

If an agent run fails halfway through, you still paid for every token it used. That money bought nothing. Multiply that across hundreds of runs and the "cheap" model you switched to can quietly become the expensive one, because it fails more often.

And then there's the scary one: an agent stuck in a loop at midnight, calling the model over and over until someone notices in the morning.

You can't fix what you can't see. So step one is seeing it.


The three numbers costcheck gives you

1. Cost per successful task. Everything you spent, divided by only the runs that worked. This is the true number. If you switch to a cheaper model that fails more, per-request cost goes down while this number goes up.

2. Failed-run tax. The share of your bill that bought nothing.

3. Overnight exposure. Your most expensive step, at your own measured speed, repeated for twelve hours. In plain words: what one stuck loop starting at midnight would cost you by morning.

Here's what a real report looks like:

Your agent cost check (12 runs · claude-sonnet)

What it measuresResultWhat it means
Runs recorded12
Succeeded8the denominator
Did not succeed4still counts on top
Total spend$1.6522
Cost per request$0.1377pays by the mile
Cost per SUCCESSFUL task$0.2065pays for arrival
The gap50% higherand it is the true number
Failed-run tax$1.196672% of spend
Most expensive step$0.0117
Your pace3.0s per model callmeasured from your runs
Twelve hours unattended$168.95if one run loops overnight

Look at that: 72% of the spend bought nothing, and the real cost per finished task is 50% higher than the dashboard suggests. And a single step that costs about one cent becomes a $169 problem if it loops overnight.


Try it in 30 seconds (no API key needed)

You need Python 3.9+ and git. On Mac, open Terminal. On Windows, open PowerShell and type py wherever you see python3.

git clone https://github.com/Here2ServeU/agent-cost-check
cd agent-cost-check
python3 -m costcheck demo
python3 -m costcheck report

That uses fake data, so you can see the report before touching your own code.

New to Python or git? These short videos set everything up from scratch: Mac setup · Windows setup.


Add it to your agent: three lines

costcheck has zero dependencies, so you can just copy the costcheck folder into your project. (Or install it with python3 -m pip install git+https://github.com/Here2ServeU/agent-cost-check.)

Here's a normal model call:

resp = client.messages.create(model="claude-sonnet-5-5", max_tokens=1024, messages=msgs)
print(resp.content[0].text)

Here it is with costcheck. Nothing else changes:

import costcheck

# 1. start measuring one task
with costcheck.run(task="my-first-task") as r:
    resp = client.messages.create(model="claude-sonnet-5-5", max_tokens=1024, messages=msgs)
    # 2. after every model call
    r.record(resp)
    print(resp.content[0].text)
    # 3. only if the task really worked
    r.succeeded()

What each line means:

  • costcheck.run(task=...): "Everything inside here is one task."
  • r.record(resp): reads the token counts and model name. Works with Anthropic, OpenAI and Gemini.
  • r.succeeded(): you decide what "worked" means. If you never call it, or your code crashes, the run counts as failed. That's the honest default.

If your agent calls the model several times in one task, call r.record(...) after each one, all inside the same block. Not using an SDK? Pass the numbers yourself: r.record(input_tokens=1200, output_tokens=300, model="gpt-4o").

Run your agent a dozen times, then:

python3 -m costcheck report

A few other handy commands:

python3 -m costcheck check  # six yes/no questions: what would stop a runaway run?
python3 -m costcheck card  # a plain-text summary you can paste into a post
python3 -m costcheck reset  # delete everything and start over

Your data stays yours

Every result goes into one file, .costcheck/runs.jsonl, in the folder you run from. One line per task. No server, no network calls, no telemetry. Open it and read it yourself. And if you want it gone, costcheck reset clears it.

The tests use only the standard library. If they pass, every number in the report is arithmetic you can check by hand.


What costcheck does not do (on purpose)

It measures. It does not stop anything. No cap, no loop detection, no kill switch, no alert.

That's deliberate. You can't pick a sensible cap until you know what a normal run costs. costcheck tells you that first.

The six mechanisms that do stop a runaway agent are what I teach in The Agent Cost Problem:

  1. Instrument it: you can see it
  2. Cap it: it is bounded
  3. Detect the loop: 6 steps, not 14
  4. Checkpoint: stopping is cheap
  5. Cost per successful task: the true number
  6. Alert on burn rate: you find out while it happens

Seven sessions, plus a capstone repo where each mechanism is proven by running it off, then on. The whole safety layer is under four hundred lines.


Your next step

Measure before you optimize. Run the demo today, add three lines to your agent this week, and look at your real numbers.

Build systems, not surprises.

Share

Technology

Meet IT Hub: Your Daily Tool for Staying Sharp in Tech

A tool built to keep your Cloud, DevOps, SRE, FinOps, AI/ML, Robotics, and Quantum Computing concepts fresh, five minutes a day, so nothing you've already learned quietly fades before the interview that needs it.

4 min read

Technology

You Are Already Behind in AI. Here Is How to Catch Up Before It Is Too Late

Most people use AI. Very few people build with AI. There is a difference, and that difference is your next career move. A chatbot answers questions. An agent gets things done. FinTech. Healthcare. Production-ready. No fluff. The course is coming soon. The books are coming too.

6 min read