One month with AI agents cost Thibaud Colas (Wagtail) 2 billion tokens: $68 on the cheap model and $150 in a single night with the wrong one.
Colas, a member of Wagtail’s core team, published the breakdown on October 2, 2026 as the conclusion of a challenge he set for himself: spend all of September coding with a single efficient open-weight model, GLM 5.3 Flash. By his own standards, the challenge failed. What’s valuable is the bill, because it shows where agentic coding money really goes.
What happened during the month?
The first two weeks went according to plan. He used only GLM 5.3 Flash, and his portion of the month cost $68. For that part he also reports about 4 kWh of energy and 365 grams of CO₂.
The second two weeks didn’t go that way. 1,000 of the 2 billion tokens went to models other than the target, for three reasons:
- A nighttime session with the wrong model. This is the expensive part, and we’ll see it below.
- Provider congestion. The open model inference providers that Wagtail uses don’t have the GPU capacity of the big labs. GLM 5.3 Flash became noticeably slower, so he had to switch to DeepSeek V4.1 Flash and Qwen 3.8 Flash.
- R&D. Wagtail is building a benchmark of models on Wagtail tasks, and that requires running many models, not just one.
Final tally: 50% of tokens on the target model, and about 35 kWh of energy versus a goal of 10.
How much does vibe coding with the wrong model cost?
In this case, $150 in one night. The wrong model was GLM 5.3, the full model and not the Flash variant. The post doesn’t name it, but Colas confirmed it in the Hacker News discussion. He was prototyping an experimental MCP server for Wagtail, and did all that build in vibe coding mode on the non-Flash model without realizing it.
The result: 450 million tokens, $150 and 5 kWh, practically overnight. In the HN thread he mentions that this single session amounted to nearly a quarter of the month’s spending and 150% of the budget. He attributes the problem to two things: the price difference between the two models, and an agent that kept running on its own, in “vibe” mode, all night, working much harder than the task required.
The price gap is obvious. These are the list prices from TensorX per million tokens as of October 3, 2026. TensorX is one of the two providers he used.
| Model | Input | Cache Read | Output |
|---|---|---|---|
| GLM 5.3 Flash | $0.20 | $0.05 | $0.50 |
| GLM 5.3 | $1.75 | $0.44 | $4.50 |
That’s a difference of approximately 9 times in each line. In a long session of an agent, that multiplier applies to every token the agent consumes, and an unsupervised loop consumes a lot of them. (If you need a refresher on how tokens are charged, we explain it in What are tokens in AI and why do they charge you for them?.)
Colas’s own estimate is that he could have reached similar results with a cost “probably 5 times lower”, without much more effort. The MCP server works and has a demo, so the money wasn’t wasted. But it yielded much less than it could have.
What’s the best AI for coding without overpaying?
According to Wagtail, it’s not a single model but a division of labor. This is their published recommendation, current as of October 2026:
- GLM 5.3 Flash for 95% of tasks, with the reasoning level set to “high” or “max” in your harness.
- GLM 5.3 for the most demanding 5%.
- Plan on switching models every two or three months.
Several developers in the HN thread describe the same pattern from the other side: use the strong model to write a detailed, unambiguous plan, and let the Flash model execute it. One of them sums it up this way: GLM 5.3 reasons more holistically about the code, while Flash, with a clear plan, executes it well for a fraction of the price.
Colas also lists what’s going to change in October. Choosing the model is just one of four points:
- Local and constant measurement of tokens, energy, and spending.
- A separate budget for experimenting, not just for day-to-day work.
- Multi-agent patterns with bounded objectives: orchestrator, explorer, implementer, and reviewer roles.
- Keep pushing toward more efficient models and techniques.
His conclusion for day-to-day work is that focusing on one or two flash-type models is perfectly viable. For him, the majority of inference work should run on those models, measured in cost or energy rather than tokens.
How much does an AI agent cost per month and how much should you budget?
Wagtail recommends $10 per month per developer for personal use and $100 per month for professional use, as published on October 2, 2026. Their position is that, by choosing the model well, those budgets are enough for programming sessions with intensive agent use.
The billing model matters as much as the number. Wagtail recommends 100% pay-as-you-go billing with spend quotas rather than subscriptions. Their reasoning: hitting the cap should push you toward smaller models or tighter prompts, with less context and more cache. A subscription hides that signal.
The same page gives approximate price levels as reference, based on cost per task:
- $1 per million tokens: flagship models, for planning and for the most complex 5% of tasks.
- $0.10 per million: a mid-range model for most tasks.
- $0.01 per million: occasional tasks, price-sensitive.
How do you see how much your agents are actually spending?
Colas measures his consumption with AgentsView. It’s a local-first tool with MIT license that reads the session files your code agents already write and converts them into a queryable file, with reports of tokens and costs. It supports over 60 agent formats, including Claude Code, Codex, Cursor, Gemini CLI, OpenCode, and Aider.
Installation on macOS or Linux:
curl -fsSL https://agentsview.io/install.sh | bash
```Windows installations, desktop app, pip, and Docker are in the quickstart guide: agentsview.io/docs/quickstart.
Three commands worth knowing:
```bash
agentsview usage daily # last 30 days by agent, model, and project
agentsview usage statusline # today's spend, for your status bar
agentsview capture run -- claude -p "fix the tests" # exact consumption of a non-interactive run
The second one would have caught the overnight session in time.
A caveat that Wagtail itself points out: AgentsView’s token counts are reliable, but its costs are indicative. They’re calculated using public pricing data from LiteLLM and OpenRouter, which doesn’t always match what your provider charges you. Use it to spot anomalies, and use your provider’s billing for the real figure.
What don’t these numbers tell you?
- It’s one developer’s month, not a benchmark. The work mix was Wagtail core development, an MCP prototype, and model R&D. Yours will be different.
- Price depends on the provider. The same model costs different amounts depending on your inference provider, and cache behavior makes a huge difference to the bill. One HN commenter argues that buying DeepSeek and GLM directly from their creators is cheaper than buying them through Neuralwatt.
- Energy figures are GPU-only. They come from Neuralwatt, one of Wagtail’s providers, which measures energy per request. At least one HN commenter questions their methodology.
- The model benchmark in the post is Wagtail’s work in progress. In it, DeepSeek V4.1 Flash leads with 95% accuracy and $0.09 per task. These are self-reported figures and not yet fully published.
What’s useful about this story doesn’t depend on any of those numbers. With pay-as-you-go billing, model selection is a budget control. The expensive mistake didn’t come from picking a bad model—it came from leaving a wrong default running without supervision.
What about you? Do you know how much your AI agents spent last month, and which model took up most of it?
- I always use the same model
- I plan with a large one and execute with a cheap one
- I switch models depending on how the month’s budget is going
- I prefer another option (tell us which)
Comment below or, if this article reached you by email, reply directly to the email: your reply gets published here.
Related: Netlify tested 11 models with the same prompt: and published the credits each run consumed