使用 GLM 5.3 Flash 编程一个月
One month coding with GLM 5.3 Flash

原始链接: https://wagtail.org/blog/one-month-on-glm-53-flash/

本月目标是让全部 20 亿个 token 都使用 GLM 5.3 Flash,但实际只有 50% 的使用量使用了目标模型。前半段保持在预算内,花费 68 美元,耗电约 4 千瓦时,排放 365 克二氧化碳;由于基础设施容量问题以及大量模型实验,剩余 10 亿个 token 转而使用了其他模型。 最大的一项意外支出是一个快速编码完成的 Wagtail MCP 原型。由于选错了模型,它在一夜之间消耗了 4.5 亿个 token、花费 150 美元,并耗电 5 千瓦时。虽然这个原型做出了有价值的演示,但如果规划得当,类似结果很可能只需约五分之一的成本即可实现。 总体而言,实际耗电量为 35 千瓦时,而目标是 10 千瓦时,因此这次挑战在技术上并不成功,但提供了很多有价值的信息。后续措施包括:持续在本地测量成本、能耗和结果;设定明确的实验预算;更加谨慎地选择原型和模型;采用有明确范围和限制的多智能体工作流;以及持续评估高效的模型。目标是让大多数常规 AI 推理使用一个或两个低成本的 Flash 模型,并以效率和结果为评估标准,而不是只看 token 数量。

一场关于使用 GLM 5.3 Flash 进行智能体编程的 Hacker News 讨论。作者透露,一个由 MCP 驱动的通宵构建任务意外使用了价格高得多的非 Flash 模型,消耗了 4.5 亿个 token,费用约 150 美元,预计耗电 5 千瓦时。仅一个会话就消耗了约四分之一的月度额度和 150% 的预算。 作者仍然认为这次实验很有价值,因为它成功做出了演示作品,并让大家更清楚地了解模型选择、智能体循环和 token 用量如何影响成本。主要教训是:要核实模型配置 closely monitor budgets and use cheaper models wherever possible; with modest additional effort, the cost might have been reduced to about one-fifth. 评论者表示,Flash 模型足以胜任编程、规划和代码审查,而非 Flash 版本在高级任务中具有更稳定的表现。建议的替代方案包括 Qwen、Kimi、DeepSeek 4.1 Flash 和 Mimo 2.6 Flash,同时将更强的模型留给更困难的任务。
相关文章

原文

Zooming in on the models split specifically:

Tree map of agents token usage over September 2026, with half going to GLM 5.3 Flash

The goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:

  • Successfully spent the first half of the month on just that model.
  • That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

The second half of the month didn’t go so well, with 1B tokens going to other models.

Unexpected hurdles

The cost of vibe coding

We’re pretty transparent that our experimental Wagtail MCP server is a vibe-coded prototype. Vibe coding isn’t quite what we normally aspire to, but for a prototype it’s spot on. Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

Infrastructure woes

Another unexpected hurdle was infrastructure availability issues. We’ve written extensively about comparing inference providers. Our choices work really most of the times, but it turns out they’re very popular, and do not have the same capacity as the big labs who hoard all the GPUs. We noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being so high up the Pareto frontier of relevant models for our work.

Scatter plot of AI models, with the drawn pareto frontier, GLM 5.3 Flash in the top left

This meant having to switch to other similar models (DeepSeek V4.1 Flash, Qwen 3.8 Flash). Which is very simple to do, but nonetheless unexpected!

The cost of experimentation and R&D

Last but not least, beyond using one model for day-to-day engineering, it felt essential to keep experimenting with a wide range of models, keeping up with what providers are releasing. This is particularly essential as we start to benchmark models’ performance on Wagtail tasks, where we need data across a wide range of models. Sneak peek of our benchmark:

Data table of 14 AI models reporting their accuracy as a percentage, energy use in Wh, Cost in $. Top of the table is DeepSeek V4.1 Flash with 95% accuracy and 14.9Wh energy use, $0.09 per task

It’s much easier to guide people towards leaner options with this kind of concrete data. And for us to make those options even more viable with agent skills, or our new CLI prototype, which is intended to work well with agents.

Takeways and what to do next

So technically this challenge was a failure. Only 50% usage on the target model, 1B out of 2B tokens. About 35 kWh of energy use instead of 10. But we did learn a lot, which is crucial for the current moment. Reflecting on this for October, here’s what will make it work:

  1. Constant, local usage measurement and reporting. Looking not just at tokens but also energy use and spend, and ideally how well this all leads to concrete positive outcomes.
  2. Budgeting for experimentation, not just day-to-day tasks. Making more concerted decisions about which prototypes are worth building, and how.
  3. Better prompt selection and multi-agent techniques. Orchestrator vs. scout vs. implementer vs. reviewer agents. Bounded goals. Not rocket science but certainly one more thing to learn.
  4. Keep pushing for more efficient techniques and models. The Jev-style decision diffusion models look very promising if they can run so efficiently. Latest flagship models also look like a step in the right direction on that front.

For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models. A viable target is probably that the majority of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That’s the goal for October! You should try it too, you’ll learn a lot in the process.


And come say hi at Wagtail Space 2026 in November to hear how that all pans out!

联系我们 contact @ memedata.com