Skip to content
0%
Opening 2 min

The math that decides

At some point, the question stops being “does it work?” and becomes “is it worth it?” The answer comes from a chain of three numbers, and almost everyone stops at the first one.

Cost per token is what shows up on the invoice, and it is the least useful of the three.

Cost per task already includes iterations, rework, and tool calls, and starts to say something meaningful.

Cost per outcome is the only one that decides. It adds human escalation (the cases the agent did not resolve and someone had to handle) and compares it with the previous baseline.

This has a practical effect. A cheaper model that makes more mistakes creates more human escalation and ends up costing more per outcome. Teams that compare models by token price often choose wrong.

This calculation only exists if stage 1 was done: without a measured baseline, cost per outcome is a loose number that proves nothing.

To discuss

How does this show up in your company today? What would you change first?