Framework

Token Spend Management

Ramp's thesis that tokens are the third pillar of business spend alongside people and vendors, unmanaged and crossing every department, where a roughly seven-hundred-times cost spread between frontier and optimally routed AI tasks creates a real management opportunity, framed as see it, understand it, control it.

The three pillars

Eric Glyman's framing names three pillars of business spend across history: people, managed since roughly the agricultural age through payroll and human resources; vendors, managed since double-entry bookkeeping emerged around 600 BCE; and now tokens, a category that in 2026 remains almost entirely unmanaged.1 Tokens are what he calls the blind spot: spend that crosses every department at once, engineering on one model, marketing on another, customer service on a third, dev tooling on a fourth, with no annual contract, no stable unit price, and no single owner inside the organization. It is structurally similar to the diffuse corporate-card spend problem that made Ramp's original product valuable, but harder, because there is no single vendor to negotiate with and the price of doing the same task can vary enormously depending on which model handles it.

The cost spread

The core optimization opportunity is a roughly 700-times spread between routing a task to a frontier model and routing the same task to an optimally chosen, cheaper model, with the spread between frontier pricing and a mid-tier model alone accounting for as much as 300 times. Organizations that route every task through the most expensive available model are leaving a large amount of avoidable cost on the table. The framing for the resulting product is see it, understand it, control it: give an organization visibility into token spend across every team and provider first, then the analytical tools to understand what the spend is actually buying, then the routing rules and budgets to act on it.

The scale of the opportunity

By 2028, industry forecasts project roughly 15 trillion dollars in business-to-business transactions handled by AI agents, on the order of 10 percent of global GDP, and Glyman's strategic bet is that those agents will eventually carry something functionally equivalent to corporate cards, routed through the same kind of spend intelligence platform that already manages human spend.1 The near-term number is already large on its own: total spend on the major frontier labs is plausibly heading past 300 billion dollars a year, which Glyman calls a reasonably conservative estimate and describes as roughly 1 percent of the entire gross domestic product of the United States being spent on tokens.2 Unlike most other categories of software spend, this one carries a real marginal cost attached to every single unit of work performed, which is what elevates it from a line item to a third mega-category alongside people and vendors, one a chief financial officer needs to be able to split between operating expense and research and development, attribute by team, and measure for return.

Part of the underlying mechanism is a predictable lag: roughly six months after the newest, most capable model ships, an open-weight model of comparable quality tends to arrive at a small fraction of the cost, sometimes on the order of a hundredth. The resulting management move is to route the advanced, expensive model only to the tasks that actually need it and shift everything else down the cost curve as the cheaper equivalent becomes available. One customer's experience shows how fast the scale arrived: Uber's own chief technology officer said publicly the company had spent its entire year's allocated AI budget in a single quarter, and, as Glyman puts it, no company budgeted for this kind of spend five years ago.2

The counter-position

Not every voice in this space agrees that the resulting number should function as a performance metric. Scott Wu of Cognition accepts the routing logic entirely and specifically rejects using the same token count to rank engineers, arguing that a token count measures an input, consumption, rather than an output, and that ranking people by consumption quietly rewards consuming more of it. He points to a project scoped at eighteen months and fifteen million dollars with an outsourced contractor that his own team instead completed internally in three months for one million dollars, the kind of measure he argues should replace token counts.3 The tension is real and belongs to both sides: Ramp's own interest is in the number mattering, since visibility is what it sells, while a company that absorbs tokens into its own product margin has an interest in the number receding from view. Read together, the more defensible position is that token spend functions well as a cost-control instrument and poorly as an individual performance instrument, and that the two get confused mainly because the cost instrument is the one that happens to come with a dashboard already attached. See output over token spend for the counter-position in full.

Practiced by

Connections

Loading connections…

References

  1. 01

    The Ramp Secret

    Eric Glyman · article · 2026

  2. 02

    The $44 Billion Company Building Self-Driving Money (Eric Glyman with David Senra)

    Eric Glyman · podcast

  3. 03

    The Future of Software & AI

    Scott Wu · podcast · 2026

Related