Buy Inference Like Inventory
Buy Inference Like Inventory
Thesis: In agentic commerce, cheaper tokens do not buy you a cheaper P&L. They buy you more agent steps per task — and the bill follows session complexity, not the model price list. Founders who still treat inference as “R&D that will sort itself out” will lose to operators who treat it like inventory: purchase limits, turn targets, spoilage rules, and kill criteria per surface.
I am building an AI-native commerce company. I want agents in merchandising, support, fraud review, and buyer-side discovery. I refuse to run them on blank-check compute.
The Signal: Prices Down, Budgets Up
The timeline is loud about two things that sound contradictory until you put them on one spreadsheet.
First, capability is getting cheaper. Bargain open and frontier inference keep racing toward commodity pricing. DeepSeek-class cost curves, analytics agents topping benchmarks at a fraction of last year’s spend, and “race to zero” model chatter all push the same headline: intelligence is no longer scarce.
Second, enterprise AI spend is still exploding. One useful thread put hard numbers on the joke: token prices down hard year over year, while annual enterprise AI budgets jumped multiples — with the majority of that budget landing on inference, not training. Another cut said the quiet part out loud: price cuts are not making agents cheaper when most agentic spend is burned on response refinement — retries, re-plans, self-correction loops, and tool thrash.
That is Jevons for operators, not economists. When a unit of intelligence gets cheaper, you do not bank the savings. You attach more agents to more workflows. A human support lead asks one good question. An agent can burn forty tool calls to look thorough.
Founder translation:
| What the market celebrates | What hits your COGS |
|---|---|
| $/M tokens falling | Tool-call fan-out per ticket or SKU |
| ”Agent that never sleeps” demos | Always-on inference with no inventory turn |
| Smarter reasoning | Longer trajectories before acceptance |
| One universal checkout button for AI | More successful purchases and more failed agent loops behind them |
I already argued that intelligence is cheap and verification is the moat. This essay is the cash-flow sibling: cheap intelligence is still expensive if you cannot cap how much unverified work you generate.
Session Complexity Is the Real Unit
Per-token pricing is a vendor UI. Your unit economics live one layer up.
Commerce agents do not complete “a chat.” They complete a task trajectory:
- Read catalog, policy, or ticket state
- Call tools (search, ERP, shipping, fraud, CRM)
- Re-plan when the world disagrees
- Write a proposed action
- Wait for acceptance — or retry until they hit a wall
Happy-path demos hide steps 3 and 5. Production is mostly 3 and 5.
That is why “agentic workflows burn 8-12x the tokens of a single completion” is not a lab curiosity. It is a margin model. If your support agent, pricing agent, or catalog-enrichment agent averages N retries before a human merge owner accepts the output, your COGS is N times the brochure number — and N is a product choice, not a law of nature.
The teams that survive this do not “minimize tokens” as a vibe. They measure:
- Cost per completed task (refund resolved, SKU fields accepted, fraud case closed)
- Acceptance rate on first proposal (how often verification says yes without another loop)
- Retry budget (hard stop before the agent thrash-spends the margin on a $40 order)
- Spoilage (inference spent on tasks that never ship to a customer or operator)
Treat those numbers the way a merchant treats inventory turns. Slow turns are not “still learning.” They are capital stuck on the shelf.
Universal Checkout Makes This Worse — and Better
Agentic commerce heat this week is not only model leaderboards. It is rails: shopping agents that research, compare, and buy inside the conversation; universal checkout for AI; platform storefronts that expose catalogs to ChatGPT, Gemini, Copilot, and peers; privacy posts noting that a machine purchase may never fire your cookie banner.
That is distribution good news if agents prefer you. It is cost bad news if you answer machine demand with unbounded internal agents.
Every time a buyer-side agent asks a harder question — total landed cost, origin claim, return window, substitute SKU — your seller-side stack will want to answer with more retrieval, more tool use, more “let it think.” Defaulting agent access without a cost model is how you industrialize margin leakage the same way swarming support agents industrializes wrong answers.
Think big: inference is a permanent line item next to payment fees, ads, and fulfillment — a channel cost of machine demand.
Step small: pick one agent surface shipping this week (catalog enrichment, L1 support, or fraud pre-screen) and put a weekly dollar cap plus a max tool-call budget on it.
Do smart: kill or throttle any surface that cannot clear a target cost per accepted outcome for two consecutive weeks. No romance. Spoilage rules.
Monday Morning Operator Playbook
Do this before you buy another model or spin another “autonomous retail” stack.
1. Rename the budget
Stop booking agent inference under “innovation” if it touches orders, tickets, or catalog truth. Put it on COGS or channel cost for that surface. What gets a P&L owner gets a kill switch.
2. Cap retries, not vibes
Vendors will keep cutting list prices. That does not cap your spend. Product does:
- Max tool calls per task
- Max re-plans after verification reject
- Escalate-to-human before the trajectory costs more than the order contribution
The X-side advice is blunt and correct: cap retries per task, not price per token.
3. Instrument acceptance as the win condition
A finished stream is not success. Success is:
- Human merge owner accepted, or
- Policy gate auto-accepted with an audit trail, and
- No reverse within the dispute window you care about
If you only log “agent responded,” you are counting inventory received, not inventory sold.
4. Portfolio the experiments
Run a few agent surfaces like SKUs: small lots, clear sell-through, clear exit. Blank-check “let it cook” is how R&D becomes a silent tax on every order.
5. Prefer leverage over autonomy theater
Embed where spend already happens (checkout, ticket queues, catalog admin). Instrument before scale. Verification before wallet. Distribution before model shopping. Same slogan, cash-flow version: do not buy infinite inference to paper over a missing acceptance path.
The Claim Worth Arguing
The winning AI-native commerce companies will not be the ones with the cheapest model API. They will be the ones with the best inference inventory discipline: which tasks earn compute, how many retries are allowed, what acceptance means, and when a surface is marked spoilage and pulled.
Counterexample I want: a merchant running high agent volume with unbounded session complexity who still expands contribution margin for four quarters. Show me the cost-per-accepted-task series, not the demo GIF.
If you are building or operating in this stack, disagree on X with a real number from your shop — best failure mode or best unit-econ win. I will update how we buy inference the same way we buy inventory: small, measured, and ready to cut.