Every lab quotes "dollars per million tokens," but no two companies' tokens are the same thing. Over the past three years, the numbers on this price list dropped by two orders of magnitude; at the same time, enterprises' actual AI bills went up. This article unpacks the pricing mechanism itself—how it's calculated, why it dropped so fast, why bills rose anyway, and what the industry is shifting toward from "per-token billing."
TL;DR · 30-second read
Token prices dropped two to three orders of magnitude in three years, but that's not a commodity price list you can directly compare—enterprises' actual bills are rising instead, and the industry is switching its unit of account from "how many tokens were used" to "how much work got done."
Counter-consensus insight: when the price war plays out, the winner won't be the one with the lowest sticker price, but the one with the highest ratio of "completed work per dollar spent"—a metric no one is currently willing to publicly disclose.
"Dollars per million tokens" sounds like a commodity price—like oil, like electricity, a universal unit of measure where different suppliers quote different prices and you just compare. But a token has never been a universal unit. Each lab uses its own tokenizer to slice the same text into different numbers of fragments; different slicing methods mean different bills.
This gap isn't theoretical. The Register's July 2026 test found that for the same 2,888-character TypeScript file, Anthropic's new tokenizer produced 73% more tokens than GPT-5.x's tokenizer, and 32% more than Anthropic's own old tokenizer—if you convert Anthropic's sticker price to an "apples-to-apples" per-character cost with OpenAI, Opus 4.8's actual price isn't the officially listed $5/$25, but $7.50/$37.50.
TensorZero's cross-vendor benchmark quantifies this even more directly: under tool-call-intensive workloads, claude-opus-4-7's actual cost is 5.3× that of gpt-5.4, while their price lists differ by only 2×. More troublingly, the rankings flip—Gemini is cheapest on plain text and structured data, but on content like tool definitions it's actually 46% more expensive than OpenAI. The question "who's the cheapest provider?" has no fixed answer; it depends on what you're sending.
The numbers on the price list are essentially "per kilogram" prices—but one kilogram weighs differently on every vendor's scale.
This isn't any vendor playing tricks; tokenization efficiency is itself part of model quality—finer slicing can express more precise semantics, especially for code and non-English text. But this means the first thing to acknowledge before comparing token prices is: this price list was never designed for direct cross-vendor comparison.
When GPT-4 launched in March 2023, running a model of equivalent capability cost $30/million input tokens, $60/million output tokens. a16z's LLMflation analysis tracked the inference cost of equivalent-quality models and found prices dropped roughly 62× since then; extending the timeline back to 2021 as a starting point, the decline rate for equivalent capability tiers is roughly 10× per year.
Two milestones smashed steps into this curve. In December 2024, DeepSeek V3 went online at $0.14/million input tokens— frontier-level capability at one-hundredth of GPT-4's launch price. A month later, DeepSeek R1 offered reasoning capability at $0.55/$2.19, while OpenAI's o1-preview had launched four months earlier at $15/$60—a 97% discount at equivalent capability.
In May 2026, DeepSeek permanently cut V4-Pro prices by 75%, from $0.0145–$3.48/million tokens down to $0.0035–$0.83—not a promotional discount, but nailing down a new floor. Reportedly, Sam Altman was considering大幅 cutting OpenAI's developer pricing in response, predicting Anthropic would match the reduction.
Update (2026-09): The "nailed-down floor" judgment didn't hold for long. DeepSeek introduced peak/off-peak pricing on August 16, 2026, raising V4-Pro output from $0.87/million to $3.96 peak / $1.98 off-peak—a 355% peak increase, rendering the earlier "permanent 75% cut" statement void. Less than a month later, DeepSeek replaced V4-Pro itself with the cheaper V4.1 Flash ($0.15–$0.30/million input, $0.60–$1.20/million output, depending on peak/off-peak): starting September 14, 2026, all requests sent to V4-Pro are automatically routed to V4.1 Flash and billed at Flash prices, until the future V4.1 Pro launches. Altman's prediction of price cuts was only half fulfilled— OpenAI cut GPT-5.6 Sol pricing by over 20% on August 22 ($5/$30 down to $4/$20, promotional period through November 21, 2026); but Anthropic didn't match the magnitude of cuts, taking a different path— canceling its planned September 1 price increase, setting Sonnet 5's introductory $2/$10 as the permanent standard price. Three companies gave three different answers in the same month—raising prices, cutting prices, and abandoning a price hike. This curve is harder to predict than either "monotonically decreasing" or "sawtooth"; the only certainty is that no floor stays nailed down for good.
But this curve isn't monotonically decreasing. An inflation index tracked using the same methodology shows that the "rate of decline" itself peaked in mid-2025 and then retreated— the July 2026 reading is 333× cheaper than GPT-4's launch price, but 3× more expensive than this index's own historical low. The reason is straightforward: each new model generation launches with prices set at the "new capability tier," causing frontier model sticker prices to jump up first, then gradually fall with competition and efficiency gains, forming a sawtooth rather than a smoothly declining curve. Analysts generally believe the 10×-per-year decline rate of 2021–2025 won't continue indefinitely—the easily captured optimization space has been largely exhausted; a more realistic expectation post-2027 is 3–5× per year, narrowing to 1.5–2× thereafter.
Remember the caveat from §01—the table below compares sticker prices, not actual bills. Tokenization efficiency differs, so real-world gaps may be larger or smaller.
| Model | Input / M tokens | Output / M tokens | Notes |
|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $50.00 | Released late Aug 2026, Anthropic top-tier new flagship |
| Claude Opus 5 | $5.00 | $25.00 | Batch $2.50/$12.50, cache hit $0.50 |
| Claude Sonnet 5 | $2.00 | $10.00 | Planned 09-01 increase to $3/$15 canceled; permanent standard price (2026-08-10 official announcement) |
| Claude Haiku 4.5 | $1.00 | $5.00 | — |
| GPT-6 Astra | $10.00 | $50.00 | Released 2026-09-03, OpenAI new flagship, priced above entire GPT-5.6 series |
| GPT-5.6 Sol | $4.00 | $20.00 | 2026-08-22 limited-time cut (from $5/$30, through 2026-11-21); long context >272K is $8/$30 |
| GPT-5.6 Terra | $2.00 | $12.00 | 2026-07-30 20% price cut |
| GPT-5.6 Luna | $0.20 | $1.20 | 2026-07-30 80% price cut |
| DeepSeek V4-Pro | $0.66→$1.32 | $1.98→$3.96 | 2026-08-16 switched to peak/off-peak (major increase from May "permanent" price); routed to V4.1 Flash from 09-14 |
| DeepSeek V4.1 Flash | $0.15→$0.30 | $0.60→$1.20 | Released 2026-09-10; took over all V4-Pro traffic from 09-14, lower pricing |
Sources: eesel AI Anthropic pricing summary · Anthropic official pricing page (Sonnet 5 price increase canceled, Fable 5.1) · OpenAI official blog (GPT-5.6 pricing) · Business Standard/Reuters (Sol price cut) · OpenAI official (GPT-6 Astra) · InfoWorld (DeepSeek May price cut) and August report in same publication (peak/off-peak price increase) · DeepSeek official (V4.1 Flash). Pricing changes frequently; the table body is a mid-August 2026 snapshot, rows with dates have been updated to September 2026, and this is still not real-time data.
Converting these numbers into a concrete scenario is more intuitive: processing 100 million output tokens per month, Claude Opus 4.7 costs about $2,500, GPT-5.5 about $3,000, and DeepSeek V4-Pro (estimated at the May price of $0.83/million) about $348—roughly 7× cheaper than Anthropic and 9× cheaper than OpenAI. OpenRouter's routing data shows that Chinese mainland models' share of token consumption on its platform grew from under 2% at the end of 2024 to over 50% by June 2026. Where Western frontier models still hold their ground is on the hardest reasoning tasks, long-chain agent work, and "get it right the first time" scenarios— when the cost of errors is high, the premium pays for itself; but on large volumes of repetitive, fault-tolerant workloads, open-weight models have pulled the price anchor very low.
Update (2026-09): The monthly comparison above uses DeepSeek's May price of $0.83/million, which itself didn't survive to this article's update—by August it had risen to a peak of $3.96/million, and in September it was replaced by the cheaper V4.1 Flash ($0.60–$1.20/million output). The direction hasn't changed (open-weight models are still dramatically pulling down the price anchor), but please use the latest numbers in the §03 table for specific multiples, and don't use the $348 figure here.
The week DeepSeek smashed through the price floor, Microsoft CEO Satya Nadella said only one thing: "Jevons Paradox strikes again"—a phrase repeatedly cited by Fortune, because it precisely captured what happened next. In 1865, economist William Stanley Jevons observed that as steam engines became more efficient, Britain burned more coal—not less—because cheaper energy unlocked previously uneconomical new uses. The same thing is replaying on AI bills: measured by the same methodology, blended inference costs dropped about 67% in one year (from $18.40/million tokens to $6.07), yet enterprise spending on large models doubled.
The main driver isn't "more chatting"—it's that the workflow itself changed shape. A single Q&A calls the model once; an agentic workflow—plan, execute, observe, adjust, iterate—may call the model 10 to 20 times for the same task. Reasoning models make this multiplier effect even more extreme: OpenAI itself disclosed that o3 uses 83× the compute of GPT-4o on the same problem.
A more hidden layer is "invisible tokens." Reasoning models like o1 and o3 generate large volumes of hidden chain-of-thought tokens before producing an answer; this content doesn't appear in the returned result, but is billed at output token prices. A third-party benchmark (not official vendor disclosure; methodology not independently audited; numbers for order-of-magnitude reference only) found that o1's hidden reasoning tokens average roughly 10× the visible output of GPT-4o; in a benchmark batch of about 31 million prompt tokens, total cost varied by over 50× across models. This explains why an answer with only 500 visible output words might have 2,000 or even 5,000 tokens on the bill.
The set of numbers that best illustrates the scale comes from Sachin Katti—who became Intel CTO in March 2025, then was poached by OpenAI to head Compute Infrastructure 10 months later. In his keynote at the 2025 OCP Global Summit he gave two baselines: Gartner predicts that by 2028, 80% of AI compute will be used for inference, not training; Google's monthly token processing volume grew from 9.7 trillion in 2024 to 1,400 trillion in 2025—a 140× increase in one year. He also provided a token consumption multiplier table converted by agent complexity: a simple chatbot is 1× baseline, adding chain-of-thought reasoning is 10×, and for true multi-step agentic systems, token consumption per query can reach up to 100×. In other words, "inference cost dropped 99% over two and a half years"—the exact words of Altimeter Capital founder Brad Gerstner in the Stanford MS&E 435 classroom— this curve is completely offset, even inverted, as application forms migrate toward agents: unit price dropped two orders of magnitude, but token consumption per task is simultaneously rising several orders of magnitude; net compute demand isn't decreasing—it's exploding.
Unit price decline and total spending increase can both be true simultaneously, and will likely remain true for a long time—cheap enough to use in new scenarios is itself the starting point for demand explosion. Enterprise AI infrastructure spending grew from $11.5B in 2024 to $37.5B in 2026, with inference spend exceeding training spend for the first time and accounting for over half—this is the manifestation of that logic at the financial-statement level.
§04 explains why bills are rising, but doesn't explain something more fundamental: why the price war is being fought by model companies like OpenAI and Anthropic, rather than by Nvidia. A new course opened at Stanford in spring 2026, MS&E 435: Economics of the AI Supercycle, was built around exactly this question—instructor Apoorv Agrawal is a partner at Altimeter Capital who led the largest single investment in the firm's history: entering OpenAI at a $150B valuation. The central question he set for the course is a single sentence: "In the AI supercycle, which layer will value accrue to?"
The answer falls out of a three-layer stack gross margin table. Apoorv splits the AI industry into semiconductor (Nvidia / Broadcom / HBM), infrastructure (Azure / AWS / GCP / CoreWeave), and application (OpenAI / Anthropic / Cursor etc.) layers, tracking annualized revenue and gross margin year by year:
| Layer | 2024 Annualized | 2026 Annualized | Gross Margin | Dominant Players |
|---|---|---|---|---|
| Semiconductor | $75B | $300B | 73% | NVIDIA 80% + Broadcom ASIC |
| Infrastructure | $10B | $75B | 55% | Azure/AWS/GCP + CoreWeave |
| Application | $5B | $60B | 33% | OpenAI+Anthropic 75% |
In two years the entire ecosystem grew from ~$90B to ~$435B, nearly 5×, but its shape barely changed: the semiconductor layer's absolute gross margin is $225B, infrastructure $40B, application $20B—the semiconductor layer alone captures 79% of the entire AI ecosystem's gross margin (87% two years ago, declining only 4 percentage points per year). At this rate, it would take far more than a decade for the application layer to catch up to the "application layer takes most of the gross margin" distribution seen in the cloud computing ecosystem. These numbers come from Apoorv's April 2026 retrospective, where he wrote what became the course's most-cited conclusion: "The semiconductor layer is a single-player game, the application layer is a two-player game, the infrastructure layer is the only fully competitive layer—the most profitable strategy in the AI era is still selling shovels."
Whichever layer is scarcest has pricing power, and whichever layer has pricing power captures most of the gross margin.
This principle has a more formal version—the Bottleneck Calculator from the course materials, whose modeling logic borrows Liebig's Law of the Minimum from ecology: effective supply growth rate is determined by the slowest-growing of chips, memory, energy, and networking; the larger the gap, the stronger the pricing power of the scarce layer. As of mid-2026, the annual growth rates of these factors are roughly energy 15%, chips 45%, interconnect 35%, while demand growth is 310%—the biggest supply-demand gap isn't in chips but in energy, which is why data center siting is shifting from "near users" to "near power." But regardless of which link is the specific bottleneck, Nvidia currently stands at the scarcest position—this is the fundamental reason it can command a 73% gross margin, independent of how good its chips are or whether alternatives exist—pure scarcity pricing.
This chain explains why OpenAI and Anthropic are caught in a "sell more, lose more" situation: the application layer they inhabit must pay a premium to the scarcest chip layer (the top 5 hyperscalers' combined capital expenditure was ~$443B in 2025, up 73% YoY, projected to exceed $600B in 2026, with ~75% going to AI infrastructure), while simultaneously fighting a price war for market share within the most competitive application layer— they have pricing power on neither side, squeezed from both ends. SemiAnalysis founder Dylan Patel in his Class #5 interview looked at the demand side alongside this situation: frontier models are the only product people are willing to pay almost unlimited amounts for—his own company's AI spending went from tens of thousands of dollars per year to $7M annualized in one year—but the supply side is simultaneously bottlenecked by three physical constraints: memory (HBM capacity), advanced process nodes (TSMC 3nm/2nm queues), and lithography equipment (ASML EUV shipment pace). His conclusion is direct: companies that can't generate value through AI will face economic reckoning—this statement can be seen as the supply-side mirror of §07's "outcome pricing" logic: whether the people buying shovels make money ultimately determines whether the people selling shovels can keep raising prices.
Both mobile internet and cloud computing took roughly fifteen years to flip from "hardware layer captures most gross margin" to "application layer captures most gross margin"— after Qualcomm/ARM came iOS/Android, and only then did Uber, Instagram, and TikTok take the lion's share. Looking back at his prediction from two years ago, Apoorv said in 2026 that he still believes this inversion will happen, "but much more slowly than I imagined in 2024." This is also why the outcome pricing and RWA-ification discussed in earlier sections read more like the application layer searching for an escape route—they can't control the chip layer's pricing power, so all they can work with is the billing method and financial engineering itself.
The headline number "$X per million tokens" is now just the most visible cell in a multi-dimensional price list. The same model, the same text, will cost different amounts depending on four variables:
The result is that any comparison table—including the one in §03—is just a snapshot. What really determines the bill is how these four variables combine for your specific workload, and two different vendors will almost never give directly comparable answers for that combination.
Per-token billing has a structural flaw: an agent that spends 5,000 tokens carefully researching and resolves the customer's issue in one shot, versus another agent that spends 20,000 tokens going in circles and solves nothing—the former gets a cheaper bill. The pricing method rewards verbosity and penalizes efficiency, with no necessary connection to "how much was this interaction actually worth." Fortune called this phenomenon the end of "tokenmaxxing" in a May 2026 article: when consumption itself becomes the KPI, consumption rises, but the work doesn't necessarily get done—a textbook Goodhart's Law failure.
Several companies have already switched their billing unit. Salesforce Agentforce launched in 2024 with flat $2/conversation billing, and the criticism was direct: whether the AI resolved the issue, failed, or escalated to a human agent, the meter kept running. After introducing Flex Credits (per-action billing, ~$0.10/standard action) in May 2025, the 2026 pricing page simultaneously lists per-conversation, per Flex Credit, and per-seat ($125–$550/user/month) systems, letting customers choose based on their usage pattern. Zendesk and HubSpot went further— Zendesk bills per "resolved ticket" (~$2 each), and HubSpot Breeze changed from "$1 per conversation processed" to "$0.50 per resolved conversation," with the official wording: "customers only pay when the agent completes its assigned task." Sierra states this logic most directly: charge only on task completion, aligning vendor and customer incentives.
Databricks CEO Ali Ghodsi, in a fireside chat at MS&E 435 Class #4, gave this shift an even grander name: "Service as Software." His argument is that traditional SaaS sells tools, with the price ceiling determined by "how much labor this tool can save," typically capping at around 33% application-layer gross margin; whereas an outcome-billed AI agent sells the completed result of the work itself, and the comparison object is no longer a similar software subscription but the consulting fees, legal fees, and customer service labor costs it replaces—a much higher ceiling. Ghodsi in the same conversation also gave a symmetric warning: companies that raised billions of dollars but have almost no revenue are "obvious bubbles"—outcome pricing opens up the price ceiling, but doesn't automatically generate revenue for every player.
This model comes at a cost to vendors—outcome pricing means the cost of failed tasks and retries must be absorbed by the vendor, rather than being passed along via the "bill anyway if you called it" safety net of per-token billing. The rule of thumb circulating in the industry is: set the outcome price at 2.5 to 3.5× the "blended unit cost per token at the expected success rate," using this premium to cover failed retries and the bad 90th-percentile months. The unspoken implication: outcome pricing is essentially transferring some of the model's uncertainty risk from the customer's bill back onto the vendor's own income statement.
The real difficulty isn't in the algorithm or billing system, but in who defines "outcome." Customers don't want to face an opaque set of acceptance criteria to determine whether this call counts—the friction currently has no standard answer, and it's the biggest unresolved question on the path from concept to large-scale deployment for this model, harder than the technical implementation itself.
Outcome pricing solves "what to bill for," but hasn't touched "how to settle"—monthly bills, manual approvals, credit card authorizations: this process was designed for humans, and agent-to-agent high-frequency small-value calls can't use it at all. In May 2025, Coinbase released the x402 protocol to fill exactly this gap: borrowing the HTTP 402 (Payment Required) status code, the server responds to an agent's request with a "payment required" response, the agent signs a stablecoin transaction (USDC etc.) and attaches it to the retry request, the server verifies and lets it through—the entire process completes in seconds, with no human intervention, no credit card, no subscription account.
The scale is already significant. Reportedly, x402 has processed over $50B in stablecoin transactions, covering 200 million payments, most with individual amounts under $0.50— exactly the scale that traditional credit card networks' fixed fee structure can't handle: the fee itself costs more than the payment. Cloudflare and Coinbase are the protocol's co-governors, Stripe and AWS have referenced the standard in their documentation, and AI-native platforms like Browserbase and Exa already support it natively. Combined with Google's AP2 (authorization layer) and MCP (tool discovery layer), it forms a payment stack for the agent economy: MCP lets agents find callable tools, x402 enables per-call payment for that invocation, and AP2 confirms whether the money is truly authorized by this agent's human principal.
Update (2026-09): This infrastructure is becoming institutionalized, but "$50 billion" type flow numbers don't withstand scrutiny. On July 14, 2026, the x402 Foundation was officially established under the Linux Foundation, with 40 member institutions including Visa, Mastercard, Ripple, Stripe, Google, and AWS—governance is indeed becoming more mainstream. But according to on-chain analysis by Chainalysis and Artemis Analytics, over 95% of the widely cited "200 million transactions" are protocol self-test machine traffic, not real buyer-seller transactions; settlement volume itself has also dropped significantly from its Q4 2025 peak— that peak was largely inflated by "paid minting" speculative activities like PING, and after the tide went out, volume fell by more than half over the past three months. The one arguably positive signal is that transaction structure is changing: the share of transactions above $1 grew from 49% in early 2025 to 95% in early 2026, suggesting that what remains is indeed closer to real commercial settlement—just at a much smaller scale than the headline numbers imply.
This means the next layer of the "token economy" isn't about pricing tokens more precisely, but about turning each API call itself into an instantly settleable cash flow—no longer a monthly bill, but pay-per-call. This infrastructure is still early, with scale data coming from protocol parties and media reports, not yet third-party audited, but it has already solved the hardest engineering problem: making the friction of machine-to-machine payments low enough to be negligible.
The token economy is colliding with another long-existing narrative: RWA (real-world asset tokenization). This line originally involved moving treasuries, private credit, and real estate on-chain—as of mid-2026, the total on-chain RWA value excluding stablecoins is between ~$27B and ~$36B, with the largest single category being tokenized US Treasuries (~$14.8B), and BlackRock's BUIDL and Ondo Finance are the two largest issuers in this track. In 2026, AI compute and inference rights started becoming a new asset class in this category.
The specific approach has two layers. The first is tokenizing GPUs themselves: Aethir packages over 440,000 GPU containers as divisible, yield-bearing tokenized assets; RWAi lets retail investors directly buy fractional shares of GPU racks running open-weight models like DeepSeek and Llama, earning dividends based on the actual inference revenue generated by the compute. The second layer goes further—using GPUs or future inference revenue as loan collateral: USD.AI lets stablecoin depositors lend to data center operators, with loans collateralized by GPUs and other AI infrastructure assets; the protocol claims $225M in executed loans and over $1.2B in approved loan capacity.
This boundary continues to blur. Projects like DIEM directly make "prepaid inference credits" into tradeable tokens—1 token corresponds to $1/day of inference capacity, and unused credits can be sold on the secondary market, essentially turning "how many tokens a model will compute for you in the future" into a priced, holdable, transferable instrument. Derivatives markets are following: CME and ICE have announced plans to launch GPU rental futures, and a March 2026 academic paper proposed a "Standard Inference Token" framework, analogizing AI tokens to commodities like electricity and carbon allowances, designing complete futures contract specifications and margin mechanisms—this is still at the proposal stage, some distance from actual listing, but the direction is clear: decompose "intelligence" into standardized contracts that can be priced, hedged, and collateralized.
Tokenization solves "how to quickly liquidate these assets," not "how real the revenue behind these assets actually is."
This narrative sounds like it found a bigger capital market for the token economy, but it's replicating exactly the kind of risk that has been repeatedly warned about over the past two years. The Bank for International Settlements (BIS), in its June 2026 annual report, listed "AI capital expenditure bubble burst" and "opaque circular financing" as the two biggest threats to the global financial system, explicitly warning that related assets "may be pledged multiple times." CoreWeave's $8.5B investment-grade loan collateralized by GPUs and customer contracts has its repayment window coinciding with a period of declining GPU secondary-market values; some analysts compare this structure—where vendors lend to their own customers to create the illusion of revenue growth—to the "vendor financing" of Nortel and Lucent around 2000— that bubble ended with loans going bad, resulting in one of the largest bankruptcies in Canadian corporate history. Not everyone agrees with this analogy: supporters point out that AI data center tenants are primarily cash-rich investment-grade hyperscalers, and underlying compute demand is real and growing—two things the 2008 mortgage bubble lacked.
Five threads are converging toward a single direction: the tokenizer efficiency race (who can express the same semantics in fewer tokens), outcome pricing (transferring uncertainty risk from the customer's bill back to the vendor's income statement), agent-to-agent micropayments (driving transaction costs themselves toward zero), compute and inference rights RWA-ification (turning these assets into quickly liquid, leveragable collateral), and the structural squeeze of the three-layer stack itself (the application layer lacks the chip layer's scarcity pricing power and must fight a price war within its own layer). They all point to the same thing—the price war ultimately isn't about sticker price, but about the ratio of "completed work per dollar spent," and no one is currently willing to publicly disclose this ratio.
But this narrative doesn't answer a more basic question: the foundational model layer's own income statement hasn't turned positive yet. Analysts estimate OpenAI's Q1 2026 adjusted operating margin at approximately -122%—for every dollar of revenue, it loses $1.22; this isn't an operational mistake, it's the direct implication of §05's gross margin distribution: the semiconductor layer takes 79% of gross margin, the application layer gets only 33%, and model companies are squeezed from both ends. Outcome pricing, agent micropayments, and RWA-ification can all optimize "how money is calculated, paid, and flows," but they can't magically fill the gap of "how much it costs to produce a dollar of intelligence," let alone move pricing power from the chip layer to the application layer. Tokenizing, collateralizing, and leveraging compute and inference rights just moves this gap from one side of the balance sheet to the other—as long as open-weight models can be hosted at near-zero cost and Nvidia still stands at the scarcest layer, the price war has a floor that keeps descending, and this is a problem no billing model or financial engineering can circumvent.
Note: All figures in this article are snapshots from public reports and third-party benchmarks, primarily from mid-August 2026; paragraphs marked "Update (2026-09)" have been synchronized to mid-September 2026; some figures are analyst estimates or vendor self-reports, not independently audited; pricing changes frequently, readers are advised to verify against vendor official pricing pages before citing.
STORYLINE
Each of this article's ten sections only excavated one segment of one thread—pricing mechanisms, three-layer stack gross margin distribution, outcome pricing, RWA-ification— and each has a deeper standalone treatment in DeepDive. The route below strings them together; this article is the entry point, not the endpoint: starting from "how tokens are priced," it extends one segment in each of three directions—supply side (inference costs, energy bills), demand side (agent economy, business model transformation), and academic frameworks (the original source of the three-layer stack).
Note: #04 Capability Sink and Compute Concentration and #09 MS&E 435 Study Guide are also members of other storylines on this site (SL·01 Power Contest, SL·07 AI Economics in the University)—an article can belong to multiple storylines simultaneously; the selection here represents the angle most relevant to the token pricing argument chain.