Subject Line: Cheap AI Is Bullish for NVDA
Preheader: Falling token prices historically expand total GPU demand. August 26 is six days away and the options market is pricing an 8.3% move.
Meta Description: Falling AI inference costs look like a threat to Nvidia. The data says the opposite: cheaper tokens drive more compute demand, and NVDA captures the hardware bill every time.
The Market Gets the Logic Backwards
Every time AI gets cheaper, investors panic about Nvidia. It happened in January 2025 when DeepSeek dropped a low-cost model and NVDA shed billions in market cap in a single session. It happened again in June 2026 as B200 GPU rental prices slid from a peak of $6.11 per hour to $4.22. The logic seems airtight: lower prices, lower revenue per unit, lower earnings for the company selling the chips.
That logic has been wrong each time it has been applied. When DeepSeek released its low-cost model in January 2025, Nvidia’s stock fell sharply. The market’s logic was straightforward: cheaper inference means less GPU demand. Then Microsoft CEO Satya Nadella posted “Jevons paradox strikes again,” arguing that as AI gets more efficient and accessible, use skyrockets, turning it into a commodity the market cannot get enough of. Nvidia stock recovered. The market had misread the economics.
Eighteen months later, the misread is repeating. The correct frame is not price per token. It is total token volume, and total volume is doing something that the per-unit price decline completely obscures.
This is not about whether Nvidia beats consensus on August 26. It is about whether the market correctly understands what falling AI prices actually do to infrastructure demand. The answer changes how you size the trade, structure the options position, and think about the Q3 guide that consensus is already pricing at $103.1 billion.
The Data Section: Where Nvidia Actually Stands
NVDA delivered $82 billion in Q1 revenue, up 85% year over year, while authorizing $80 billion in new buybacks at a 75% gross margin. That gross margin figure is the one that tends to get buried. Nvidia’s 75% gross margin compares to Microsoft’s 67.9%, Alphabet’s 59.7%, and Amazon’s 50.3%. The company is not just growing faster than its customers. It is doing so at a margin structure that is structurally superior to the people writing it checks.
Microsoft, Alphabet, Amazon, and Meta spent a combined $166 billion on capital expenditures in the June quarter, up 87% from a year ago and 27% from the March quarter alone. Over the ten quarters since the start of 2024, combined hyperscaler capex has risen 272%. Despite zero China H20 shipments versus $5 billion a year ago, management still guided Q2 revenue to $91 billion with China stripped out entirely.
NVIDIA reports earnings on August 26. Consensus revenue sits at $91.85 billion, with EPS of $2.08, against the company’s own guide midpoint of $91 billion. The number that will move the stock is the Q3 forward guide against a consensus of $103.1 billion. A guide at or above consensus signals the capex curve still compounds. A guide below marks the first inflection in the cycle. Everything else on August 26 is noise.
For full fiscal year 2026, Nvidia revenues hit $215.9 billion, up 65% from the prior year, with GAAP operating income of $130.4 billion and net income of $120.1 billion. Fiscal year 2026 delivered record-breaking revenue of $215.94 billion, representing a 65% year-over-year increase that validates the company’s dominant position in the AI infrastructure market.
Strategic Interpretation: What Cheaper AI Actually Does to Nvidia’s Business
The central reframing that most investors are missing is this: Nvidia does not sell tokens. It sells the hardware that produces tokens. When tokens get cheaper, Nvidia’s customers buy more of them. To produce more of them, they need more hardware. The company capturing that hardware order is Nvidia.
Inference cost for comparable capability tiers has fallen from about $0.06 per 1,000 tokens in early 2025 to about $0.006 by mid-2026, a roughly 10x decline over the period. That is a compression that sounds bearish in isolation. It is not. Lower serving cost is not bearish for infrastructure by default. It often produces the opposite effect. Cheaper serving expands the set of economically viable products, which increases demand for GPUs, networking, storage, and memory. That feedback loop is why cost compression at the model layer can still support capital spending across the semiconductor and cloud stack.
The academic term for this dynamic is Jevons Paradox. It is the idea that efficiency gains often lead to greater overall consumption rather than less. Jevons Paradox is the counterintuitive observation that improvements in fuel efficiency tend to increase, rather than decrease, overall fuel use. Applied to AI compute, the mechanism is already visible in the spending data. Inference costs fell roughly 1,000-fold over the period measured. Demand rose roughly 10,000-fold. The lower unit price did not reduce spend. It authorized more consumption than it saved.
The demand elasticity on token consumption confirms this is structural, not cyclical. Simulation data yields a super-elastic response: a 1% decrease in price generates a 1.42% increase in volume. Because the demand elasticity exceeds 1, the price drop leads to an increase in total industry revenue. This confirms that the demand for digital intelligence is currently unsatiated; efficiency gains are fully offset by increased consumption intensity.
The companies that benefit from Jevons Paradox are not the ones selling tokens. They are the ones selling the picks and shovels: GPUs, HBM, substrates, and electrical power. This is why Nvidia’s margins are expanding while Anthropic’s are deeply negative, and why memory manufacturers have pricing power despite every efficiency paper published. The efficiency gains lower the price of inference, which creates demand, which requires more hardware, which increases hardware prices.
Sector Implications: The Inference Shift Is Already Priced Into Hardware, Not Software
The training-to-inference ratio has inverted. In 2026, inference accounts for approximately two-thirds of all AI compute demand, up from roughly one-third in 2023. That shift has a direct consequence for where the hardware dollars flow. 80% of AI GPU spend is now inference. Training was a one-time event per model. Inference is recurring, scales with users, and grows as AI embeds into more products.
Nvidia’s CFO stated that Blackwell has been the fastest product ramp in company history, and its demand is propelled by its lowest token generation cost at inference. CEO Jensen Huang said the company is growing its inference share rapidly, and that Vera Rubin is already off to a tremendous start, positioned to be far more successful than Grace Blackwell.
Rubin is designed for inference-heavy workloads essential to agentic AI, offering 10x higher throughput per watt compared to Blackwell. Nvidia’s backlog exceeding $500 billion for Blackwell and Rubin provides clear revenue visibility into 2027. That backlog exists because the customers ordering this hardware have already done the Jevons math. They expect cheaper tokens to produce higher total token volume. To serve that volume, they need the next generation of silicon, and they are pre-committing to it 6 to 12 months before delivery.
Enterprise AI bills rose around 320% in the 2024-2025 window even as per-token prices kept falling. Deloitte’s 2026 TMT predictions argue that AI’s next phase demands more computational power, not less, with inference on track to be about two-thirds of all compute in 2026. That is the Jevons rebound, working at the bill level rather than the chip-hour level.
The sector read is straightforward. Falling token prices are a headwind for AI application companies whose revenue is priced per token. They are a tailwind for infrastructure companies whose revenue is priced per GPU shipped. Nvidia sits firmly in the second category. Its customers absorb the margin compression from cheaper inference. Nvidia captures the volume rebound in hardware demand.
Options Market Analysis: What the August 26 Earnings Setup Looks Like
Implied volatility ramps hard into NVDA earnings in what traders call the IV rush, and collapses immediately after in the IV crush. NVIDIA’s typical earnings move averages plus or minus 8.3%, with a more recent tendency toward plus or minus 5.4%.
Implied volatility ramps hard into NVDA earnings. Combined with a sub-25% beat rate against the implied move, holding long options through NVDA earnings has historically been difficult: traders pay peak IV, and the actual move usually does not clear it. NVDA has cleared its implied move only about a quarter of the time, with a plus or minus 20.1% 95th-percentile figure marking the realistic worst case for any position.
Implied volatility is elevated into the event, which means post-earnings IV crush can significantly impact option prices. For long premium strategies such as call spreads and put spreads, collapsing implied volatility can become a headwind if the stock does not move enough. For short premium strategies such as iron condors, collapsing implied volatility can become a tailwind if the stock remains inside the expected range.
Nvidia earnings are no longer just another company update. The company has become one of the central pillars of the AI infrastructure trade, meaning its results often influence sentiment across semiconductors, mega-cap technology stocks, and the broader Nasdaq. That systemic weight means a strong Q3 guide does not just lift NVDA. It reprices the entire AI infrastructure complex: SMCI, VRT, EQIX, and MU all move on the read-through.
Structured Trade Framework
Bull Case: The Guide Clears $103B
If you believe the hyperscaler capex cycle has another leg, and the Q3 guide meets or exceeds the $103.1 billion consensus, a defined-risk bull structure captures the upside while managing the IV crush risk that punishes naked long calls. A call spread targeting the expected move range, for example buying the at-the-money call and selling a strike approximately 8-10% above current price on the August 29 expiry, captures directional upside while limiting the damage from IV compression post-announcement. The defined max loss is the premium paid. The defined max gain is the spread width minus premium. This structure works when the stock moves decisively in one direction, which the Q3 guide tends to produce.
Bear Case: The China Gap Widens or Gross Margin Compresses
The primary bear case is not falling token prices. It is a Q3 guide that explicitly adjusts lower for China export controls or signals that the 75% gross margin is compressing as Vera Rubin production costs ramp. Nvidia shipped zero H20 units to China in Q1 against $4.6 billion in the year-ago quarter, and Q2 guidance excludes China data center compute entirely. If management signals incremental restriction risk in the forward guide, the read-through is negative. A defined-risk bear structure using a put spread on the August 29 expiry, buying slightly out-of-the-money puts and selling a lower strike at the expected move boundary, limits cost to premium paid while targeting downside if the guide disappoints. For traders who believe the China headwind is underpriced in consensus, this structure costs less than a naked put while still participating in a 5-8% downside move.
Neutral Case: The Move Stays Inside the Implied Range
Given the historical data showing NVDA clears its implied move only about one quarter of the time, a short volatility structure that profits from IV crush without requiring directional accuracy is coherent. An iron condor built outside the plus or minus 8.3% average expected move range on the August 29 expiry collects premium on both the call and put side. The maximum profit is the net premium received. The risk is a move beyond the short strikes. For traders who have no strong directional conviction but expect IV to collapse post-announcement regardless of direction, this is the structure to consider. Width of the condor wings should reflect the 20% 95th-percentile tail risk referenced above.
Risk Analysis
Three risks are worth sizing before any position is built.
First, the China variable. Key risks include high valuation multiples that assume continued perfect execution, intensifying competition from AMD and custom AI chips, geopolitical export restrictions limiting China market access, and potential cyclicality in AI capital expenditure. China was a $4.6 billion quarterly line item a year ago. It is zero today. If export restrictions tighten further or the political environment shifts, the Q4 guide would have to absorb another step-down that consensus has not modeled.
Second, the gross margin trajectory. Gross margin has been the quiet variable while revenue absorbs all the attention. Vera Rubin’s production ramp carries initial yield risk. The VR200 NVL72 rack draws roughly 190-230 kilowatts, up from 120-130 kilowatts for Blackwell. Higher power density raises the infrastructure cost for customers and increases the complexity of Nvidia’s own manufacturing ramp. If Rubin’s early yield is lower than expected, gross margin could compress even as revenue grows, which tends to produce an outsized negative stock reaction.
Third, the custom silicon encroachment is real even if it is not yet decisive. Custom silicon now represents 20.9% of the AI chip market in 2025 and is expected to expand to 27.8% by 2026, posing a long-term threat to Nvidia’s market share. This is not a Q2 story. It is a 2027-2028 story. But options traders pricing the stock at a forward multiple above 30x are implicitly assuming that custom silicon does not accelerate faster than Nvidia’s architecture roadmap can absorb.
Forward Outlook: The Inference Economy Is Just Getting Started
If AI becomes embedded into office suites, development tools, browsers, customer service systems, security consoles, and industrial monitoring platforms, the demand curve shifts from building the model to running the model everywhere, all the time, cheaply enough that users do not notice the meter running. That is why Nvidia’s roadmap is so focused on platform throughput, networking, and cost per token.
The biggest shift is that inference has moved from lab expense to product gross-margin line. VentureBeat’s April 30, 2026 article framed the change as a move from a small number of scheduled model jobs toward thousands of concurrent inference workloads, with agentic use accelerating token demand.
Goldman Sachs projects agentic AI will multiply consumer-side token consumption 24-fold by 2030, with combined demand reaching roughly 120 quadrillion tokens processed monthly. Every one of those tokens requires compute. Compute requires hardware. Hardware requires Nvidia’s silicon, its networking stack, and increasingly its software layer. TIKR’s mid-case, realized in early 2031, values NVDA at around $505, a total return of roughly 162% over approximately 4.6 years, driven by continued hyperscaler capital spending and the Vera Rubin ramp reaching enterprise, sovereign, and AI-cloud customers, supporting mid-case revenue growth of around 23%, with a data center operating leverage holding a net income margin of around 55%.
The Jevons dynamic does not eliminate risk. It reframes which risk matters. The existential question for NVDA is not whether tokens get cheaper. It is whether demand grows faster than prices fall, and whether Nvidia retains the architectural position to capture that demand with hardware that cannot be substituted in the near term. The last 18 months of data answer both questions in Nvidia’s favor. August 26 is where the market gets another data point.
Action Checklist
- Earnings date confirmed: August 26, 2026, after the close.
- Consensus to watch: Q2 revenue $91.85 billion, EPS $2.08. Q3 forward guide consensus $103.1 billion is the number that moves the stock.
- Key data center metric: Data center revenue was $75.2 billion of $81.6 billion in Q1, roughly 92% of total. Watch for any share shift.
- Gross margin signal: 75% in Q1. Any compression below 73% on the Vera Rubin ramp would be a negative read-through.
- China disclosure: Zero H20 revenue in Q1. Any change to the China assumption in the forward guide, positive or negative, will move the stock independently of the headline beat.
- Options structure for bulls: August 29 call spread, buy at-the-money, sell 8-10% above current price. Defined max loss equals premium paid.
- Options structure for bears: August 29 put spread, buy slightly out-of-the-money, sell at the expected move lower boundary. Defined max loss equals premium paid.
- Options structure for neutral traders: August 29 iron condor built outside the historical plus or minus 8.3% expected move. Profit from IV crush if the stock stays inside the range. Size wing width to accommodate the 20% 95th-percentile tail risk.
- The Jevons check: Before building any bear position on falling token prices alone, confirm whether the Q3 guide reflects volume acceleration, not just price per unit. Volume is what matters for Nvidia’s revenue line.
- Risk budget: Defined-risk structures only into a binary event. NVDA has moved as much as 20% in a single session on earnings. Uncovered short options carry risk that exceeds most traders’ intent. Use spreads.
