torygreen

vip
Active for: 3.1y
Peak Tier 0
No content yet
GPU-hours might be turning into something you hedge.
The CFTC is already exploring compute derivatives, and Liquid Compute just raised $15M to build a regulated market for GPU futures, options + swaps.
Makes sense if AI compute really grows from ~$360B in 2025 to ~$2.3T by 2030.
The weird part is standardization.
An H100-hour in one region isn’t the same product as a B200-hour somewhere else with different networking, power costs and cluster size.
So before compute gets a real futures market, someone has to invent the benchmark contract everyone agrees to trade around.
Once that exists, labs c
post-image
H1000.00%
B200-1.25%
.@OpenAI spent $6.5B building the device team and now it may be spending another $300M+ on the part that decides what the model gets to see in the first place (Glass imaging).
Glass is working below the camera app directly on raw sensor data. Its proposed mobile system uses a sensor ~2x larger than today’s biggest smartphone sensors and claims up to 10x more light capture.
For an always-on assistant, that matters a lot.
Anything the sensor misses is gone. The model has to guess around it later, which means more uncertainty, more compute, and worse understanding of whatever is happening around
post-image
Query counts are about to become a pretty bad proxy for AI demand.
Over 8 weeks, one Claude Code user typed 1,138 prompts. Those turned into ~14k model calls and 3.2B processed tokens. Estimated energy use came out to ~150 Wh per typed prompt, around 625x a median Gemini text prompt.
Most of the compute happens after the user stops typing. The agent rereads context, calls tools, retries, branches, and keeps working in the background.
So one human request is starting to behave more like a small batch of jobs than a single query.
For infra, the better unit might end up being agent-hours, not pro
post-image
The market is already routing around the benchmark gap.
7 of OpenRouter’s top 10 models by weekly token volume are Chinese, accounting for ~75% of usage across that top 10.
US labs still have the edge at the frontier, but most workloads don’t need the absolute best model. They need something that’s good enough, cheap enough, and easy to actually ship.
Chinese labs seem to be getting that part right and they're getting better with each release. Stay close to the frontier, push cost down, make deployment easy, and usage can scale way before you ever take the #1 spot.
2 years ago, Google was processing 9.7T AI tokens a month.
Today it is above 3.2 quadrillion. That’s a 330x increase.
There is no reliable global count so Google’s number can be best treated as a disclosed floor.
Even if growth slows to 2x a year, Google alone would reach 51 quadrillion tokens a month by 2030.
And that workload will look very different from someone opening a chatbot today. Long-running agents will hold context, divide jobs into subtasks, call other models, retry failures and verify each other’s work. One instruction could keep dozens of models active for hours.
Imagine a compa
post-image
GOOGL+0.47%
One Anthropic contract is pulling $3.5B of financing and an entire GPU buildout into the present.
Nscale reportedly signed a roughly $45B deal with Anthropic, is seeking $3.5B before an IPO, and has shown investors $103B in projected revenue from signed leases. Nvidia may provide $2B of the raise.
Anthropic’s commitment gives lenders something concrete to underwrite and Nvidia’s capital helps pull its own GPU sales forward. Nscale converts both into capacity before most of the revenue arrives.
The timing mismatch is where this gets interesting.
The leases can run for years, while GPU economic
NVDA+1.23%
  • 3
Data centers are turning into a major credit market.
JLL expects North American permanent debt originations to jump from $80B in 2025 to $175B this year, then $322B by 2028. That’s roughly $722B across the three forecast years.
But the cost of that capital is already splitting the market. Hyperscaler-backed construction can reach ~85% loan tocost at low-200 bps spreads, while projects tied to AI labs and neoclouds can price 200–300 bps wider with only 70–80% leverage.
So two identical megawatts can have very different economics before a single GPU is installed, purely because of who sits behin
post-image
JLL-0.13%
30x more agent work from the same megawatt.
That’s what @nvidia says Vera Rubin NVL72 delivered versus GB300 NVL72 on DeepSeek V4 Pro, with up to 35x lower cost per million tokens.
The workload is what makes this more than a normal throughput benchmark. AgentX keeps the ugly parts of real agent usage intact: growing context, repeated tool calls, and sub-agents branching off mid-task.
If those early numbers hold up under SemiAnalysis review, the infra implication is pretty big. The same contracted megawatt can suddenly support far more billable agent work without waiting for another substation
post-image
NVDA+1.23%
DEEPSEEK-8.08%
  • 1
Memory is becoming a much bigger part of the AI scaling problem.
J.P. Morgan estimates DRAM prices could be up 400%+ from early 2024 through end-2026, while SK hynix expects the shortage to persist through 2030.
That pressure is already shaping both sides of the stack. Models are leaning harder on MoE, quantization and KV-cache efficiency, while Nvidia’s Rubin pushes to 288GB of HBM4 and 22 TB/s of bandwidth per GPU.
We spent years treating FLOPs as the main scaling variable. For long-context agents and inference-heavy workloads, squeezing more intelligence out of every GB may matter just as m
SK Hynix+6.41%
SKHY+2.61%
SKHYV-0.98%
NVDA+1.23%
McKinsey’s latest State of AI survey has a stat that should make a lot of SaaS companies uncomfortable:
32% of companies say they’ve already skipped buying at least one software product or feature because coding agents could build it internally.
That doesn’t mean companies stop buying software. It means the “feature” layer gets a lot harder to monetize.
If an internal team can recreate a narrow workflow in a day or two, vendors need to justify the sub model with something agents can’t cheaply reproduce: distribution, embedded data, compliance, integrations, network effects, or owning the syste
post-image
Inference speed is starting to matter for agents in a much bigger way than just making responses feel faster.
Cerebras’ new CS-4 claims up to 30x faster inference than GPU systems, with 4,400+ tok/s on GPT-OSS-120B.
For agents, that speed becomes extra reasoning budget.
In the same 30-sec window, a faster system can explore more paths, retry failures, run more verification, or use a larger model without making the user wait longer.
So tokens/sec is slowly becoming more than a latency metric.
It can directly change how much useful work an agent gets done before the clock runs out.
There’s a point where vertical AI companies stop looking like API customers and start looking like model companies.
Harvey went from roughly 1T -> 14.5T tokens/month in six months. Now it’s post-training its own legal model on top of Kimi K3, using domain-specific tasks, expert feedback and proprietary workflow data.
Thomson Reuters is moving in the same direction with its own legal model.
That feels like a pretty natural evolution.
At low vol, frontier APIs are the obvious choice. At 10T+ tokens/month, model selection starts becoming a gross-margin decision. And if you already own the workflo
post-image
KIMI-0.50%
Coding may have just been the first agent market, not the biggest one.
Since February, weekly active Codex users grew :
> 108x in legal,
> 41x in sales,
> 41x in recruiting
> 26x in marketing.
Engineering only grew 5x.
And by June, agentic usage already accounted for 64% of combined Codex + ChatGPT output tokens across enterprise customers.
Software was probably the easiest place to start because the work is structured, tool-heavy and relatively easy to verify.
Now that agents are getting better at navigating messy workflows, the growth is showing up in functions where the underlying labor
LLMs had the internet waiting for them. Embodied AI doesn’t.
Robots need huge amounts of real-world interaction data, and ACE Robotics estimates the whole sector has only ~100K useful hours today.
China is already trying to manufacture that dataset: training centers buy robots, teleoperate them through real tasks, collect the trajectories, then feed that data back to robot makers.
So some early humanoid deployments are doing more than testing hardware. They’re generating the training data the next generation of physical AI will need.
That makes the race for embodied AI look a lot more like a r
post-image
Pretty insane how fast capability is compounding without the hardware budget moving nearly as much.
> Qwen3.5-27B scored 35 on Artificial Analysis in February.
> Qwen3.6-27B moved that to 38 in April.
> Qwen3.8-27B is now at 52, while staying in basically the same 27B parameter envelope.
And its agentic performance is now sitting around models that are dramatically larger and far more expensive to serve.
That feels more important to me than another leaderboard reshuffle.
For years, better AI mostly meant scaling the hardware underneath it. Now capability is jumping much faster than the deploym
AI networking is starting to look like its own mini capex cycle.
@Cisco closed FY2026 with $9.3B in hyperscaler AI infrastructure orders and is guiding AI infrastructure revenue from ~$4B to $7.5B next year. Arista just grew 37.7% YoY to $3.04B
I think the interesting shift is where the marginal AI dollar is starting to go.
The first phase of the buildout was dominated by getting as many accelerators online as possible. Once clusters get large enough, keeping those GPUs fed with data and talking to each other efficiently becomes a much bigger part of the economics.
An expensive GPU sitting id
CSCO-0.61%
  • 1