Anthropic agreed to pay Nscale $45B over six years for compute capacity at Nscale's West Virginia data center — the fourth mega compute deal Anthropic has signed in as many months, following six-year commitments worth $50B with Fluidstack, $10B with Volta, and roughly $45B with SpaceX. The West Virginia site runs on Nvidia's Vera Rubin systems and won't come online until late 2027, meaning Anthropic is pre-buying capacity years before it can use it — evidence that compute scarcity, not model quality, is now the binding constraint on how fast frontier labs can grow. (TechCrunch)
Private Companies
Nvidia agreed to buy Hugging Face, the open-source model and dataset repository, for $12.9B — more than 80 times Hugging Face's reported $150M in annualized revenue — according to The Information. The deal, still unconfirmed by either company, follows Hugging Face's rejection of a $500M Nvidia investment last year that would have valued it at $7B. It gives Nvidia a foothold in the open-source ecosystem that rival labs building in-house chips, including Anthropic and OpenAI, increasingly depend on. (The Information)
Celera Semiconductor closed a $30M Series B funded entirely by Maverick Silicon, its largest investor, to expand its AI-driven analog chip design platform. The Santa Clara startup's Nesto technology compiles full-custom and standard analog ICs — the voltage-regulation circuitry that's become a bottleneck as individual AI accelerators draw more than 1kW — far faster than traditional design flows. The round follows Celera's acquisition of a Portugal-based design-automation team as it scales engineering headcount. (PR Newswire)
Public Markets
Nvidia$209.66▼ 1.6%Mkt Cap: $5.0T
Nvidia reported $96.2B in quarterly revenue, up 106% year-over-year, with Data Center revenue of $89B and Q3 guidance of $108B versus a $104.2B estimate — a beat that sent shares up nearly 5% after hours, snapping a seven-session losing streak, with the outlook assuming zero China data-center revenue. (CNBC)
Emerging
Chip design: A new paper details TFA, a synthesizable INT8 accelerator for transformer inference that handles both prompt processing and token generation on one datapath, reporting roughly 20x speedup and a projected 1,000x energy reduction per token versus a 22-thread CPU baseline, in just 2.73mm² of logic. It's a small-team, open design rather than a hyperscaler product — evidence that competitive inference silicon no longer requires Nvidia-scale R&D budgets to prototype. (arXiv)
Memory computing: Samsung detailed LPDDR5X-PIM at Hot Chips, a drop-in memory chip that places compute logic inside standard LPDDR5X DRAM using the same JEDEC package, delivering 8x the bandwidth and roughly 3x the AI token throughput of ordinary LPDDR5X — including a jump from 27 to 81 tokens per second running Llama 3.1 — without a system redesign. Unlike HBM-based processing-in-memory efforts that need new packaging, this targets server, mobile, and client parts alike. (Tom's Hardware)