Anthropic agreed to pay Nscale $45B over six years for compute capacity at Nscale's West Virginia data center — the fourth mega compute deal Anthropic has signed in as many months, following six-year commitments worth $50B with Fluidstack, $10B with Volta, and roughly $45B with SpaceX. The West Virginia site runs on Nvidia's Vera Rubin systems and won't come online until late 2027, meaning Anthropic is pre-buying capacity years before it can use it — evidence that compute scarcity, not model quality, is now the binding constraint on how fast frontier labs can grow. (TechCrunch)
Private Companies
Hugging Faceother
Nvidia agreed to buy Hugging Face, the open-source model and dataset repository, for $12.9B — more than 80 times Hugging Face's reported $150M in annualized revenue — according to The Information. The deal, still unconfirmed by either company, follows Hugging Face's rejection of a $500M Nvidia investment last year that would have valued it at $7B. It gives Nvidia a foothold in the open-source ecosystem that rival labs building in-house chips, including Anthropic and OpenAI, increasingly depend on. (The Information)
Celera Semiconductorfundraise
Celera Semiconductor closed a $30M Series B funded entirely by Maverick Silicon, its largest investor, to expand its AI-driven analog chip design platform. The Santa Clara startup's Nesto technology compiles full-custom and standard analog ICs — the voltage-regulation circuitry that's become a bottleneck as individual AI accelerators draw more than 1kW — far faster than traditional design flows. The round follows Celera's acquisition of a Portugal-based design-automation team as it scales engineering headcount. (PR Newswire)
Public Markets
Nvidia reported $96.2B in quarterly revenue, up 106% year-over-year, with Data Center revenue of $89B and Q3 guidance of $108B versus a $104.2B estimate — a beat that sent shares up nearly 5% after hours, snapping a seven-session losing streak, with the outlook assuming zero China data-center revenue. (CNBC)
Emerging
Chip design: A new paper details TFA, a synthesizable INT8 accelerator for transformer inference that handles both prompt processing and token generation on one datapath, reporting roughly 20x speedup and a projected 1,000x energy reduction per token versus a 22-thread CPU baseline, in just 2.73mm² of logic. It's a small-team, open design rather than a hyperscaler product — evidence that competitive inference silicon no longer requires Nvidia-scale R&D budgets to prototype. (arXiv)
Memory computing: Samsung detailed LPDDR5X-PIM at Hot Chips, a drop-in memory chip that places compute logic inside standard LPDDR5X DRAM using the same JEDEC package, delivering 8x the bandwidth and roughly 3x the AI token throughput of ordinary LPDDR5X — including a jump from 27 to 81 tokens per second running Llama 3.1 — without a system redesign. Unlike HBM-based processing-in-memory efforts that need new packaging, this targets server, mobile, and client parts alike. (Tom's Hardware)