Before we dive in: a special invite for our readers
On September 15, we’re hosting Funding AI Compute - a morning of conversations on building in compute, backing startups, and financing the ecosystem. We have an incredible slate of speakers, including Kai Mak (CRO of Together), Erik Bernhardsson (CEO of Modal), Max Hjelm (SVP of Revenue at Coreweave), Steve Hou (Silicon Data) and many more. We’d love to have you in the room.
Sign up here, spots are limited.

OpenAI shared the first independent benchmarks for Jalapeño, its first custom inference chip, at Hot Chips 2026 today. Co-designed with Broadcom on TSMC's N3P node and built around HBM4 memory, the chip went from a design start in mid-2024 to tape-out in roughly nine months — an unusually fast cycle for a first-generation ASIC. SemiAnalysis, which reviewed the results, found Jalapeño beating Nvidia's Blackwell on tokens generated per watt across nearly every tested scenario and roughly matching the still-unreleased Vera Rubin on cost per output token, even without the speculative-decoding support Jalapeño doesn't have yet — an edge its analysts expect to widen once it does. For OpenAI's own account of why it built the chip, read hardware lead Richard Ho's announcement with Broadcom. The team will be presenting this afternoon at Hot Chips. More to come tomorrow. (SemiAnalysis)
Private Companies
Emerald AIfundraise
Emerald AI closed a $150M round at a $1.05B valuation, led by DCVC and Energize Capital, as pushback against new data centers hardens into a real constraint on the AI buildout — more than 75 projects worth roughly $130B were disrupted or blocked in the first quarter of 2026 alone over water use, noise, and rising utility bills. Led by former Biden administration energy official Varun Sivaram, the company's Emerald Conductor platform lets data centers throttle or reschedule compute in real time to ease grid stress, pitched as the fix that keeps local opposition from becoming a hard cap on new capacity. (DealBook)
Public Markets
Nvidia's own disclosure that AI server prices are rising more than 15% on memory costs extended a seventh straight losing streak into Wednesday's earnings, compounded today by benchmarks showing a customer's in-house chip closing the performance gap. (Tom's Hardware)
Micron fell with the rest of the memory complex even though the same cost inflation squeezing Nvidia's margins should lift its own — a sign the market is pricing Wednesday's earnings risk into every AI supplier at once. (The Motley Fool)
Emerging
Memory-bound serving: A new paper proposes KARAT, a processing-near-memory system that pulls the KV cache for long-context inference out of GPU memory entirely, splitting work between GPU nodes that hold model weights and separate memory-side nodes that handle retrieval and cache lookups. It's a direct technical answer to the same bottleneck now showing up in AI server price sheets — put the cache where memory is cheap, and reserve the GPU for what only a GPU can do. (arXiv)