Anthropic is in talks to acquire Decart AI, the Tel Aviv real-time-video and GPU-optimization startup, for roughly $6B — the largest acquisition in Anthropic's history and a premium of about 50% over the roughly $4B valuation Decart carried after its last round. Decart's software squeezes more inference out of the same silicon, and Anthropic wants that efficiency inside its own stack as Claude demand keeps outrunning the compute it can rent or build. The deal signals where frontier labs now see the highest-value spend: not the next GPU order, but software that makes the GPUs already on order do more work. Decart co-founder Dean Leitersdorf's conversation joined Sequoia's Training Data podcast for a relevant conversation on what Decart journey. (Bloomberg)
Private Companies
FirmusACQUISITION
Firmus Technologies agreed to acquire the fabrication, design, and projects businesses of Benmax, the Queanbeyan manufacturer that has built its modular HyperCube AI Factory system since 2020, for A$300M. The deal folds a six-year external manufacturing partnership into Firmus's own operating structure, more than doubles its global headcount to 340, and brings Benmax managing director Scott Polsen in as Firmus's chief development officer. It's a vertical-integration move into physical manufacturing capacity — the same logic neoclouds have applied to orchestration software — five months after Firmus raised $2B at a $10.5B valuation. (SMBtech)
Public Markets
Q2 revenue of $180M missed the $194M Street estimate and margins compressed as Cerebras rented back its own systems from cloud customers to cover demand — even as its chips began powering a new "Ultrafast" tier of OpenAI's API at up to 14x normal speed. (Yahoo Finance)
Management raised its 2026 revenue outlook for the second time this year, to €43B–€45B, and lifted its long-term margin target to 54–56% on an order backlog stretching into 2028 — evidence that EUV lithography capacity, not GPU output, is now the AI buildout's binding constraint. (Yahoo Finance)
Emerging
Edge inference gets an MoE trick: EdgeXpert combines speculative decoding with mixture-of-experts inference on-device, resolving the usual incompatibility between the two techniques through prompt-wise expert reuse and depth-aware expert coalescing. Synthesized in Samsung's 28nm process, the design cuts latency by up to 56% and energy by up to 44% against existing edge-inference approaches — a data point for how much specialized silicon can still squeeze out of on-device LLM inference before a query needs to reach a datacenter GPU at all. (arXiv)
❝
Join us at the NYC Tech Summit on September 16. Learn from the founders of Etched, Cloudflare, and Runway. Connect with the leaders shaping what's next in tech. Applications now open at primarysummit.vc.