Callosum raised $100M in seed funding — one of the largest seed rounds in European history — led by Atomico, with Plural, DCVC, and the UK's Sovereign AI Fund also participating in the government fund's first-ever startup investment, according to Bloomberg. The London startup's Tailored Inference platform breaks AI workloads into smaller pieces and routes each to whichever chip — Nvidia, AMD, Cerebras, or a specialist accelerator — fits best without requiring code changes, and it paired the round with new partnerships with chipmakers Cerebras and Rebellions. The bet is that routing across chip vendors is how enterprises will actually run AI workloads at scale. (Bloomberg)
Private Companies
Fractilefundraise
Fractile is in talks to raise roughly $600M at a $6.5B pre-money valuation — more than six times the roughly $1B mark it set three months ago — after reportedly signing an initial $250M deal to sell its chips to Anthropic, according to Bloomberg. Fractile most recently announced a $220M fundraise in May 2026, valuing the company north of $1B. The British startup's chips are planning to ship in 2027. (Bloomberg)
Thunder Computefundraise
Thunder Compute raised a $13M Series A led by Matrix Partners, with Y Combinator and CEAS Investments participating. Its GPU virtualization software treats GPUs as a shared network resource, targeting the roughly $200B of AI compute capacity sitting idle inside data centers at any given time. The round marks a shift from proving the technology on Thunder's own cloud to selling it directly to cloud providers and enterprises that already run GPU fleets at scale. (SiliconANGLE)
Public Markets
Nebius fell after launching a $4.5B convertible-note offering, split across 2030 and 2034 tranches, to fund data-center construction and GPU purchases — debt financing priced into a week already nervous about how the AI buildout gets paid for. (Bloomberg)
Emerging
Edge inference: A new paper, EdgeXpert, accepted at MICRO 2026, combines mixture-of-experts routing with speculative decoding on-device, directly targeting the memory constraints that keep large language models off edge hardware today. As more inference work moves out of hyperscale data centers and onto local devices, architectures that treat memory — not raw compute — as the binding constraint are where the next round of efficiency gains is likely to come from. (arXiv)