Modal is in talks to raise new financing at a roughly $15B valuation, about triple the mark it set in a round four months ago, as investors chase the infrastructure layer that runs AI models in production rather than trains them. The jump reflects a broader repricing already underway: spending on inference compute is on pace to overtake training spend as agentic workloads multiply the number of model calls per task, and inference-serving platforms are capturing that shift first. (Bloomberg)
Private Companies
Basetenfundraise
Baseten is in talks for a new funding round that could value the AI inference platform at roughly $26B, up from the $13B mark it set in June — a roughly 2x jump in three months. The pace reflects surging enterprise demand for dedicated inference infrastructure as businesses shift AI spending from training toward serving models in production; neither round has closed and terms remain undisclosed. (Bloomberg)
Firmusfundraise
Firmus is in talks with lenders for roughly $10B in financing — about $7.5B in debt and $2.5B in equity — to buy the Nvidia chips for its 360-megawatt "AI factory" in Batam, Indonesia, part of a $30B chip deal it signed with Nvidia. The facility is separate from Firmus's own late-October IPO and would fund roughly 170,000 GPUs delivered through 2027–2028 ahead of a first-quarter-2027 go-live — lenders underwriting GPU purchases directly rather than waiting on IPO proceeds. (Bloomberg)
Public Markets
IonQ$40.74▲ 11.0%Mkt Cap: $16.5B
IonQ said it will install its Superion 256 quantum processor at NVIDIA's new Accelerated Quantum Research Center, a day after unveiling what it called the industry's first real-time quantum error-correction decoder running on off-the-shelf hardware. (IonQ)
Emerging
Cooling and compute, scheduled together: A new paper, "ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters," from researchers at Hong Kong Polytechnic University, HKUST and the University of Macau, jointly schedules a datacenter's cooling setpoints alongside each GPU's clock speed and batch size, rather than treating cooling as a fixed facility constraint. The joint scheduler cut inference energy by up to 33% while keeping latency-target violations below 0.7% in testing — evidence the next efficiency gains come from treating compute and cooling as one control problem, not two. (arXiv)