Modal is in talks to raise new financing at a roughly $15B valuation, about triple the mark it set in a round four months ago, as investors chase the infrastructure layer that runs AI models in production rather than trains them. The jump reflects a broader repricing already underway: spending on inference compute is on pace to overtake training spend as agentic workloads multiply the number of model calls per task, and inference-serving platforms are capturing that shift first. (Bloomberg)
Private Companies
Basetenfundraise
Baseten is in talks for a new funding round that could value the AI inference platform at roughly $26B, up from the $13B mark it set in June — a roughly 2x jump in three months. The pace reflects surging enterprise demand for dedicated inference infrastructure as businesses shift AI spending from training toward serving models in production; neither round has closed and terms remain undisclosed. (Bloomberg)
Firmusfundraise
Firmus is in talks with lenders for roughly $10B in financing — about $7.5B in debt and $2.5B in equity — to buy the Nvidia chips for its 360-megawatt "AI factory" in Batam, Indonesia, part of a $30B chip deal it signed with Nvidia. The facility is separate from Firmus's own late-October IPO and would fund roughly 170,000 GPUs delivered through 2027–2028 ahead of a first-quarter-2027 go-live — lenders underwriting GPU purchases directly rather than waiting on IPO proceeds. (Bloomberg)
Public Markets
IonQ said it will install its Superion 256 quantum processor at NVIDIA's new Accelerated Quantum Research Center, a day after unveiling what it called the industry's first real-time quantum error-correction decoder running on off-the-shelf hardware. (IonQ)
Emerging
Cooling and compute, scheduled together: A new paper, "ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters," from researchers at Hong Kong Polytechnic University, HKUST and the University of Macau, jointly schedules a datacenter's cooling setpoints alongside each GPU's clock speed and batch size, rather than treating cooling as a fixed facility constraint. The joint scheduler cut inference energy by up to 33% while keeping latency-target violations below 0.7% in testing — evidence the next efficiency gains come from treating compute and cooling as one control problem, not two. (arXiv)