Blackfuel came out of stealth with more than $250M in contracted revenue under multi-year customer agreements, and says it contracts demand first and builds capacity against it. Its first inference-only cluster sits in Digital Realty’s Barcelona data center, running AMD Instinct MI355X accelerators in liquid-cooled Dell racks, with further clusters planned in Spain, Finland and France and no outside investors disclosed. Selling reserved token output backed by multi-year contracts is a bid to make inference capacity financeable on AMD hardware, and it tests whether buyers will commit to tokens rather than GPU-hours. (Business Wire)
Private Companies
CScalefundraise
CScale exited stealth with a $145M Series C led by Atreides Management, with Valor Equity Partners and Premji Invest as co-leads and Nvidia and Intel Capital joining as its first strategic investors, bringing total funding to $188M. The Palo Alto company, founded in 2023, is building an integrated optical light engine for scale-up networks, designed so that a failed laser does not interrupt compute. CEO Martin Lund’s framing: “Lasers will fail. Compute shouldn’t.” Optical reliability across thousands of accelerators now draws the same strategic capital as bandwidth. (The Next Web)
GMI Cloudfundraise
GMI Cloud closed a $223M Series B led by ARCHIV with Nvidia participating, and added a credit facility led by CTBC that takes the total above $660M. The company reports contracted ARR above $600M, up about 9x since the end of 2025, and processes roughly 4 trillion tokens a week, with Fireworks among its named customers. Capacity expands beyond the U.S. into Taiwan, Japan and Southeast Asia. Debt now funds most of the build, which ties GMI’s growth to how long its customer contracts hold. (Business Wire)
Public Markets
Micron$1,065.11▲ 0.0%Mkt Cap: $1.2T
Micron reported fiscal fourth-quarter revenue of $54.2B and guided the next quarter to $61.5B, but its gross margin guide of about 86.3% sits below the 87.0% just posted, so further growth leans on volume more than on price increases. (Micron)
Emerging
Power per phase: A paper from Jae Gon Kim and colleagues, Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving, sets separate GPU power limits for the prefill and decode phases of a serving system instead of one cap per node. On an 8x B200 node running Qwen3-Coder-480B and Qwen3-235B-A22B, it improved energy efficiency 20.4% over Nvidia’s Max-Q profile at a 3.5% latency cost, with gains about five times smaller on dense models than on mixture-of-experts models. Power limits are becoming a per-model, per-phase scheduling decision, which shifts value toward serving software that can tune them. (arXiv)