Friday, October 2, 2026
AI 인프라 · 뉴스 & 분석
홈 › 자본시장 › 리포트
자본시장 · 리포트

Nebius announced an acquisition of Inferize, a stealth-mode inference-optimization startup, for up to $150 million, less than nine months after Inferize's $10 million seed round.

The deal expands Nebius's inference capabilities and GPU utilization software, consolidating the fragmented inference-optimization sector and positioning Nebius against hyperscaler in-house solutions.
업계 전문지Slicast · 2026년 10월 1일 17:55 UTC · 미국 · 출처: calcalistech.com
중요도 75

Israeli AI infrastructure startup Inferize has been acquired by Nebius for up to $150 million in a deal that underscores the premium the industry places on solving inference economics at scale. The acquisition is particularly striking given Inferize's youth: founded earlier this year, the company has been operating largely in stealth, built a team of just 17 people in Tel Aviv, and completed its $10 million Seed round in less than nine months.

CEO and co-founder Guy Bortnikov disclosed the previously unreported funding round on LinkedIn, crediting TLV Partners for leading the investment and thanking Shahar Tzafrir and other backers. Bortnikov and CTO Lior Gorbonos, both members of the founding team at infrastructure optimization company Granulate—which Intel acquired in 2022 for approximately $650 million—developed technology aimed at a critical bottleneck in AI operations: model loading time.

AI models can take tens of minutes to load onto GPUs and other computing infrastructure, a delay that becomes particularly costly during inference, when models must respond to rapidly changing workloads. Inferize claims to compress that process to seconds, transforming what Bortnikov calls slow-loading inference engines into "true elastic workloads" that can span different architectures. The company worked with large AI cloud providers, inference platforms, and startups while developing the technology, though it has not publicly identified these customers.

The implications are substantial. As AI companies operate workloads that fluctuate dramatically from moment to moment, slow GPU provisioning leaves infrastructure idle or forces users to wait for capacity. Efficient model loading directly addresses a fundamental tension in cloud AI: balancing the capital cost of constantly available GPUs against the operational expense and latency of spinning them up on demand. Inference workloads, which occur every time an AI application responds to a request, are particularly sensitive to these dynamics. Small inefficiencies multiply at scale.

Bortnikov emphasized the speed of execution across the team, noting that Inferize's work demonstrated what could be achieved "when there were no limits." The Nebius acquisition brings that technology into Nebius Token Factory, the company's platform for production AI inference, marking the integration of Israeli infrastructure expertise into a broader cloud AI strategy.

원문 보기
Nebius announced an acquisition of Inferize, a… · Slicast