Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
컴퓨트·클라우드리포트
컴퓨트·클라우드 · 리포트

Lumai는 Oxford 광학 컴퓨팅 스핀아웃 회사로, 빛을 이용한 전용 하드웨어 가속기 Iris Nova를 발표했습니다.

GlobeNewswire 보도자료 — 원본
공식 공시Slicast · September 11, 2026 · 글로벌 · 출처: GlobeNewswire

Frontier AI companies expect compute demand to grow roughly 1,000-fold over the next five years. Meeting that with conventional digital accelerators would require an estimated 100 trillion dollars in infrastructure investment and around 1,000 gigawatts of additional electrical capacity. As power becomes the binding constraint in data centers, efficiency at every stage of inference matters.

AI inference involves two distinct computational phases with different characteristics. Prefill processes the input context before generation begins and is dominated by highly parallel matrix multiplication, making it compute-bound. Decode generates output tokens and is driven more heavily by memory bandwidth and data movement. Leading infrastructure providers are disaggregating these workloads into separate pools, but Lumai argues the next step is specializing hardware for each stage.

Phil Burr, Head of Product at Lumai, explained the thinking: "If the workloads are fundamentally different, it makes sense to stop asking the same hardware to do both jobs. Disaggregating prefill and decode is an important step. But the bigger opportunity is to match the compute architecture to the workload."

The shift toward longer context windows, retrieval-augmented generation, multimodal inputs, and agentic workflows is increasing the information processed before token generation begins. A single user request can trigger multiple model calls, each carrying accumulated context, making prefill efficiency as critical as decode to overall inference economics. Today, hardware optimized for two jobs adequately often underperforms at either.

This creates three infrastructure challenges. Power becomes the binding constraint, converting a fixed megawatt budget into inference capacity. GPU capacity ends up stranded as the same fleet absorbs more prefill computation better suited for token generation. And inference economics becomes constrained, as the cost of processing context can make long-context and agentic applications difficult to price competitively.

Lumai Iris Nova uses optical computing to address the prefill workload. The hardware completes each vector-matrix multiplication in a single optical cycle, treating prefill as a purpose-built workload rather than running it on general-purpose silicon. In testing, Iris Nova ran billion-parameter language models in real time and demonstrated approximately 10 times more compute per watt than GPU-based equivalents on prefill workloads.

The approach frees GPU capacity for token generation, enabling existing infrastructure to become more productive without replacement. Burr stated: "Our view is not that every accelerator needs to be replaced, it is that we should stop asking every accelerator to solve every problem. The future AI data center will be increasingly heterogeneous, with different technologies optimized for different stages of inference."

Lumai, spun out of University of Oxford optics research in 2021, has received the Falling Walls Award for Science Breakthrough of the Year 2025 and won Best Overall Technology at the OCP Future Technologies Symposium. The company claims its technology delivers up to 90 percent lower energy consumption than conventional GPU architecture.

원문 보기