Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

Researchers propose using optical interconnects to build more efficient and high-bandwidth AI supercomputers.

Emerging interconnect technologies could reduce power consumption and latency bottlenecks in future high-performance clusters.
Trade pressSlicast · April 4, 2024 · Global · Source: wired.com
importance 60

Artificial intelligence experts widely agree that achieving the next major breakthrough in the field will require building supercomputers at unprecedented scale. At an event hosted by venture capital firm Sequoia last month, the CEO of startup Lightmatter presented a technology designed to enable this hyperscale computing shift by allowing chips to communicate directly using light rather than electrical signals. Currently, data moves through computers and between chips in data centers via electrical signals, with some interconnections converted to fiber-optic links for higher bandwidth. However, the repeated conversion between optical and electrical formats creates a communications bottleneck that limits speed. Lightmatter aims to directly connect hundreds of thousands or even millions of GPUs—the silicon chips essential to AI training—using optical links, with the conversion bottleneck reduction potentially enabling data movement between chips at far higher speeds than currently possible.

Lightmatter's technology, called Passage, consists of optical, or photonic, interconnects built in silicon that allow its hardware to interface directly with transistors on silicon chips like GPUs. The company claims this approach makes it possible to shuttle data between chips with 100 times the usual bandwidth. According to Lightmatter CEO Nick Harris, Passage will be ready by 2026 and should allow more than a million GPUs to run in parallel on the same AI training run. For context, GPT-4—OpenAI's most powerful AI algorithm and the technology behind ChatGPT—is rumored to have run on more than 20,000 GPUs. Harris confidently states that "Passage is going to enable AGI algorithms," referring to artificial general intelligence, the vaguely-specified goal of programs that can match or exceed biological intelligence in every way.

The intersection of Lightmatter's work with OpenAI's ambitions suggests the urgency of solving interconnection challenges at scale. Sam Altman, CEO of OpenAI, attended the Sequoia event and has repeatedly focused on building larger, faster data centers to advance AI. In February, The Wall Street Journal reported that Altman has sought up to $7 trillion in funding to develop vast quantities of AI chips, while The Information reported that OpenAI and Microsoft are drawing up plans for a $100 billion data center, codenamed Stargate, with millions of chips. Since electrical interconnects are extremely power-hungry, connecting chips on such a scale would require extraordinary amounts of energy and new connection methods like those Lightmatter proposes. Lightmatter is "working with the largest semiconductor companies in the world as well as the hyperscalers," Harris says, referring to companies like Microsoft, Amazon, and Google. GlobalFoundries, which makes chips for AMD and General Motors, previously announced a partnership with Lightmatter.

The scaling of computational capacity has proven fundamental to AI advances. More computation was essential to the breakthroughs that led to ChatGPT, and many AI researchers view further scaling-up of hardware as crucial to future algorithm improvements and to reaching AGI. If Lightmatter or another company can reinvent how giant AI projects are wired, a key bottleneck in smarter algorithm development could disappear. Current AI data centers consist of racks filled with tens of thousands of computers running specialized chips connected by mostly electrical wiring. Harris describes the problem: "Normally you have a bunch of GPUs, and then a layer of switches, and a layer of switches, and a layer of switches, and you have to traverse that tree" to communicate between two GPUs. In Lightmatter's vision, "every GPU would have a high-speed connection to every other chip."

Lightmatter's initiative reflects a broader industry trend of companies large and small attempting to reinvent hardware to advance AI. Nvidia, the leading GPU supplier for AI projects, held its annual conference last month, where CEO Jensen Huang unveiled Blackwell, the company's latest chip for training AI. Nvidia will sell Blackwell in a "superchip" consisting of two Blackwell GPUs and a conventional CPU processor, all connected using Nvidia's new high-speed communications technology, NVLink-C2C. The Blackwell GPUs are twice as powerful as their predecessors but consume much more power due to being bolted together. This trade-off, combined with Nvidia's efforts to glue chips together with high-speed links, suggests that upgrades to other key components for AI supercomputers, like those proposed by Lightmatter, could become increasingly important.

Read the original
Researchers propose using optical… · Slicast