Thursday, September 24, 2026
AI 인프라 · 뉴스 & 분석
자본시장리포트
자본시장 · 리포트

Basecamp Research, which trains AI models on genetic material and evolutionary data, secures $140 million in funding from investors including Nvidia and Anthropic's Anthology Fund.

Cross-investor backing from Nvidia and Anthropic signals convergence between AI infrastructure players and foundation model labs on specialized data and computational approaches.
업계 전문지Slicast · 2026년 9월 23일 13:52 UTC · 글로벌 · 출처: The Decoder
중요도 70

Founded in 2020, Basecamp plans to deploy the funding to develop its biological AI models, called EDEN, and advance its own therapy candidates toward clinical development. The company's strategy centers on cell therapies—treatments that genetically modify a patient's cells to fight cancer and other diseases directly inside the body.

The company sources microorganisms from habitats that researchers have barely studied, training EDEN to learn patterns from their genomes that could have medical applications. "Given the complexity of biology, we just need orders of magnitude more data," says Philip Lorenz, Basecamp's CTO. The key, he emphasizes, is pairing that massive data collection with smaller, targeted experiments. So far, the work remains at the laboratory stage—mice studies and lab results—without evidence that the approach can produce safe and effective therapies in humans.

**The Scale Problem**

Pharma companies have poured billions into AI in recent years, yet solid evidence of materially higher success rates in clinical development remains scarce. Lorenz rejects this as reason for skepticism. Scaling delivered rapid gains for language models; biology is advancing too, but harder problems persist—speeding up clinical trials, for example. Biology, he argues, is "way, way, way more complicated and way larger."

The data landscape differs fundamentally. Language models learn from vast text collections; pharma companies typically work with small, siloed datasets from individual studies. Using an estimate from Epoch AI, Lorenz illustrates the gap: existing words total roughly five quadrillion tokens. Basecamp estimates the nucleotides on Earth at 10 to the power of 37—"If you took a stack of cards, like poker cards, of 10 to the power of 37 cards, that stack would surround the observable universe a million times," he says. While the company doesn't need a dataset that large, it does need one far larger than currently exists.

Public genome databases paint an incomplete picture. According to Basecamp's research on its BaseData database, roughly 68 percent of the sequence volume in the Sequence Read Archive comes from just five species. Humans alone account for about 54 percent. This bias makes sense for medical research but limits a model meant to capture as many biological processes as possible. "If you were to train an LLM only on newspaper articles from 1975, it would be a really, really bad model," Lorenz said in a Microsoft case study. "That is kind of where we are in biology."

Basecamp collects its own samples through a network of local research partners, sourcing material from rainforest soil, volcanic ground, and the deep sea off Antarctica. Its latest funding announcement reports the network spans more than 30 countries and all seven continents; Microsoft puts participating organizations at 208 across 31 countries. At the time of Lorenz's interview, the dataset held about 15 trillion tokens—comparable to text datasets used to train models like Claude or GPT, though here each token is a single DNA building block. The AI processes strings of characters describing genetic material rather than words.

The first EDEN generation trained on 9.7 trillion DNA building blocks from over a million newly sequenced species, with training conducted on Azure at computational levels reportedly matching GPT-4's. Over the next year and a half, the dataset is projected to grow roughly hundredfold, exceeding one quadrillion tokens. This expansion rests on the Trillion Gene Atlas announced in March, where Basecamp, Anthropic, Nvidia, PacBio, and Ultima Genomics plan to assemble genetic data at the scale of one trillion genes.

**Evolution as Drug Blueprint**

Lorenz views evolution—the force driving all biological processes—as the bridge between environmental samples and medicine. Similar selective pressures shape whether a bacterium adapts to heat or a cancer cell metastasizes. Most vertebrate genomes are sequenced; the rest of biodiversity remains largely uncatalogued. Yet many drugs originated in plants, bacteria, and fungi.

Competing microorganisms offer concrete examples. Powerful therapeutic molecules emerge from "biological warfare" between organisms. When a bacterial species seeks to dominate an ecosystem, it sometimes develops compounds that inhibit or kill rivals, securing food and habitat. This dynamic repeats billions of times globally, especially where nutrients are scarce—yielding antibiotics and other useful compounds. Lorenz also cites phages, viruses that infect bacteria and inject their DNA, producing tools for gene editing through evolutionary arms races.

Basecamp aims to learn from these evolved solutions. EDEN isn't merely rediscovering known compounds; it's designed to invent new candidates from learned patterns. The company records chemical, physical, and ecological conditions at each site alongside genetic material, focusing on long, continuous DNA stretches that reveal which genes sit adjacent and might cooperate. Short, isolated fragments lose this critical context.

**Early Medical Proof**

Lorenz points to EDEN's antibiotic designs as early validation. "We prompt on a pathogen and then the model designs an antibiotic that kills it," he says. An unpublished research paper on the EDEN model family reports that 97 percent of tested antimicrobial peptides—short protein chains targeting bacteria—showed activity in the laboratory. That rate applies only to the tested selection, not to everything the model might generate. Yet Basecamp went beyond test-tube work.

A June company announcement describes tests with EDEN-7, a candidate that in mice infected with multidrug-resistant bacteria performed as well as a last-resort antibiotic reserved for hard-to-treat infections. The model generated the candidate directly, without iterative tweaking before testing, in collaboration with University of Pennsylvania researchers led by Cesar de la Fuente. Lorenz emphasizes this design-to-experiment step: Basecamp isn't only about collecting data and training models. What matters is whether molecules actually work.

Beyond antibiotics, EDEN generated a synthetic gut microbiome of roughly 9,000 bacterial species, validated by UC Berkeley microbiome researcher Jill Banfield under strict criteria. "It is good to ask a skeptical human and not just AI," Lorenz said.

**The Therapy Strategy**

For its own therapy development, Basecamp is using EDEN to design biological tools that insert larger DNA stretches at chosen spots in the human genome. These include large serine recombinases—enzymes that rejoin DNA. Basecamp wants to engineer them to insert therapeutically useful genetic information into cells at precise locations. This approach tackles a core challenge in gene therapy: many inherited diseases stem from different mutations depending on the patient. Rather than fixing each individually, these insertion tools would insert a healthy gene copy regardless of the specific mutation. These tools derive from the same biological conflicts [text ends incomplete].

원문 보기
Basecamp Research, which trains AI models on… · Slicast