전용 AI 워크로드 제공업체 마이크로원은 전담 추론 인프라에 대한 급증하는 수요를 반영하여 연환산 기준 5억 달러의 매출을 달성했다.
Micro1 reports that its gross annual run rate has surged from $100 million to $500 million in just eight months. The AI data startup has achieved this growth by connecting machine learning developers with subject-matter experts who create and evaluate the increasingly specialized datasets required to refine their systems. This expansion occurs alongside soaring investments in AI infrastructure, underscoring another rapidly escalating cost in building advanced models: the data itself.
While the open internet provided ample material to train early-generation AI models, sourcing the precise data needed to push today’s more sophisticated systems further has become significantly more difficult.
Micro1 has positioned itself at the center of this challenge. Rather than simply providing mass labor for basic data labeling, the company recruits professionals with deep technical expertise and brokers partnerships between those specialists and AI firms seeking to enhance their models.
This value proposition is not unique to Micro1. Competitors such as Mercor and Handshake, along with several other firms in the sector, are also scaling businesses built on supplying human expertise to AI developers. The industry’s rapid expansion indicates that advancing model capabilities has not diminished the need for human involvement in development; rather, it has shifted the nature of the expertise required.
Micro1 was founded in 2022 by Ali Ansari, who has been vocal about where the company’s data ultimately resides. Last month, Ansari posted on X that Micro1 does not sell its data to Chinese model developers. He wrote, “Some human data companies work with foreign adversaries. [A]nd the results show today in Kimi K3. We believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.”
Financial interest extends beyond human expertise. Mercor reportedly surpassed $2 billion in gross annualized revenue this summer, while Handshake crossed the $1 billion mark earlier this year. Concurrently, existing proprietary datasets are attracting substantial commercial attention.
In a recent example, Micro1 offered $12.5 million for Spirit Airlines’ corporate data following the carrier’s bankruptcy, outbidding Google’s $10 million proposal. This unusual auction highlights another dimension of the AI data boom: enterprises are not only funding the creation of new training material, but are also actively acquiring proprietary datasets that cannot be scraped from the public internet.
The premium on proprietary data stems largely from its exclusivity. Leading AI developers generally share access to the same public web content and possess the capital to invest heavily in the computing power necessary to train and deploy their models. Corporate internal data, by contrast, is distinct. It represents the cumulative output of years of transactions, operational processes, and customer interactions—assets that cannot be easily replicated by competitors.
This distinction grows more critical as AI developers hunt for data capable of squeezing further performance gains from increasingly capable models. Although vast amounts of information exist online, a significant portion has already been exhaustively consumed for AI training. Because this public data is equally accessible to rivals, relying solely on it makes it increasingly difficult to secure a competitive advantage.
Unique private datasets can introduce novel patterns and insights that models have not previously encountered. Consequently, the race for superior AI is increasingly becoming a race to acquire information that competitors do not possess.
For organizations outside the artificial intelligence sector, this shift raises a fundamental question: what is the actual monetary value of the information they have already accumulated?
Historically, most enterprise data was generated to support daily business operations or satisfy regulatory mandates, not to train machine learning models. The emergence of dedicated buyers willing to purchase unconventional datasets may effectively assign a secondary commercial purpose to this legacy information.
Transforming corporate records into AI-ready assets, however, remains highly complex. Privacy regulations, intellectual property disputes, and commercially sensitive content impose strict boundaries on what companies can or will disclose. As a result, the utility of any given dataset will depend heavily on its quality and the specific objectives of the AI developer purchasing it.
Despite these hurdles, data is gradually shifting from a mere operational byproduct to a standalone asset class with its own marketplace. Micro1’s recent trajectory offers an early preview of how that market may evolve.