Micro1's 80–90% margins come from reselling the same training data to many AI labs, a practice its 24 year old founder says now stops at the Chinese border.
Micro1 just told the world it runs at a $500 million gross annual run rate. The number is the line every AI training-data round will be measured against for the rest of the year.
The four-year-old San Francisco startup sells annotated examples to AI labs: humans and automated systems producing tagged text, images, and video that other companies use to train and evaluate their models. Data labeling is the work that sits under every frontier model, and the category has gone from back-office cost line to a $100 billion total addressable market, by the founder's estimate in the Aug. 20 TechCrunch piece. Micro1 is one of the most visible beneficiaries.
The 80% to 90% gross margin on Micro1's off-the-shelf data resale accounts for most of the $500 million gross run rate, according to a person familiar with the startup's finances, cited in the Aug. 20 TechCrunch piece. Once an AI lab licenses a labeled dataset, the same data can be sold to other labs at a marginal cost close to zero, and resale is now the bulk of the mix. The same source puts net run rate at $150 million to $200 million on a 60% to 70% customer-retention rate.
Micro1 is not the leader by gross run rate. Mercor reported a $2 billion gross annualized rate earlier this summer, and Handshake crossed $1 billion earlier in 2026. Micro1 moved from $7 million in gross run rate at the start of 2025 to $100 million by TechCrunch's December 2025 reporting to $500 million in August 2026, against a $35 million Series A that priced the company at $500 million in late 2025.
Founder and CEO Ali Ansari posted on X this week that Micro1 does not sell data to Chinese model makers, and named Kimi K3, a frontier model from Beijing-based Moonshot AI, as the cautionary example of where off-the-shelf U.S.-built training data can end up. The same post framed peer companies, by implication, as selling to "foreign adversaries." The $500 million run rate sits next to a buy-side decision: the resale structure that produces 80% margins is the same structure that puts U.S.-annotated data in adversarial hands, and Ansari is publicly choosing not to take the Chinese customer.
Off-the-shelf training data flowing to Chinese model builders has been a known pattern for at least a year, and rivals in the labeling category have not made the same public refusal. Ansari's bet is that picking customers, at the margin, will become table stakes for U.S. labeling shops, and that being first to do so in public buys reputational cover when the question is asked of the others. The data-labeling floor does not currently ask buyers to disclose their model lineage or end-use customer, and the Aug. 20 TechCrunch piece frames Ansari's post as the first public refusal of its kind from a major U.S. labeling shop.
Automated video description, synthetic data without human annotators, is a growing share of Micro1's mix, and the company has been hiring data partnership managers at $100,000 to $200,000 to license operational data for reinforcement-learning environments. The shift does not change the resale problem. A synthetic dataset is still a dataset, and it still goes to every buyer who can pay.
Micro1 started as an AI recruiting business in 2022 and pivoted to labeling after customers used the platform to vet annotators. Founder Ali Ansari is 24. The company was named in TechCrunch's December 2025 reporting as a Microsoft and Fortune 100 supplier for reinforcement-learning and post-training data. None of that explains the August 20 number. The resale structure does. The open question for the rest of the labeling shops is whether the China line Ansari drew on X becomes a category requirement, or stays a personal stance.