How U.S. Data Startups Are Silently Supercharging Chinese AI Giants
When U.S. sales teams from Silicon Valley’s big data-labeling startups visited this year’s International Conference on Machine Learning in Seoul, Korea, they arrived prepped to court the industry’s big spenders. The data companies—collectively worth tens of billions of dollars and generating billions in annual revenue supplying training data to customers like OpenAI and Anthropic—are used to chasing AI labs that are notoriously demanding, fickle and difficult to satisfy.
They found another eager customer waiting for them: China’s AI industry.
Some Chinese companies have shopping lists. Tencent—which has previously been designated by the U.S. government as associated with the Chinese military, a characterization the company disputes— circulated with prospective vendors a detailed request for training data spanning finance, cybersecurity and one of AI’s most coveted research goals: AI systems capable of improving themselves.
While the U.S. has long prohibited China from buying the chips that power America’s top-tier AI models, it hasn’t taken the same precautions with the data required to train them. This packaged human expertise, produced by specialist networks and shaped by task designs, grading rubrics and quality controls, is part of the expensive data infrastructure that teaches models to handle difficult professional tasks like financial modeling and coding. The business is worth hundreds of millions of dollars annually and it’s helping Chinese AI models close the gap with their American rivals.
“After the compute needed to train the models, good quality data is the most important thing,” Nathan Lambert, a former research scientist at the Allen Institute for AI, told Forbes . “As we try to get models to move to new, more challenging domains, getting good data is the highest-leverage item that can make the model become viable. “
Documentation and communications from buyers at Chinese AI labs and employees from top Silicon Valley data platforms such as Surge AI and Mercor reviewed by Forbes reveal a quiet multi-hundred-million-dollar trade in training datasets. The same U.S. startups serving as suppliers to OpenAI, Anthropic, and, in some cases, U.S. federal entities, are also supplying Beijing’s top labs. Increasingly, Silicon Valley is training both sides of the AI arms race.
For Chinese AI labs chasing American AI performance, there are really only three fast-track moves. The first is to poach researchers. The second is distillation, harvesting outputs from ChatGPT or Claude and using them to train Chinese models. While not illegal, the practice explicitly violates OpenAI’s and Anthropic’s Terms of Service and recently was flagged by the Trump administration’s science and technology advisor Michael Kratsios. A third is easier and comes without potential regulatory headaches: buy the same training data the American AI labs are buying from the same data vendors they are buying it from.
Teaching an AI model to handle complex accounting or web lab protocols isn’t just about collecting the raw work output. The true secret sauce sits inside the data specifications: what type of data to focus on, as well as intricate rubrics written by human professionals that push and prod the model into learning to navigate white-collar work. By purchasing these datasets, the Chinese labs efficiently buy the hard-won, proprietary judgement powering Silicon Valley’s best models.
AI data consultant Sean Cai, who publishes an eponymous data and reinforcement learning blog, puts it bluntly: “If you actually study the shape of the data industry in China and then the entire data supply chain in general, you’ll realize that yes, a lot of the same U.S. data that powers U.S. model advancement is getting sold to Chinese labs as well.”
The $500 million secret data supply chain in China
The pipeline runs straight from the American data labeling companies to China’s biggest tech giants, including firms like Tencent, the Chinese company flagged by the Pentagon. In one message seen by Forbes , a Tencent data procurement executive noted that the company relies on Surge AI–whose customers have included the U.S. Army, the U.S. Air Force and Anthropic– for data provision. Mercor, which late last month told its employees it had won a contract with the U.S. federal government, also works with Tencent.
Other communications reviewed by Forbes show a similar overlap. In text exchanges, an executive from Ant Group, the Chinese tech and finance company behind Alipay, acknowledged a close working relationship with U.S. AI training data outfit AfterQuery. Audio transcripts from conversations with a Chinese ecommerce giant Alibaba employee point to purchase agreements with AfterQuery and Mercor. Internal project documentation suggests US data labeling startup Turing has worked with ByteDance, TikTok’s Chinese parent.
ByteDance, Alibaba, Moonshot, Tencent and Ant Group didn’t respond to requests for comment; Mercor declined to comment. AfterQuery and Surge said they don’t disclose customer information. In a statement, Turing said: “We power the frontier of AI, which includes both proprietary and open source models. We believe both will keep growing, and some of the strongest open source models today come from labs outside the U.S.”
All in all, the top six Chinese AI labs spend about $500 million with American data labeling companies per year, according to two data labeling entrepreneurs who have received market size estimates from Tencent and ByteDance executives. According to one person who has interacted with a variety of them, Chinese AI labs don’t view themselves as being in direct procurement competition with each other; rather they see themselves as a collective buyer pool looking for top-tier Western datasets.
U.S. data labeling companies are happy to sell to them. According to a source familiar with the matter, Mercor’s Q2 revenue from Chinese AI labs was 2% of total revenue (Mercor’s annualized revenue crossed $2 billion in June), and the company has hired a specific engagement manager for China. A larger portion of AfterQuery’s business, at least $50 million in recurring revenue, comes from Chinese AI labs, sources told Forbes. And Surge AI, an early pioneer in the region, has actively courted these buyers, with CEO Edwin Chen traveling to China to meet lab executives directly.
Data labelers sell Chinese AI labs custom, one-off projects as well as what are known as "Off-the-Shelf" (OTS) datasets: standardized, pre-packaged training sets that are meant to be resold to multiple labs. Sounds harmless enough, but off-the-shelf datasets are far from innocuous. They are manufactured with the same "knowledge pipelines" developed during custom engagements for OpenAI or Anthropic. That means the Chinese companies that buy them benefit from the same PhD expert networks, quality control algorithms, and post-training rubrics that help inform the training done at the big U.S. AI labs. And that is enormously helpful. In many cases, Chinese labs can bypass months of trial-and-error by buying OTS datasets with battle-tested reasoning structures.
"You've got to realize the shapes of this data sometimes come from the same companies that tailored their initial scaling directions for Anthropic," Cai, the AI data consultant, explained. "It is extremely easy for a data vendor that has served Anthropic to use the same stack to produce data for a Chinese lab."
This makes off-the-shelf datasets a tempting financial engine for data vendors. Because they can be resold to multiple AI labs, they carry the highest profit margins. And Chinese AI labs explicitly tell U.S. data labelers that they want to buy all the data that the U.S. labs have already purchased, said one executive who has spoken to Chinese AI labs.
Max Gazor, founder of Striker VC, which has invested in many of the American AI model builders, acknowledged this market reality comes with certain complicating factors. “I would distinguish it as more of an ethical decision of whether the company would sell the data to Chinese AI companies,” he said.
Washington has spent years building a geopolitical fortress around advanced chips. But the human expertise used to train AI manages to cross borders with almost no scrutiny at all.
That’s helped make Silicon Valley’s top data platforms multi-billion-dollar titans. Founded in 2020 and 2023, Mercor is currently eyeing a valuation of $20 billion, while Surge AI is reportedly worth at least $25 billion. Meanwhile, relative newcomer AfterQuery has grown from $100 million to hundreds of millions in annual recurring revenue in three months. While the vast majority of it still comes from established American giants, foreign sales are likely an irresistible secondary revenue stream.
Some say selling datasets to China is a matter of national security. “Some human data companies work with foreign adversaries. and the results show today in Kimi K3,” Ali Ansari, the company’s 25-year-old CEO, recently posted to X. ”We believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.” Micro1 says it does not sell data outside the U.S. Scale AI also flirted with a ByteDance contract before scrapping it in 2024 over national security concerns.
Defenders of the practice argue clamping down on data sales would choke off global open-source innovation.
“One cannot easily restrict exports of data or AI in a way that does not materially hurt American open source,” said Cai, noting that he does not have a position on the issue. “It gives an unfair duopoly to OpenAI and Anthropic, and that has its own problems.”
Loading article...