“Data is the new code. The quality of your data determines the ceiling of your AI system.”
“We are in a technology race with China that is as consequential as the Cold War. And right now, we are not winning fast enough.”
“The biggest bottleneck in AI is not compute or algorithms. It's high-quality training data.”
“Every major AI breakthrough in the last five years has had a data innovation at its core, not just an algorithmic one.”
“The US government needs to treat AI like the Manhattan Project -- massive investment with clear national security purpose.”
“RLHF requires humans who are smarter than the model. As models get smarter, this gets harder. That is the core alignment data challenge.”
Alexandr Wang has built Scale AI into the critical infrastructure layer of the AI industry — the company that provides the high-quality human-labeled data that trains the world’s most capable AI models. His public statements reflect both this data-centric worldview and an increasingly prominent position as a voice for American AI competitiveness, particularly relative to China.
Wang’s emphasis on data quality as the primary bottleneck in AI development is self-interested (Scale AI sells data services) but also supported by the technical evidence. The difference between GPT-3 and ChatGPT was largely RLHF — a data innovation more than a model architecture change. His framing of data as “the new code” positions data preparation as a skilled engineering discipline rather than a commodity.
His geopolitical hawkishness on US-China AI competition has made Wang an influential voice in Washington, where he regularly testifies before Congress and advises the Department of Defense. He frames AI development as a national security imperative comparable to the nuclear arms race, arguing that American AI leadership is essential for maintaining global stability.
Wang became the youngest self-made billionaire in America at age 25, which gives his policy pronouncements an unusual combination of youthful energy and substantial economic credibility. His access to both Silicon Valley and Washington allows him to bridge two worlds that often misunderstand each other.
His observation about the RLHF challenge — that training data requires humans smarter than the model — identifies a fundamental scaling problem that the industry has not solved. As models approach and exceed human capabilities in more domains, the pool of humans qualified to provide training signal shrinks, creating a data bottleneck that Scale AI is uniquely positioned to address.