Scale AI: The Data Engine of AGI?
- Jul 19, 2024
- 2 min read
Updated: Jun 24
Scale AI, founded in 2016, has grown into one of the biggest suppliers of labeled data for AI development. The company works with leading organizations including OpenAI, Anthropic, Microsoft, and others, turning raw data into structured training datasets. While it has reported strong revenue growth — tripling ARR in 2023 and projecting around $1.4 billion in ARR by the end of 2024. — there are serious doubts about the sustainability of its core approach.
A major concern is the heavy reliance on human labeling. Human annotation is costly, slow, and often inconsistent. It can introduce biases and scaling problems that become harder to manage as models grow. Many in the AI field believe that at some point, extensive human labeling will no longer be required. Advances in synthetic data, self-supervised learning, and automated evaluation methods are likely to reduce — or even eliminate — the need for large-scale human-powered data labeling in the future.

The data pillar: AI development rests on three main pillars: data, compute, and algorithms. Scale AI focuses on the data side, helping transform messy raw data into usable assets. Its platform includes tools for labeling, quality control, reinforcement learning from human feedback (RLHF), and safety evaluations.
However, as models continue to scale, it remains unclear whether human-curated data will stay as critical as it is today. The path from current systems like GPT-4 toward much more capable models may depend less on ever-larger amounts of human-labeled data and more on new paradigms that move beyond today’s labor-intensive methods.
Scale AI currently holds a strong position in the market, but the rapid evolution of AI suggests its role could diminish over time as the industry shifts away from dependence on human labeling. The long-term importance of traditional data labeling companies remains an open question.


