AI Power Map. #10: Precarious Data Workers

By Javier Surasky

Spanish version

Digital workers using laptops beneath an artificial intelligence cloud connected to data centers and global networks.

The Hidden Labor Infrastructure of Artificial Intelligence

Artificial intelligence learns, responds, classifies, translates, recommends, and moderates content based on enormous volumes of data and computational capacity. True, but insufficient.

Behind the automation of AI models lies an invisible labor infrastructure, made up of thousands of people who label images, transcribe audio, correct responses, compare outputs, detect bias, review toxicity, moderate violent content, or produce examples to train models, and it is a piece of the AI power map.

These workers, a crucial part of the AI production chain, make it possible for models to be trained, fine-tuned, and evaluated. Much of the safety and reliability of these systems rests on their labor and, yet, that labor is precarious and exploited to the very edge of legality, and even a little beyond it.

They are workers who perform their tasks in fragmented, outsourced, poorly paid ways and within a diffuse framework of legal protection.

Calling them “precarious data workers” makes it possible to name a part of the dominant narrative on innovation that, like any other narrative, orders the facts and people involved through stories of intentional exposure and concealment.

In this case, the power structure is simple: at the top are the large AI companies, which concentrate capital, infrastructure, intellectual property, brand, and decision-making capacity; one step below are supplier companies, subcontractors, and microwork platforms; and at the base are real people performing labeling, annotation, moderation, evaluation, or data-cleaning tasks. Economic and reputational value is captured at the top, while instability, opacity, and risk prevail below.

Precariousness serves a fundamental economic function: to externalize, at least partially, the labor costs of producing AI models and systems, sustained by the mythical narrative of an automatic AI, while concealing the dense network of human labor on which it depends.

Precariousness takes different forms, and the first is contractual: workers are excluded from labor recognition by the company that ultimately benefits from their work and appear as contractors, freelancers, platform workers, or outsourced personnel, which reduces job stability, access to labor and social protection rights, severance payments, and the possibility of collective bargaining.

Another form of precariousness is economic: payment is usually organized by task, batch, or productivity, shifting risk onto the worker. If there are no tasks available, if the platform changes its criteria, if the work is rejected, or if an account is blocked, these workers lose their income.

This way of “organizing” work ends up giving rise to an economy of permanent availability, in which the people subjected to this system compete for tasks, accept conditions imposed unilaterally, and remain always available, waiting to receive work assignments.

Precariousness also takes a form that we may call “informational,” since those who do the work may not know which company they are working for, which model they are training, how it will be used, how their performance is evaluated, or why a task is rejected. Much less do they know how to appeal a sanction applied to them, whether fairly or capriciously.

This brings us to a fourth element, the psychological dimension of precariousness. In content moderation or safety tasks, people are exposed for long hours to violent, sexual, racist, suicidal, or traumatic images, texts, or audio. They see what the models consider the rest of us should not see because of its possible effects on us (and on our attachment to those models, of course). The cornerstone of the safety that systems offer their users is a transfer of harm onto invisible workers.

A clear geopolitical dimension can also be identified, where patterns typical of the analog world are repeated: the most harmful tasks, and those carried out under the worst conditions, are displaced toward countries of the Global South or peripheral economies, where wages are lower, labor markets are more fragile, and labor protection standards are weaker, in addition to a broad market of workers in search of income.

To put it clearly: the global data chain reproduces and intensifies the international division of labor. Some actors design models, control infrastructure, and capture value; others label, clean, moderate, and correct under conditions of exploitation.

All of this builds a curtain of opacity that is, in itself, a form of expression of power, and it brings me to the hidden side of one of the most common questions in debates on AI: just as we ask which jobs artificial intelligence will replace, we should be asking what jobs it is generating now and under what conditions they are being developed.

Today, when people speak of the “knowledge society,” I see that data workers possess indispensable practical knowledge that is not being respected but, on the contrary, is being forced into conditions that keep it poorly paid within structures that smell of modern digital slavery, mediated by subcontractors that serve as legal and reputational buffer intermediaries, ranging from large exploiters to small microwork platforms that openly present themselves on the web and on social media.

Few States show interest in regulating this digital and cross-border labor framework, but unions and digital rights organizations are already emerging that seek to turn a dispersed labor force into a collective subject with its own capacity.

The tension between the speed and capacity of technological innovation and decent work is exposed in its greatest cruelty, shoulder to shoulder with the externalization of costs and the accumulation of wealth by companies we all know.

When AI ethics is discussed without including these workers, the discussion is poisoned from the start: it is difficult to understand how we can ask whether a model is safe, transparent, and explainable, but not about the labor conditions that made its training possible. An AI built through precarious work should not be called ethical or socially responsible.

A democratic governance of artificial intelligence should include, alongside rules on data, models, and risks for end users, transparency obligations regarding labor supply chains, access to rights, and psychosocial protection for those who perform tasks that put their health at risk. An AI that presents itself as aligned with human rights must allow full access to verification of labor relations that ensure decent work and the right to unionize, among others, within its own “production line.”

AI is a technical infrastructure, but it is also a way of organizing work, and if that dimension remains hidden, the public debate will continue to be partial and biased against the weakest within our societies.

Basic Data

  • Precarious data workers perform the labeling, annotation, transcription, moderation, evaluation, cleaning, and correction tasks needed to train, fine-tune, or supervise AI systems.
  • The value of the labeling market is measured by estimate; for data collection and labeling, it was estimated at USD 3.77 billion in 2024, with a projection of USD 17.1 billion by 2030 (Grand View Research), and for content moderation at USD 11.63 billion in 2025 and USD 13.31 billion in 2026, with estimates of USD 26.09 billion in 2031 (Mordor Intelligence).
  • The data labeling and moderation markets are growing rapidly, and their economic value is concentrated in the United States and in large technology companies, while their labor force is located mainly in countries such as India, the Philippines, Kenya, Colombia, Ghana, Venezuela, Brazil, and Madagascar.
  • Traditional labor statistics do not adequately capture work distributed across platforms, subcontractors, freelance contracts, business process outsourcing companies, and occasional tasks. The broadest estimates speak of between 154 and 435 million online digital platform workers, but the number of labelers, annotators, evaluators, and moderators remains unknown.
  • The main forms of precariousness are contractual, economic, informational, psychological, and geopolitical.
  • Outsourcing allows companies to benefit from labor without assuming direct labor responsibility and by externalizing costs: labor precariousness is part of the AI supply chain and must be treated as a central problem of technological governance, especially when its ethics and alignment with human rights are discussed.