By Javier Surasky
The Hidden Labor Infrastructure of Artificial Intelligence
Artificial
intelligence learns, responds, classifies, translates, recommends, and
moderates content based on enormous volumes of data and computational capacity.
True, but insufficient.
Behind the
automation of AI models lies an invisible labor infrastructure, made up of
thousands of people who label images, transcribe audio, correct responses,
compare outputs, detect bias, review toxicity, moderate violent content, or
produce examples to train models, and it is a piece of the AI power map.
These
workers, a crucial part of the AI production chain, make it possible for models
to be trained, fine-tuned, and evaluated. Much of the safety and reliability of
these systems rests on their labor and, yet, that labor is precarious and
exploited to the very edge of legality, and even a little beyond it.
They are
workers who perform their tasks in fragmented, outsourced, poorly paid ways and
within a diffuse framework of legal protection.
Calling
them “precarious data workers” makes it possible to name a part of the dominant
narrative on innovation that, like any other narrative, orders the facts and
people involved through stories of intentional exposure and concealment.
In this
case, the power structure is simple: at the top are the large AI companies,
which concentrate capital, infrastructure, intellectual property, brand, and
decision-making capacity; one step below are supplier companies,
subcontractors, and microwork platforms; and at the base are real people performing labeling, annotation, moderation, evaluation, or data-cleaning tasks. Economic and reputational value is captured at the top, while
instability, opacity, and risk prevail below.
Precariousness
serves a fundamental economic function: to externalize, at least partially, the
labor costs of producing AI models and systems, sustained by the mythical
narrative of an automatic AI, while concealing the dense network of human labor
on which it depends.
Precariousness
takes different forms, and the first is contractual: workers are excluded from
labor recognition by the company that ultimately benefits from their work and
appear as contractors, freelancers, platform workers, or outsourced personnel,
which reduces job stability, access to labor and social protection rights,
severance payments, and the possibility of collective bargaining.
Another
form of precariousness is economic: payment is usually organized by task,
batch, or productivity, shifting risk onto the worker. If there are no tasks
available, if the platform changes its criteria, if the work is rejected, or if
an account is blocked, these workers lose their income.
This way of
“organizing” work ends up giving rise to an economy of permanent availability,
in which the people subjected to this system compete for tasks, accept
conditions imposed unilaterally, and remain always available, waiting to
receive work assignments.
Precariousness
also takes a form that we may call “informational,” since those who do the work
may not know which company they are working for, which model they are training,
how it will be used, how their performance is evaluated, or why a task is rejected.
Much less do they know how to appeal a sanction applied to them, whether fairly
or capriciously.
This brings
us to a fourth element, the psychological dimension of precariousness. In
content moderation or safety tasks, people are exposed for long hours to
violent, sexual, racist, suicidal, or traumatic images, texts, or audio. They
see what the models consider the rest of us should not see because of its
possible effects on us (and on our attachment to those models, of course). The
cornerstone of the safety that systems offer their users is a transfer of harm
onto invisible workers.
A clear
geopolitical dimension can also be identified, where patterns typical of the
analog world are repeated: the most harmful tasks, and those carried out under
the worst conditions, are displaced toward countries of the Global South or
peripheral economies, where wages are lower, labor markets are more fragile,
and labor protection standards are weaker, in addition to a broad market of
workers in search of income.
To put it
clearly: the global data chain reproduces and intensifies the international
division of labor. Some actors design models, control infrastructure, and
capture value; others label, clean, moderate, and correct under conditions of
exploitation.
All of this
builds a curtain of opacity that is, in itself, a form of expression of power,
and it brings me to the hidden side of one of the most common questions in
debates on AI: just as we ask which jobs artificial intelligence will replace,
we should be asking what jobs it is generating now and under what conditions
they are being developed.
Today, when
people speak of the “knowledge society,” I see that data workers possess
indispensable practical knowledge that is not being respected but, on the
contrary, is being forced into conditions that keep it poorly paid within
structures that smell of modern digital slavery, mediated by subcontractors
that serve as legal and reputational buffer intermediaries, ranging from large
exploiters to small microwork platforms that openly present themselves on the
web and on social media.
Few States
show interest in regulating this digital and cross-border labor framework, but
unions and digital rights organizations are already emerging that seek to turn
a dispersed labor force into a collective subject with its own capacity.
The tension
between the speed and capacity of technological innovation and decent work is
exposed in its greatest cruelty, shoulder to shoulder with the externalization
of costs and the accumulation of wealth by companies we all know.
When AI ethics is discussed without including these workers, the discussion is poisoned from the start: it is difficult to understand how we can ask
whether a model is safe, transparent, and explainable, but not about the labor
conditions that made its training possible. An AI built through precarious work
should not be called ethical or socially responsible.
A
democratic governance of artificial intelligence should include, alongside
rules on data, models, and risks for end users, transparency obligations
regarding labor supply chains, access to rights, and psychosocial protection
for those who perform tasks that put their health at risk. An AI that presents
itself as aligned with human rights must allow full access to verification of
labor relations that ensure decent work and the right to unionize, among
others, within its own “production line.”
AI is a
technical infrastructure, but it is also a way of organizing work, and if that
dimension remains hidden, the public debate will continue to be partial and
biased against the weakest within our societies.
Basic Data
- Precarious data workers perform the labeling, annotation, transcription, moderation, evaluation, cleaning, and correction tasks needed to train, fine-tune, or supervise AI systems.
- The value
of the labeling market is measured by estimate; for data collection and
labeling, it was estimated at USD 3.77 billion in 2024, with a projection of
USD 17.1 billion by 2030 (Grand View Research), and for content moderation at USD
11.63 billion in 2025 and USD 13.31 billion in 2026, with estimates of USD
26.09 billion in 2031 (Mordor Intelligence).
- The data labeling and moderation markets are growing rapidly, and their economic value is concentrated in the United States and in large technology companies, while their labor force is located mainly in countries such as India, the Philippines, Kenya, Colombia, Ghana, Venezuela, Brazil, and Madagascar.
- Traditional labor statistics do not adequately capture work distributed across platforms, subcontractors, freelance contracts, business process outsourcing companies, and occasional tasks. The broadest estimates speak of between 154 and 435 million online digital platform workers, but the number of labelers, annotators, evaluators, and moderators remains unknown.
- The main forms of precariousness are contractual, economic, informational, psychological, and geopolitical.
- Outsourcing allows companies to benefit from labor without assuming direct labor responsibility and by externalizing costs: labor precariousness is part of the AI supply chain and must be treated as a central problem of technological governance, especially when its ethics and alignment with human rights are discussed.
.png)