By Javier Surasky
Original in Spanish; translated with ChatGPT assistance
For years, one key problem in artificial intelligence was getting enough data. The Internet and social media seemed to offer a perfect solution: millions of books, photographs, articles, conversations, videos, code repositories, and web pages could be turned into raw material for training ever-larger models.
Generative
AI is changing that equation because, although data continues to be produced at
an ever-increasing pace, a growing share of it no longer comes directly from
people but is artificially generated by the same systems that later need new
data in order to keep learning.
The paradox
deserves attention because it suggests that certain kinds of human data may
become more valuable.
This is not
to say that we are about to “run out of humanity” on the Internet, and I do not
mean that synthetic data is necessarily bad. I am looking instead at a more
concrete issue: even when accessing information, or producing it synthetically,
is cheap, knowing where data comes from and under what conditions it was
produced is beginning to acquire a value that was once taken for granted, but
no longer can be.
A Machine Learning from Machines
The problem
has a well-known technical dimension: a study published in Nature
examined what happens when successive generations of models are trained
indiscriminately on content produced by other models. Its authors called the
result model collapse, defined as the “degenerative process affecting
generations of learned generative models, in which the data they generate end
up polluting the training set of the next generation” (Shumailov et al., 2024,
p. 755).
Among the
effects of this collapse is the loss of information located at the edges of a
distribution: less frequent cases, rare variations, information that a previous
model failed to reproduce with sufficient fidelity. In less technical terms, we
might think of it as a certain “flattening” of difference, which led Shumailov
and his coauthors to a conclusion that I find particularly interesting: in an
environment saturated with model-generated content, data derived from genuine
human interactions becomes increasingly valuable.
I do not
want to overstate what this means, and to avoid misunderstandings I want to
make clear from the outset that synthetic data has important uses. Recent work
on medical data shows that it can facilitate the reuse of information and help
address certain privacy risks, although its usefulness depends on how fidelity,
utility, and data protection are balanced (Kaabachi et al., 2025).
This leads
me to think that, while “data” as a resource may be abundant, data whose human
provenance we know and can verify is not quite so abundant, and is becoming
increasingly difficult to separate from our everyday “data salad.”
The Price of Knowing Where Data Comes From
There are
signs that this shift already has an economic expression.
In
September 2026, Snorkel, whose business has shifted increasingly toward
producing complex datasets and reinforcement-learning environments for AI labs
and other organizations, reached a valuation of USD 3.5 billion. What is really
interesting here is that the company combines automation with specialized human
knowledge in areas such as programming, law, and medicine. Its strength
therefore lies not in selling large quantities of text, but in producing data
designed to solve problems that require particular forms of expertise (Hu,
2026).
Data
provenance is also attracting more attention.
Longpre et
al. (2025) show that reconstructing the origin, conditions of use, licenses,
and restrictions associated with datasets used in AI systems can reveal
problems that are not visible when looking only at the final dataset.
We are, as
I see it, approaching a point at which a dataset may be technically valuable
while simultaneously becoming a legal or governance problem if no one can
clearly reconstruct where it came from, under what conditions it was obtained,
or what restrictions accompany its use. Recent research on licensing compliance
reinforces this point by showing that the visible license attached to a dataset
is not always enough to understand its legal status, because its lifecycle and
provenance also need to be reconstructed (Kim et al., 2025).
This
changes the logic of the data economy. During the first expansion of the
digital economy, the dominant incentive was to accumulate more users, more
interactions, more data. But generative AI has introduced a new variable: if
producing enormous amounts of synthetic information is cheap, volume is no
longer the main advantage, and value begins to shift toward data quality,
provenance, and the conditions for safe use: a demonstration performed by a
specialist, a conversation in a language underrepresented in training data,
decisions made in real-life situations, records obtained with consent for their
use, information whose chain of provenance can be traced.
The
economic question then shifts toward the kinds of human experience that are
actually needed to train a model.
What Kind of Humanity Does AI Need?
If it is
true that human experience converted into data acquires economic value, it is
also true that not all people or communities will participate in that market
under the same conditions. Certain forms of knowledge will be sought because
they improve a model’s reasoning; others will be valuable because they are
underrepresented in training data, because of their high degree of
specialization, or even because of the status of the person who produced them.
This raises
a familiar question, although the answers may be about to change: who captures
the value created by that data?
At least
since the rise of the Internet, people have handed over enormous amounts of
information to platforms in exchange for services without having any real idea
of its value. The use of data to train AI models expanded that logic by drawing
massively on content available on the web. But the next stage may prioritize
quality over quantity and, as a result, become more selective. In other words,
companies may no longer need just “any human data,” but rather data capable of
adding something models cannot produce on their own.
Seen from
outside the major technological centers, this could mean that countries and
communities occupying peripheral positions in the global AI infrastructure may,
precisely because of their marginalized position, be sources of data that is
less represented in existing models. But that does not guarantee them
bargaining power, as the history of commodity ownership has repeatedly shown.
The analogy, I recognize, has its limits: human data does not exist separately
from the people who produce it, and those people, it is worth remembering here,
have privacy rights and their own interests in the experiences being turned
into raw material for AI systems.
After Abundance
There is a
certain consensus that three elements lie at the center of the race for
artificial intelligence: computing capacity, digital infrastructure, and data.
I have always been inclined to add a fourth: the human talent that makes the
creation of AI models possible.
We have now
reached a point at which human data and human talent form an unusual bridge,
one in which “talent” includes the experiences particular to each person and
community. Synthetic data can produce new variations and combinations, but it
cannot replace human experiences that were never adequately represented in the
original data. Outliers, so often treated as noise or as obstacles to identifying
general patterns, may now acquire a different kind of value precisely because
they contain what the average fails to represent.
As a
logical consequence, human experiences outside the typical pattern are becoming
more valuable in the digital economy, which may be moving toward a phase in
which the more content machines are capable of producing, the greater the value
attached to data that falls outside the average and whose human provenance can
be demonstrated. And if that happens, the debate will shift over who decides
which experiences are worth capturing, who can authorize their use, and who
receives the benefits.
In a data
economy saturated with averages, difference may become a valuable asset. In
that economy, hungry for atypical and traceable human data, people return to
the center from a different place, although the conditions under which that
shift will take place still remain to be defined.
References
Hu, K.
(2026, September 22). Snorkel AI valued at $3.5 billion amid surging demand
for complex AI training data. Reuters. https://www.reuters.com/legal/transactional/snorkel-ai-valued-35-billion-amid-surging-demand-complex-ai-training-data-2026-09-22/
Kaabachi,
B., Despraz, J., Meurers, T., Otte, K., Halilovic, M., Kulynych, B., Prasser,
F., & Raisaro, J. L. (2025). A scoping review of privacy and utility
metrics in medical synthetic data. npj Digital Medicine, 8, Article 60. https://doi.org/10.1038/s41746-024-01359-3
Kim, J.,
Sohn, S., Jo, G. J., Choi, J., Bae, K., Lee, H., Park, Y., & Lee, H.
(2025). Do not trust licenses you see: Dataset compliance requires
massive-scale AI-powered lifecycle tracing. arXiv. https://arxiv.org/abs/2503.02784
Longpre, S., Singh, N., Cherep, M., Tiwary, K., Materzynska,
J., Brannon, W., Mahari, R., Obeng-Marnu, N., Dey, M., Hamdy, M., Saxena, N.,
Anis, A. M., Alghamdi, E. A., Chien, V. M., Yin, D., Qian, K., Li, Y., Liang,
M., Dinh, A., . . . Kabbara,
J. (2025). Bridging the data provenance gap across text, speech, and video. International
Conference on Learning Representations. https://openreview.net/forum?id=R2Qd8ZK5uF
Shumailov,
I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024).
AI models collapse when trained on recursively generated data. Nature, 631,
755–759. https://doi.org/10.1038/s41586-024-07566-y
.png)