When Being Human Becomes Valuable Data

By Javier Surasky

Original in Spanish; translated with ChatGPT assistance

An editorial conceptual composition showing a human figure surrounded by a large stream of repetitive synthetic data.

For years, one key problem in artificial intelligence was getting enough data. The Internet and social media seemed to offer a perfect solution: millions of books, photographs, articles, conversations, videos, code repositories, and web pages could be turned into raw material for training ever-larger models.

Generative AI is changing that equation because, although data continues to be produced at an ever-increasing pace, a growing share of it no longer comes directly from people but is artificially generated by the same systems that later need new data in order to keep learning.

The paradox deserves attention because it suggests that certain kinds of human data may become more valuable.

This is not to say that we are about to “run out of humanity” on the Internet, and I do not mean that synthetic data is necessarily bad. I am looking instead at a more concrete issue: even when accessing information, or producing it synthetically, is cheap, knowing where data comes from and under what conditions it was produced is beginning to acquire a value that was once taken for granted, but no longer can be.

A Machine Learning from Machines

The problem has a well-known technical dimension: a study published in Nature examined what happens when successive generations of models are trained indiscriminately on content produced by other models. Its authors called the result model collapse, defined as the “degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation” (Shumailov et al., 2024, p. 755).

Among the effects of this collapse is the loss of information located at the edges of a distribution: less frequent cases, rare variations, information that a previous model failed to reproduce with sufficient fidelity. In less technical terms, we might think of it as a certain “flattening” of difference, which led Shumailov and his coauthors to a conclusion that I find particularly interesting: in an environment saturated with model-generated content, data derived from genuine human interactions becomes increasingly valuable.

I do not want to overstate what this means, and to avoid misunderstandings I want to make clear from the outset that synthetic data has important uses. Recent work on medical data shows that it can facilitate the reuse of information and help address certain privacy risks, although its usefulness depends on how fidelity, utility, and data protection are balanced (Kaabachi et al., 2025).

This leads me to think that, while “data” as a resource may be abundant, data whose human provenance we know and can verify is not quite so abundant, and is becoming increasingly difficult to separate from our everyday “data salad.”

The Price of Knowing Where Data Comes From

There are signs that this shift already has an economic expression.

In September 2026, Snorkel, whose business has shifted increasingly toward producing complex datasets and reinforcement-learning environments for AI labs and other organizations, reached a valuation of USD 3.5 billion. What is really interesting here is that the company combines automation with specialized human knowledge in areas such as programming, law, and medicine. Its strength therefore lies not in selling large quantities of text, but in producing data designed to solve problems that require particular forms of expertise (Hu, 2026).

Data provenance is also attracting more attention.

Longpre et al. (2025) show that reconstructing the origin, conditions of use, licenses, and restrictions associated with datasets used in AI systems can reveal problems that are not visible when looking only at the final dataset.

We are, as I see it, approaching a point at which a dataset may be technically valuable while simultaneously becoming a legal or governance problem if no one can clearly reconstruct where it came from, under what conditions it was obtained, or what restrictions accompany its use. Recent research on licensing compliance reinforces this point by showing that the visible license attached to a dataset is not always enough to understand its legal status, because its lifecycle and provenance also need to be reconstructed (Kim et al., 2025).

This changes the logic of the data economy. During the first expansion of the digital economy, the dominant incentive was to accumulate more users, more interactions, more data. But generative AI has introduced a new variable: if producing enormous amounts of synthetic information is cheap, volume is no longer the main advantage, and value begins to shift toward data quality, provenance, and the conditions for safe use: a demonstration performed by a specialist, a conversation in a language underrepresented in training data, decisions made in real-life situations, records obtained with consent for their use, information whose chain of provenance can be traced.

The economic question then shifts toward the kinds of human experience that are actually needed to train a model.

What Kind of Humanity Does AI Need?

If it is true that human experience converted into data acquires economic value, it is also true that not all people or communities will participate in that market under the same conditions. Certain forms of knowledge will be sought because they improve a model’s reasoning; others will be valuable because they are underrepresented in training data, because of their high degree of specialization, or even because of the status of the person who produced them.

This raises a familiar question, although the answers may be about to change: who captures the value created by that data?

At least since the rise of the Internet, people have handed over enormous amounts of information to platforms in exchange for services without having any real idea of its value. The use of data to train AI models expanded that logic by drawing massively on content available on the web. But the next stage may prioritize quality over quantity and, as a result, become more selective. In other words, companies may no longer need just “any human data,” but rather data capable of adding something models cannot produce on their own.

Seen from outside the major technological centers, this could mean that countries and communities occupying peripheral positions in the global AI infrastructure may, precisely because of their marginalized position, be sources of data that is less represented in existing models. But that does not guarantee them bargaining power, as the history of commodity ownership has repeatedly shown. The analogy, I recognize, has its limits: human data does not exist separately from the people who produce it, and those people, it is worth remembering here, have privacy rights and their own interests in the experiences being turned into raw material for AI systems.

After Abundance

There is a certain consensus that three elements lie at the center of the race for artificial intelligence: computing capacity, digital infrastructure, and data. I have always been inclined to add a fourth: the human talent that makes the creation of AI models possible.

We have now reached a point at which human data and human talent form an unusual bridge, one in which “talent” includes the experiences particular to each person and community. Synthetic data can produce new variations and combinations, but it cannot replace human experiences that were never adequately represented in the original data. Outliers, so often treated as noise or as obstacles to identifying general patterns, may now acquire a different kind of value precisely because they contain what the average fails to represent.

As a logical consequence, human experiences outside the typical pattern are becoming more valuable in the digital economy, which may be moving toward a phase in which the more content machines are capable of producing, the greater the value attached to data that falls outside the average and whose human provenance can be demonstrated. And if that happens, the debate will shift over who decides which experiences are worth capturing, who can authorize their use, and who receives the benefits.

In a data economy saturated with averages, difference may become a valuable asset. In that economy, hungry for atypical and traceable human data, people return to the center from a different place, although the conditions under which that shift will take place still remain to be defined.


References

Hu, K. (2026, September 22). Snorkel AI valued at $3.5 billion amid surging demand for complex AI training data. Reuters. https://www.reuters.com/legal/transactional/snorkel-ai-valued-35-billion-amid-surging-demand-complex-ai-training-data-2026-09-22/

Kaabachi, B., Despraz, J., Meurers, T., Otte, K., Halilovic, M., Kulynych, B., Prasser, F., & Raisaro, J. L. (2025). A scoping review of privacy and utility metrics in medical synthetic data. npj Digital Medicine, 8, Article 60. https://doi.org/10.1038/s41746-024-01359-3

Kim, J., Sohn, S., Jo, G. J., Choi, J., Bae, K., Lee, H., Park, Y., & Lee, H. (2025). Do not trust licenses you see: Dataset compliance requires massive-scale AI-powered lifecycle tracing. arXiv. https://arxiv.org/abs/2503.02784

Longpre, S., Singh, N., Cherep, M., Tiwary, K., Materzynska, J., Brannon, W., Mahari, R., Obeng-Marnu, N., Dey, M., Hamdy, M., Saxena, N., Anis, A. M., Alghamdi, E. A., Chien, V. M., Yin, D., Qian, K., Li, Y., Liang, M., Dinh, A., . . . Kabbara, J. (2025). Bridging the data provenance gap across text, speech, and video. International Conference on Learning Representations. https://openreview.net/forum?id=R2Qd8ZK5uF

Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y