By Javier Surasky
The
original article was written in Spanish. This English version is a translation
prepared with ChatGPT and reviewed by the author.
The Social Infrastructure of Open Models
Hugging
Face occupies a place on the AI power map as a technology company that also
serves as the social infrastructure that allows open models to circulate, gain
legitimacy, and be adopted. Code, datasets, and applications are hosted there
under conditions that give them visibility, reusability, comparability, and
political significance.
The
platform operates as a meeting point for research, business development, open
source communities, major labs, startups, universities, and users looking for
alternatives to closed models. The Hugging Face Hub brings together models,datasets, and “Spaces”: repositories for models, datasets, and demoapplications that make it possible to test, document, and share machine
learning systems.
Hugging
Face’s power comes precisely from its ability to organize code, model weights,
datasets, documentation cards, demos, download metrics, leaderboards, licenses,
community spaces, and enterprise services, helping shape a critical part of the
contemporary AI supply chain.
The key
point is that, for an open model to circulate, it needs hosting, a format,
metadata, documentation, a license, reputation, discovery mechanisms, a user
community, integration with libraries, and, increasingly, safety conditions.
Hugging Face appears as a bridge across that intermediate zone between the
technical artifact and its institutional trajectory.
Thomas Wolf
et al. described “Transformers” as an open source library designed to put
state-of-the-art architectures within reach of the machine learning community,
combining a unified API with a collection of pretrained models available to
researchers and developers (Wolf et al., 2020). This perspective helps show how
Hugging Face reduces technical friction so that complex models can be used by a
broader community.
However,
that technical democratization also produces new infrastructural dependencies
and creates related risks.
The term
“foundation models” was popularized by the report On the Opportunities and
Risks of Foundation Models, prepared by Rishi Bommasani et al. (2021, p.
3), who defined them as models trained on broad data, generally through
self-supervised learning at scale, that can be adapted to a wide range of
downstream tasks and therefore create incentives toward homogenization: when a
base model is reused in many contexts, its flaws can spread. Hugging Face
facilitates exactly that spread.
As a
result, from the perspective of power analysis, Hugging Face creates
opportunities for researchers, small companies, educators, public
organizations, and technical communities that could not train large-scale
models from scratch. But it does so at the cost of turning a private platform
into an almost unavoidable passage point for a significant part of the open
ecosystem: openness reorganizes concentration, but it does not eliminate it.
This points
to another important issue: Hugging Face’s internal tension as both community
infrastructure and company.
The community contributes models, datasets, documentation, and other content, but Hugging Face, as a company, organizes the platform where these materials are
concentrated, defines tools, offers paid services, sets access conditions, and
produces a layer of private governance within open source software. This gives
the platform intermediary power, insofar as it directly influences how its
components are discovered, evaluated, and reused.
Similarly,
by functioning almost as a “Strait of Hormuz of open source,” relevance signals
such as likes, downloads, model cards, rankings, Spaces, tags, and library
integrations end up directing attention, investment, and experimentation, while
generating trust or distrust toward models.
Documentation
plays a key role in that process. Hugging Face’s official documentation
incorporates model cards and dataset cards as relevant components of the Hub,
helping create transparency around intended uses, limitations, data, metrics,
and evaluation conditions. Yet recent studies on documentation in Hugging Face
show unevenness: Yang, Liang, and Zou (2024) analyze 7,433 dataset cards and
find that descriptive and structural sections tend to be more developed than
those addressing use considerations, limitations, and social impacts; Liang et
al. (2024), based on 32,111 model cards, find that sections on limitations,
evaluation, and environmental impact have lower completion rates than technical
sections; Erfan, Ryan, and Rahman (2026) extend that concern to AI Bills of
Materials in Hugging Face repositories, showing persistent gaps in information
on datasets, risks, limitations, safety, and traceability.
This
documentation gap is politically relevant: a model can be available, downloaded
thousands of times, and have a functional demo while still remaining opaque
with respect to some of its fundamental elements, including the possible social
or environmental consequences of its use. Although the information uploaded
depends on those who publish it, the platform still determines how those
information fields are structured.
Hugging Face also participates in the dispute over what “open” means in AI. In free
software, openness generally refers to access to source code and the
possibility of modifying and redistributing it. In generative AI, however, the
issue is more complex, because code, weights, architecture, training data,
evaluations, logs, training recipes, or only some of those components may be
opened. This allows a model to be “open” in an operational sense while
remaining closed in areas that matter for accountability. Hugging Face operates
within that space of ambiguity when it allows models and datasets to be shared,
but cannot guarantee that the published artifacts are fully auditable, legally
safe, socially appropriate, or scientifically reproducible.
The
geopolitical dimension is no less important. Open models allow companies,
states, and communities to reduce their dependence on a handful of proprietary
APIs and enable linguistic, sectoral, and regional adaptations. But the
infrastructure that makes this possible remains concentrated among actors with
resources, servers, capital, technical talent, and alliances with major cloud
providers. Hugging Face helps decentralize model production, but not
necessarily democratize the material layers of AI.
This is why
Hugging Face can be understood as a soft form of infrastructural power: it does
not impose binding obligations, but positions itself through interfaces,
practical standards, repositories, documentation, permissions, rankings, APIs,
terms of use, and enterprise services grounded in its organizational authority.
Hugging
Face’s role in the AI power map unfolds through five dimensions.
- The infrastructural dimension, where its Hub emerges as a space for hosting, searching, downloading, documenting, and deploying models, datasets, and applications, turning Hugging Face into a global circulation node for AI artifacts.
- The community dimension, sustained by the participation of developers, researchers, organizations, and users who publish, comment on, adapt, and reuse models through the platform.
- The reputational dimension, which operates through visible metrics, rankings, likes, downloads, leaderboards, and demos to guide technical attention and turn it into symbolic capital and, in some cases, economic opportunity.
- The normative dimension, exercised through model cards, dataset cards, licenses, restricted repositories, community guidelines, and access-control mechanisms, which produce expectations of conduct and de facto standards.
- The economic dimension, expressed in the combination of open infrastructure and enterprise services, since Hugging Face enables community and scientific uses while also offering organizations solutions for deployment, access control, storage, and services linked to corporate AI adoption.
While
Hugging Face’s public narrative rests on openness, collaboration, and the
democratization of access, its position as a power actor in the AI field shows
it as a platform that distributes practical capacities for understanding,
evaluating, adapting, or challenging models.
Hugging
Face makes visible a central paradox of contemporary AI: openness can be a form
of resistance to concentration, but it can also become a form of global private
infrastructure.
Key Facts
- Hugging Face was founded by Clément Delangue, Julien Chaumond, and Thomas Wolf. The company itself identifies Delangue as CEO, Chaumond as CTO, and Wolf as Chief Science Officer.
- It functions as a central AI hub that hosts widely circulated open models and datasets, as well as Spaces, where demos and machine learning applications are hosted directly on personal or organizational profiles.
- Transformers, one of its most influential libraries, was presented as an open source library designed to facilitate the use of Transformer architectures and pretrained models by researchers, developers, and industrial environments.
- The platform incorporates documentation mechanisms such as model cards and dataset cards, as well as access-control functions, restricted models, download statistics, library integration, and deployment tools.
References
Bommasani,
R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M.
S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card,
D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky,
D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J.,
Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K.,
Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J.,
Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D.,
Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W.,
Krass, M., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee,
T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C.
D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A.,
Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko,
J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance,
E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y.,
Ruiz, C., Ryan, J., Ré, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A.,
Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tramèr, F., Wang, R. E.,
Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia,
M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., &
Liang, P. (2021). On the opportunities and risks of foundation models.
arXiv. https://arxiv.org/pdf/2108.07258
Erfan, M.,
Ryan, A., & Rahman, M. R. (2026). A large-scale measurement of AI Bill
of Materials completeness in Hugging Face models. arXiv. https://arxiv.org/pdf/2607.17242
Gebru, T.,
Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H.,
& Crawford, K. (2021). Datasheets for datasets. Communications of the
ACM, 64(12), 86–92. https://dl.acm.org/doi/epdf/10.1145/3458723
Johnson, D.
G. (2022). Algorithmic accountability in the making. Social Philosophy and
Policy, 38(2), 111–127. https://www.cambridge.org/core/journals/social-philosophy-and-policy/article/algorithmic-accountability-in-the-making/6F3CE994EC96C65392C5374B3CDE3C51
Liang, W.,
Rajani, N., Yang, X., Ozoani, E., Wu, E., Chen, Y., Smith, D. S., & Zou, J.
(2024). What’s documented in AI? Systematic analysis of 32K AI model cards.
arXiv. https://arxiv.org/pdf/2402.05160
Mitchell,
M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer,
E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings
of the Conference on Fairness, Accountability, and Transparency, 220–229. https://dl.acm.org/doi/epdf/10.1145/3287560.3287596
Wolf, T., Debut,
L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf,
R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite,
Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., & Rush,
A. M. (2020). Transformers:
State-of-the-art natural language processing. Proceedings of the 2020
Conference on Empirical Methods in Natural Language Processing: System
Demonstrations, 38–45. https://aclanthology.org/2020.emnlp-demos.6.pdf
Yang, X., Liang,
W., & Zou, J. (2024). Navigating dataset documentations in AI: A large-scale analysis of
dataset cards on Hugging Face. International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2024/file/8c67fc501a50977947c5bebbc39ca8f6-Paper-Conference.pdf
.png)