By Javier Surasky
Generative artificial intelligence can be described as a technology capable of producing answers by exploring a latent space of possibilities, which has social impacts.
Borges’s
metaphor of the Library of Babel is useful for understanding that latent space
because, much like the infinite library in the story, language models can
generate texts that are true, false, plausible, or accidentally correct: thinking about AI through Borges’s story therefore helps us better understand
its hallucinations, the limits of verification, and the difference between
producing language and producing knowledge, while also shedding light on an
extremely dense space of textual, visual, and conceptual options that we do not
get to see, but from which generative models produce their answers.
Within that
space, which is internal to each AI model’s working structure and which we will soon refer to as “latent space,” one of the greatest problems models face when
responding to our queries is finding their way among forms of language that may
be true, false, partial, plausible, absurd, or accidentally correct.
I am
convinced that one of the most useful metaphors for understanding that space
and how it works comes from Borges, and more specifically from his story “The
Library of Babel,” where the Library appears as a complete world, presented to
the narrator as “The universe (which others call the Library)” (Borges, 1984,
p. 465). This makes it a total environment of experience, not an archive
external to the subject moving through it.
From the Library of Babel to AI’s Latent Space
Borges’s
Library belongs to literature: it takes the form of a building, and its
inhabitants move through galleries, search for books, read signs, and try to
find some principle of order.
The latent
space of an AI belongs to a different environment, since its architecture is
computational: it can be understood as an internal zone of the model where it organizes patterns learned during training, a mathematical space in
which words, images, concepts, and relations are transformed into vectors,
numerical coordinates that allow the AI to locate and compare meanings within
this latent space and group them according to proximities, differences, andassociations.
Recent work
on language models has increasingly focused on latent space as an internal
level of processing and representation. This approach makes it possible to see
that relevant processes in AI models take place on a plane that is “invisible”to the user, hidden behind the curtain of the words generated, and configured
as a multidimensional and continuous space where relations, trajectories, and
possibilities for generating answers are organized (Yu et al., 2026). These
answers are no longer produced as repetitions of stored phrases, but from the
possibilities made available to the system by what it has learned: terms that
tend to appear together, language structures that are more likely in certain
contexts, relations among lexical units, among many others.
If Borges’s
Library is a building that organizes textual possibilities, latent space is a
computational architecture that organizes generative possibilities. They
therefore share a kind of structured interiority that exceeds the human
capacity to traverse it.
In Borges,
the books are laid out in a total combinatorial order; in AI, answers appear
during the generative act. Even so, the experience can seem similar: a question
is asked, and a text appears that seems to have been found in some region of
meaning within a multidimensional universe that escapes our full understanding.
Why Generative AI Resembles the Library of Babel
As we said,
latent space can be thought of as a generative Library of Babel insofar as both
architectures link possibility, search, truth, error, and orientation.
Borges is
explicit in saying that the library’s shelves contain “all possible
combinations” of the symbols (Borges, 1984, p. 467), a totality that creates a
paradox: if everything is there, then truth is there, along with its false
versions, correct refutations, and erroneous refutations. As with an answer
generated by AI, the mere appearance of a text is no guarantee of truth. The
form of meaning can separate from meaning, and the form of truth can separate
from truth.
The Library
contains the correct catalog of its contents, but also “thousands and thousands
of false catalogs” (Borges, 1984, p. 467). That image is central to thinking
about generative AI: a system that can produce convincing but erroneous
answers, nonexistent citations, or syntheses that mix real facts with fragile
inferences, not because truth is unavailable to the system or because the
system cannot say that it does not find it, but because truth has become
dissociated from the question being answered, lost in a space of millions of
vectors, true and false, applicable and inapplicable.
The
literature on hallucinations in large language models, for example, describes a
problem close to that intuition: LLMs can produce “seemingly plausible yet
factually unsupported content” (Huang et al., 2025, p. 1), even when doing so
takes the form of knowledge. Content becomes dissociated from reality in the
wall-less labyrinth of latent space.
In Borges,
finding a false catalog inside the Library promises orientation and leads to
being lost, while in AI, an answer can present itself as an explanation and
rest on weak associations, erroneous data, or invalid inferences. In both
cases, there is falsehood dressed up as knowledge.
AI Hallucinations, Accidental Truth, and Verification
There is an
even more unsettling possibility: accidental truth.
A reader of
the Library may happen upon a true book, an exact biography, a correct
prediction, but that discovery does not prove that the reader has understood
the order of the Library. The reader may have arrived at the true book through
a false or random path.
The same
happens with AI, where a model can produce a correct answer for the wrong
reasons: a statistical regularity coincides with the true datum, a superficial
association leads to the right result, the prompt activates a dense area of
valid information located within latent space.
The result
may be useful, but the process will be unreliable. This forces us to
distinguish between truth and knowledge, because a correct answer is not enough
when we do not know why it is correct. To speak of knowledge, truth needs
justification beyond the element of chance from which mathematics and deep
learning can never fully free themselves.
In
response, external criteria of verification become relevant, because the textcannot be its own tribunal of truth.
A randomly
found book that asserts a historical fact needs to be checked against something
else: documents, archives, testimonies, records, or evidence. And an AI answer
that states a date, cites a resolution, attributes a phrase, or explains a rule
requires verifiable sources.
Here an
important difference appears between the structure of the Library of Babel and
AI: the former presents itself as a closed universe that contains everything,
including true and false versions of the criteria for evaluating the books it
contains, giving rise to an almost metaphysical self-enclosure.
AI, by
contrast, does allow us to step outside the generated text to review a source,
consult a database, verify a law, and so on. This opens a way out of the
generated text, but one that is not free from contamination: if the information ecosystem fills up with synthetic texts, fabricated citations, recycled content, and false references, it encloses us in the same space, now digital,
proposed by the Library.
Explainability: The Impossible Catalog of Language Models
The
comparison with Borges also sheds light on the debate over explainability. The
narrator tells us that he has journeyed in search of a book, perhaps the
“catalog of catalogs” (Borges, 1984, p. 465), a search that closely resembles
the contemporary impulse to open AI’s black box, which Zhao et al. describe,
when speaking of LLMs, as complex black-box systems: “their inner working
mechanisms are opaque” (Zhao et al., 2024, p. 2).
And yet,
understanding why a system produced an answer, what patterns it activated, what
biases it carries, and what trajectories are reliable is necessary. But there
are limits that seem insurmountable: it is unlikely—because I leave here a
space for chance and doubt—that we will find the “total catalog” that contains
the solution to every possible case of lack of transparency.
We need
more reliable systems, fewer errors, better sources, greater recognition of
uncertainty, the capacity to abstain, traceability, and auditing. But for an AI
to be capable of always delivering true answers assumes that truth is available
as a stable object, and this is not the case.
Conclusion: More Language Does Not Mean More Knowledge
AI runs the
risk of becoming an oracle when it is credited with the capacity to resolve the
relationship between language, world, and truth.
There are
empirical truths that allow for relatively clear verification, and there are
interpretive questions involving conceptual frameworks, values, interests, and
contexts. Treating both planes as equivalent leads to error.
The Library
of Babel warns of the disproportion between information, meaning, and truth by
stressing that textual abundance can both orient and lead astray or, in
Borges’s words, that the certainty that everything has been written “annuls us
or turns us into phantoms” (Borges, 1984, p. 470). Borges did not anticipate
embeddings, weights, or Transformers, but he did play, in literary terms, with
the idea of an architecture in which textual abundance destroys the illusion
that more language, or more data, is equivalent to more knowledge.
Latent
space thus becomes a Library of Babel without visible shelves: an architecture
of possibilities that delivers texts, but not necessarily truths or knowledge.
Seen this
way, generative AI confronts us with a new condition of reading: we are
surrounded by texts that seem to be the result of a search through an immense
field of knowledge and that, in reality, have been generated by mathematical
operations that cannot compute truth value or knowledge value. They are texts
that have an explanatory appearance when they are informed approximations, and
that are taken as true when they require verification.
If Borges
turns the Library of Babel into a literary nightmare about the excess of
meaning, generative AI can turn that nightmare into a digital infrastructure of
everyday use, one we turn to willingly, and at times under illusion.
References
Borges, J.
L. (1984). La biblioteca de Babel. In Obras completas 1923–1972 (pp.
465–471). Emecé. (Original
work published 1941)
Huang, L.,
Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X.,
Qin, B., & Liu, T. (2025). A survey on hallucination in large language
models: Principles, taxonomy, challenges, and open questions. ACM
Transactions on Information Systems. https://doi.org/10.1145/3703155
Yu, X.,
Chen, Z., He, Y., Fu, T., Yang, C., Xu, C., Ma, Y., Hu, X., Cao, Z., Xu, J.,
Zhang, G., Tao, J., Zhang, J., Ma, S., Feng, K., Huang, H., Li, Y., Chen, R.,
Wang, H., Wu, C., et al. (2026). The latent space: Foundation, evolution,
mechanism, ability, and outlook. arXiv. https://doi.org/10.48550/arXiv.2604.02029
Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H.,
Wang, S., Yin, D., & Du, M. (2024). Explainability for large language models: A survey. ACM Transactions
on Intelligent Systems and Technology, 15(2), Article 20. https://doi.org/10.1145/3639372
