What is an LLM?
And why should I care?


For many people, artificial intelligence is something rather "new". Ask someone when AI appeared, and there's a fair chance they would point to the early 2020s.
Instead, the modern field is generally traced back to a summer workshop in 1956, and the term itself appeared a year earlier, in the proposal for it.1 At the time of writing, that makes the discipline 70 years old.

Surprising, right? The reason for the confusion is less surprising: you might be conflating AI with chatbots such as ChatGPT, Claude, or Gemini. These products are built around LLMs (short for large language models).
Before we go any further, let's make a distinction:
So, what are they?
A text-based large language model is trained around one deceptively simple task: given some text, predict what token is likely to come next.
Repeat that prediction over and over and you get sentences, answers, code, and much more. Modern chatbots add further training, instructions, search, and tools around that core.

Let's take the name apart:
- Language: it works with text, chopped into chunks called tokens. A token is roughly a word, or a piece of one.
- Model: it's a mathematical function with a lot of adjustable numbers inside it, called parameters. Humans design the overall structure, but the values of those parameters are learned during training rather than set by hand. The particular design used by most modern LLMs is called a transformer, and it dates from 2017.3
- Large: that's how many parameters it has. Meta's openly published Llama 3.1 has 405 billion of them.4

And here's how one gets made. You show it an enormous and varied collection of text. You hide the next token. You ask it to guess. When it guesses wrong, the training process adjusts its parameters a tiny bit in the direction that would have made it less wrong. Then you do that again. And again. Trillions of times.
Now, the part that matters
To predict the next token well across an enormous range of human writing, the model cannot rely only on memorizing a handful of phrases. Completing a geography question rewards geographical knowledge. Continuing a proof rewards mathematical patterns. Finishing a joke rewards a feel for how jokes work. Completing a sentence in Spanish rewards knowledge of Spanish.
Next-token prediction is a Trojan horse. Getting exceptionally good at it appears to produce internal machinery useful for a great many other tasks. That machinery is learned during training, one adjustment at a time, rather than written line by line by a human.
One last ingredient. A freshly pretrained model is mainly good at continuing text. Ask it a question and it might respond with more questions, because that is a plausible continuation. A further phase of training (call it the finishing school) teaches it to follow instructions, answer questions, and refuse certain requests.6 That post-training helps turn a text-completion model into a chatbot.
And why should I care?
Three reasons, in ascending order of inconvenience.
1. You're already encountering them. They may be in your search results, your inbox, your writing tools, or the summary sitting on top of the page you were trying to read. You opted in to very little of that.
2. They are most convincing exactly where you are least equipped to check them. Ask about something you know well and you'll catch the errors instantly. Ask about something you don't (a legal question, a medication, a country's history) and the same confident tone reads as authority. The tone is free. It costs the model nothing to produce.
3. Knowing what it actually does changes how you use it. A machine built to produce plausible text can be genuinely excellent at shaping material you give it: summarize this, rewrite this, translate this, find possible holes in my argument. It is far shakier as an unaided oracle you ask for facts. Many disappointing uses get this backwards.
What we still don't know
No one should expect to grasp all of this after one short article. Researchers around the world are working at a breakneck pace to understand what these systems learn, which abilities hold up, which ones fail, and why.
In short: we know how these systems are built and trained, but we do not fully understand how their learned internal machinery produces many of their abilities and failures. The people who build them say so plainly: models "arrive inscrutable to us, the model's developers."8 It's like walking into a dark cave: you know what a cave is, you know how to explore one, but you don't yet know how big it is, where it ends, or what lives in it.
Which brings me to this blog
Predictions that this technology will save the world, or end it, are easy to write and mostly useless in my humble opinion.
What I'll do here is take one thing at a time and explain how it actually works, plainly, in both English and Spanish, with an honest note wherever the truthful answer is "nobody knows yet."
You don't need to be a math whiz to understand what a model can do, what it can't, and when it's best to be skeptical of its answers.
Footnotes
-
J. McCarthy, M. L. Minsky, N. Rochester and C. E. Shannon, A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, dated 31 August 1955; the workshop ran the following summer. Reprinted in AI Magazine 27(4), 2006. https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/1904 . Dartmouth's own account of the event is at https://home.dartmouth.edu/about/artificial-intelligence-ai-coined-dartmouth
-
Grace Solomonoff, The Meeting of the Minds That Launched AI, IEEE Spectrum. https://spectrum.ieee.org/dartmouth-ai-workshop
-
A. Vaswani et al., Attention Is All You Need, 2017. https://arxiv.org/abs/1706.03762
-
Meta AI, Introducing Llama 3.1: Our most capable models to date, 2024. The 405B weights are published openly, which is why the number is quotable at all; the leading commercial labs no longer disclose theirs. https://ai.meta.com/blog/meta-llama-3-1/
-
Exact token splits vary from one model to another. Vocabularies typically run from roughly 50,000 to 200,000 tokens depending on the model. You can paste your own text into OpenAI's tokenizer and watch it split: https://platform.openai.com/tokenizer . Background at https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
-
L. Ouyang et al., Training language models to follow instructions with human feedback, 2022. This paper describes one influential post-training technique, usually shortened to RLHF. Modern systems may use RLHF, other techniques, or a combination of them. https://arxiv.org/abs/2203.02155
-
L. Huang et al., A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, 2023. https://arxiv.org/abs/2311.05232
-
Anthropic, Tracing the thoughts of a large language model, 2025. https://www.anthropic.com/research/tracing-thoughts-language-model