You Already Know 70% of How LLMs Work
A guided tour for ML students who want to see how their basics connect to large language models

If you've trained a logistic regression or a small neural network, you already know about 70% of what's inside an LLM. An LLM isn't a new species of thing. It's the same ML toolkit with a few very specific upgrades.
This post climbs from what you know to GPT-style models, one step at a time. At every step, the question is the same: what problem does this fix in the previous step?
Step 1: Start with a softmax classifier
You've seen this pipeline:
features x → linear layer → logits → softmax → probabilities → cross-entropy loss → gradient descent
The final layer of an LLM is exactly this. The only twist: the "classes" are the tokens in the vocabulary (say, 50,000 of them), and the question being classified is:
Given the text so far, what is the next token?
So an LLM is a 50,000-class classifier. Everything else in the architecture exists to answer one question: what should x be, and what function turns raw text into x?
Step 2: The simplest possible language model
Take the previous token, one-hot encode it, and feed it to softmax regression that predicts the next token. That's a bigram model.
- It learns that "New" is often followed by "York".
- It sees only one token of context, so it's weak.
- But it's a complete, working language model you can write in about 20 lines.
The weight matrix is vocab × vocab. Each row says: "given this word, here's the distribution over what comes next." Keep that in mind, because the rest of the story is about replacing this crude lookup with something smarter.
Step 3: Embeddings, learned features
One-hot vectors treat "cat" and "dog" as completely unrelated. Instead, give every token a dense learned vector (say 256 dimensions) from a lookup table.
Under the hood this is just a weight matrix multiplied by a one-hot vector, trained by ordinary backprop. Tokens that behave similarly end up with similar vectors, because that helps prediction.
Now extend the context: concatenate the embeddings of the last 8 tokens, pass them through an MLP, then softmax. This is a neural language model (Bengio et al., 2003). It's literally an MLP you already know how to build.
Its flaws set up everything that follows:
| Problem | Why it hurts |
|---|---|
| Fixed window | The 9th token back is invisible |
| No weight sharing across positions | "the cat" at positions 1-2 and 5-6 is learned separately |
Step 4: Variable context, and why RNNs aren't enough
An RNN processes tokens one at a time and carries a hidden state summarizing everything so far. It handles any length, but:
- The entire past is squeezed into one fixed-size vector (a bottleneck).
- Gradients struggle to flow across long distances.
- It's inherently sequential, so you can't parallelize training across the sequence.
The real question becomes:
Can each token directly access whichever earlier tokens are relevant, without squeezing through a bottleneck?
Step 5: Attention, the key idea
Think of attention as a soft dictionary lookup (or kernel regression, if you've seen that).
For each token, the network produces three vectors via learned linear maps:
| Vector | Role | Intuition |
|---|---|---|
| Query (Q) | What am I looking for? | The question this token asks |
| Key (K) | What do I contain? | The label other tokens match against |
| Value (V) | What do I hand over if picked? | The actual information passed along |
The recipe:
- Dot the current token's query with every earlier token's key → scores.
- Softmax the scores → weights that sum to 1.
- Output = weighted average of the values.
In compact form:
So each token builds its new representation as a content-dependent weighted average of other tokens' information. In "The animal didn't cross the street because it was tired," the query for "it" can learn to match the key for "animal."
Two practical details:
- Causal mask: a token may not look at the future. Otherwise "predict the next token" would be cheating.
- Positional information: attention ignores order by itself, so position signals are added to the embeddings.
Compare with the earlier problems:
| Earlier problem | How attention fixes it |
|---|---|
| Fixed window | Any earlier token is reachable |
| Long-distance dependencies | Any distance is just one hop |
| No weight sharing | Same Q/K/V weights at every position |
| Sequential training | Everything is matrix multiplication, so it parallelizes on GPUs |
Step 6: The transformer block
A transformer block stacks two sublayers, each wrapped in a residual connection (the ResNet trick) plus normalization:
- Attention: tokens exchange information.
- MLP: each token privately processes what it gathered. This is the plain feedforward net you already know.
Repeat this block 12, 32, or 100+ times. Roughly speaking, early layers capture local patterns like syntax, and later layers capture more abstract relationships. At the very top, the last token's vector goes through the softmax layer from Step 1.
The full GPT stack
Step 7: Training is nothing new
| Ingredient | In an LLM |
|---|---|
| Loss | Cross-entropy between predicted distribution and the true next token |
| Optimizer | Adam (a gradient descent variant) |
| Labels | Free: the text supervises itself |
| Regularization / validation | Same overfitting intuitions apply |
The "free labels" point is the quiet superpower. Every position in a document is a training example: given tokens 1..t, predict token t+1. No human labeling is needed, which is why you can train on a huge slice of the internet. This is called self-supervised learning.
One sentence yields many training examples.
Step 8: Generation
To produce text:
- Run the model and get a probability distribution over the vocabulary.
- Sample a token from it.
- Append it to the input.
- Repeat.
Temperature simply rescales the logits before softmax: low temperature makes outputs more deterministic, high temperature makes them more random.
Step 9: Why does such a simple objective produce something smart?
Predicting the next token well on diverse text forces the model to learn grammar, facts, code semantics, and patterns of reasoning, because all of them reduce prediction error. The loss is simple, but what minimizes it is rich.
Empirically, loss improves smoothly and predictably as you scale parameters, data, and compute. These are the scaling laws.
Chat assistants add one more stage on top: continue training the pretrained model on instruction-following examples and human preference feedback. Same architecture, same gradient descent, just different data and a different loss.
The mapping, in one view
| What you already know | Its LLM counterpart |
|---|---|
| Softmax classifier | Output layer over the vocabulary |
| One-hot → learned features | Token embeddings |
| MLP | Feedforward sublayer in each block |
| Weight sharing (CNN) | Same attention weights at every position |
| ResNet skip connections | Residual connections in every block |
| Cross-entropy + SGD/Adam | Pretraining, unchanged |
| Supervised labels | Next token, free from the text itself |
| Kernel regression / soft lookup | Attention |
How to make it stick: build the ladder
Reading isn't enough. Build it yourself, in this order:
- Bigram model on a text file (softmax over the vocabulary).
- Embeddings + MLP over a fixed window.
- Replace the window with a single self-attention layer.
- Stack blocks with residuals and layer norm.
Andrej Karpathy's makemore and Let's build GPT videos follow precisely this path. By the end, you'll have a tiny GPT that is structurally the same as the giants. The difference between yours and the frontier models is mostly scale.
Key takeaways
- An LLM is a huge classifier whose classes are tokens and whose features come from a transformer.
- Attention replaces fixed windows and RNN bottlenecks with content-based, parallelizable lookup.
- Training is ordinary ML: cross-entropy, Adam, backprop. The twist is that the labels come for free.
- Capabilities come from scale + diverse data, not from exotic new math.
If you understand logistic regression, MLPs, and backprop, you're closer to understanding LLMs than it probably feels.
Where to go next
If parts of this post felt shaky, like softmax, cross-entropy, or backpropagation, that's a sign to strengthen your foundations, because everything in an LLM is built on them. A good place to start is the Machine Learning Fundamentals course. Work through the basics there, then come back to the ladder in this article: build your bigram model, add embeddings, write a single attention layer, and stack your first transformer block. The stronger your fundamentals, the less mysterious the giants of the field will look.
Machine Learning Fundamentals (Free to enroll)
Comprehensive Machine Learning Fundamentals (with Python) certification program covering core algorithms, data preparation, predictive modeling, model validation, performance optimization, and real-world machine learning applications.