# The transformer block
Part 9 of the tutorial "tiny-gpt".
Canonical: https://learn.welldun.ai/tiny-gpt/09-block/
Attention gathers and never thinks. Thinking turns out to be a bank of if-then rules you can watch fire — and once one round is visibly not enough, stacking becomes a survival problem with a two-part fix.
---
Keep one picture in mind for this whole part. The model's working material is
**one row of numbers per position** — that position's description. Attention
just rewrote every row as a blend of the rows it chose to read; that was the
gathering. From here on nothing looks at two positions at once. Everything
below happens to each row on its own: it arrives holding gathered evidence,
and it has to leave holding more than evidence.
Because gathering is not thinking. A position that ends the attention layer
holding *"an `a`, shortly after `cat` and `sat`"* has collected the facts and
done nothing with them. A model that only collects is a filing cabinet.
This part builds the piece that thinks, finds out why thinking once is not
enough, and makes the answer to that — stacking — survivable.
## One if at a time
Thinking, at its smallest, is a rule: *if the evidence looks like this, then
conclude that.* The thinking piece is built from exactly that shape.
One **unit** owns two rows of weights — the model's own numbers, shaped by
training like every matrix so far:
- a **pattern**: the evidence it looks for,
- a **contribution**: what it adds to the position's description when it
finds it.
How well the evidence matches the pattern is their dot product — the same
measuring stick attention scored with. The match is then clipped at zero:
```python
fires = max(pattern · evidence, 0)
```
The clip is the **bend** (its standard name is **ReLU**): plotted, a straight
line bent flat at zero. A poor match does not fire a little — it does not
fire at all, and the unit stays out of the way. With pattern `[1, −1]`:
| evidence | `pattern · evidence` | after the bend |
|---|---|---|
| `[0.8, −0.6]` — the pattern | +1.4 | **1.4** |
| `[−0.8, 0.6]` — its opposite | −1.4 | 0 |
| `[0.5, 0.5]` — unrelated | 0.0 | 0 |
And whatever fires, the unit acts on: it adds `fires × contribution` to the
description. An if, followed by its then. Here are three units under your
hand — the first one is the table's:
## A bank of ifs
One rule is not thinking. The **feed-forward** layer is 128 of these units —
four times the model's width, by convention — all reading the same evidence,
each with its own pattern and its own contribution. Stack the 128 patterns as
the columns of one matrix, `W1`, and the 128 contributions as the rows of
another, `W2`, and the whole bank runs in one line:
The widget's three units, written that way:
```
W1 = [[ 1.0, 0.0, -1.0], W2 = [[ 1.0, 0.5],
[-1.0, 1.0, -1.0]] [-0.5, 1.0],
[ 0.5, -1.0]]
```
Walk the table's evidence through. `[0.8, -0.6] @ W1` asks all three ifs at
once: `[1.4, -0.6, -0.2]`. The bend silences the two poor matches,
`[1.4, 0, 0]`, and `@ W2` sums what fired: `1.4 * [1.0, 0.5] = [1.4, 0.7]` —
the widget's own output. At full size the only change is the count:
```python
def feed_forward(z):
return np.maximum(z @ W1, 0) @ W2 # every if, then every fired then
```
The bend is what makes the bank worth having. Delete it and every "if"
becomes "always" — and 128 always-on contributions collapse into a single
matrix: `(x @ W1) @ W2` equals `x @ (W1 @ W2)`, checked to the digit in the
lab. However many linear layers you stack, they are one matrix wearing
costumes. One kink breaks the collapse, and depth starts buying something.
No position sees another anywhere in this — mixing is attention's job,
exclusively:
Attention is the only place positions meet. The feed-forward is each position
alone at its desk, thinking about what it gathered.
## Why once is not enough
One round of gather-and-think lets every position collect facts and draw a
conclusion. What it cannot do is collect anyone else's *conclusions* — they
did not exist until the round was over.
Take the sentence from the attention part: *the cat sat on the mat because it
was soft*. In one round, the position holding `it` can gather `cat` and `mat`
and think its way to a suspicion about which one is meant. Acting on that
suspicion — going back for what was said *about the mat* — needs a second
round of gathering, because only now is there a suspicion to act on.
One round collects facts. Only the next round can collect conclusions.
So the gather-and-think pair gets stacked, conclusions building on
conclusions, one round per storey. And stacking is where the trouble starts.
## Add, never replace
A learned layer has no reason to come out exactly signal-preserving. Suppose
each of twenty-four storeys is merely a little timid — passing on about 0.7 of
what it receives. If every storey *replaces* what it was given, the signal
arrives at size **0.0001**, because 0.7²⁴ is 0.000192. Nothing survives;
one dud storey anywhere would have the same effect.
So a block never replaces. It **adds** what it computed to what it received:
```python
z = z + feed_forward(z)
```
A notebook passed down a line of people, where each may add a margin note and
none may erase. A useless note — all zeros — costs nothing, and the original
is still there at the bottom of the stack. The same lane matters in reverse:
the slope flowing back during training always has the untouched `z` path,
however unhelpful the notes were.
## Tidy after every addition
Adding has its own failure: the numbers only ever accumulate. The same
twenty-four storeys, adding instead of replacing, deliver a signal of size
**108** — nothing lost, but far from any scale the next storey expects. Same
family of trouble as the √width divisor in attention.
The repair is to rewrite each position to a standard size after every
addition: subtract its mean, divide by its spread. Mean 0, spread 1, same
information — this is **layer norm**, and it works on each position's own
numbers only, never across positions.
The full version keeps a learned scale and shift per feature, born as 1 and 0
so they start by doing nothing. Training can then relax the strict standard if
it turns out to be wrong. Nothing trains on this page, so the lab leaves them
out.
## The block
All of it, in the order tidy → gather → add, tidy → think → add:
```python
def block(z):
z = z + multi_head(layer_norm(z), heads)
z = z + feed_forward(layer_norm(z))
return z
```
Shape in equals shape out — the lab runs three blocks deep on the same input
and the shape never moves. That property is the entire trick: anything that
keeps its shape can be stacked as deep as the budget allows, and the depth of
a model is just how many of these sit in a pile.
Real blocks add one more learned matrix after the heads are glued; ours
concatenates straight into the addition. Worth adding when this is rebuilt
with a library — not needed to understand the idea.
## What the stack still cannot know
A row can now cross twenty-four storeys and arrive intact, gathered and
thought about at every one. But nothing in any of it — not the heads, not the
ifs — ever asks *where* a character sits. The next part proves that blindness
with one experiment, fixes it with one table, and assembles the whole model.
## Build it yourself
Six cells: the gathering packed into functions, one if-unit and then the bank
of them, the collapse the bend prevents, the stack that dies by replacing,
the drift that tidying holds at 1.00, and the block — three deep, shape
unchanged.