rumblr Work in progressWIP

● The AI Primer · Lesson 0 · Before you begin

Math notation, from zero

This lesson covers Every symbol in an ML formula, as a short loop

Free lesson 26 min5 figures and diagrams12 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Σ is a loop that adds; Π is a loop that multiplies.
  2. A dot product multiplies matching entries and adds them up; it measures how much two vectors agree.
  3. A matrix multiply is a grid of dot products, and it is most of the work a neural network does.
  4. log undoes e, and it turns products into sums, which is why models work with log-probabilities.
  5. The gradient points uphill; training steps the other way.

Level 2

How it works, from scratch

Machine learning papers look impenetrable mostly because of notation: Greek letters, big sigmas, little superscript Ts. Almost every symbol is shorthand for a short loop you could write in a few lines of Python. This lesson writes each one out as that loop, so that when a formula appears elsewhere in this primer you can read it aloud.

Every function below is deliberately written the slow, obvious way, with plain Python loops, so the code is the definition. The tests check each one against NumPy, which does the same thing fast.

Chapter 1

Lists and tables of numbers: vectors and matrices

Everyday picture A vector is a list of numbers, like a shopping receipt: (apples 3, bread 1, milk 2). A matrix is a table of numbers, like a spreadsheet with rows and columns.

Tiny example is a vector with 3 entries. Its entries are written with a small subscript: , , . A matrix with 2 rows and 3 columns has shape , and means "row 2, column 3".

You see Say Python
(lowercase, sometimes bold x) "the vector x" x = [3, 1, 2]
"x sub i": the i-th entry x[i - 1] (maths counts from 1, Python from 0)
(uppercase) "the matrix A" A = [[1, 2, 3], [4, 5, 6]]
or row i, column j A[i - 1][j - 1]
"the set of all lists of d real numbers" any list of d floats
"x is a list of 768 numbers" len(x) == 768
shape: n rows, d columns np.zeros((n, d))

In practice, a word's embedding is a vector, a batch of embeddings is a matrix (one row per word), and a model's weights are mostly matrices.

Chapter 2

Σ (capital sigma): add them all up

Run the loop yourself first; the symbol will read as what it is.

Everyday picture Totting up a receipt.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
"sum": add up everything that follows
(below) start a counter called at 1
(above) stop after the counter reaches
the thing being added on each step
"and so on, following the same pattern"

In words: "for i from 1 to n, add up x sub i."

With the numbers: . In code, Σ is a for loop with a running total (summation below).

Level 3: in Python
x = [1, 2, 3, 4]
total = 0
# Σ: visit each x_i, from i = 1 to n ...
for x_i in x:
    # ... and add it to a running total
    total += x_i
total  # → 10
# Python's built-in sum is the same loop
sum(x)  # → 10

Π (capital pi) is the same idea with multiplication: . It shows up whenever independent chances combine. For example, ten steps that each succeed 95% of the time all succeed with probability . This single fact explains why long chains of AI agent steps fail so often (primer.agents.planning).

In Python:

x = [1, 2, 3, 4]
product = 1
# Π: the same loop, multiplying instead of adding
for x_i in x:
    product *= x_i
product  # → 24
# ten steps that each succeed 95% of the time
round(0.95 ** 10, 2)  # → 0.6

In code: product is Π written the same way: a loop with a running product that starts at 1.

Chapter 3

The dot product: how much two lists agree

Everyday picture A recipe needs 2 eggs, 3 cups of flour and 1 cup of sugar, and eggs cost $1, flour $0.50 and sugar $2 per unit. The total cost is 2·1 + 3·0.5 + 1·2 = $5.50: multiply matching items, then add. That is a dot product.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
two vectors with the same number of entries,
the -th entries multiplied together
"dot": the whole multiply-then-add operation

In words: "multiply the lists position by position and add up the products."

With the numbers: .

Level 3: in Python
a = [1, 2]
b = [3, 0.5]
# Σ over i of a_i times b_i
sum(a_i * b_i for a_i, b_i in zip(a, b))  # → 4.0

Geometrically, the dot product is large when two arrows point the same way, zero when they are at right angles, and negative when they point apart. That is why it is the standard similarity score for attention and for embeddings.

Figure 1 · Diagram

Reading it: each pair of matching positions meets in a multiply box, first entries with first entries and second with second. Every product then flows into one addition. A dot product is always this shape: many multiplications feeding one sum, however long the lists get.

In code: dot pairs up matching entries, multiplies them and adds the products with summation, refusing lists of different lengths.

Chapter 4

‖x‖: the length of a vector

Everyday picture Walk 3 blocks east and 4 blocks north. As the crow flies, you are 5 blocks from where you started (Pythagoras).

Level 3: the formula and its symbols

Symbols

Symbol Meaning
the norm (length) of , also written
the -th entry squared ()
square root

In words: "square every entry, add them up, and take the square root."

With the numbers: and .

Level 3: in Python
import math
def length(x):
    # √ of Σ x_i²
    return math.sqrt(sum(x_i ** 2 for x_i in x))
length([3, 4])  # → 5.0
length([1, 2, 2])  # → 3.0

Dividing a vector by its length gives a unit vector of length 1 that points the same way. Embedding systems do this constantly, because then the dot product measures only direction (see primer.ml.embeddings.similarity).

In code: norm is the square root of a vector's dot product with itself.

Chapter 5

Matrix multiply and transpose

Transpose, written ("A transpose"), flips a table on its diagonal so rows become columns: a matrix becomes .

Matrix multiply, written with nothing in between, is a whole grid of dot products. The cell in row , column of the answer is row of dotted with column of .

Level 3: the formula and its symbols

Symbols

Symbol Meaning
an matrix
an matrix; its row count must equal 's column count,
the answer's cell at row , column ; the answer is
the counter that walks along row of and down column of together

In words: "to fill cell (i, j), walk across row i of A and down column j of B at the same pace, multiplying and adding."

With the numbers: $\begin{pmatrix}1&2\3&4\end{pmatrix}\begin{pmatrix}5&6\7&8\end{pmatrix} = \begin{pmatrix}1\cdot5+2\cdot7 & 1\cdot6+2\cdot8\ 3\cdot5+4\cdot7 & 3\cdot6+4\cdot8\end{pmatrix} = \begin{pmatrix}19&22\43&50\end{pmatrix}$.

Level 3: in Python
A = [[1, 2], [3, 4]]
B = [[5, 6], [7, 8]]
# A has m columns, B has m rows: they must match
m = len(B)
# (AB)_ij = Σ_k A_ik B_kj
AB = [[sum(A[i][k] * B[k][j] for k in range(m))
       for j in range(len(B[0]))]
      for i in range(len(A))]
AB  # → [[19, 22], [43, 50]]
# the transpose: rows become columns
[list(column) for column in zip(*A)]  # → [[1, 3], [2, 4]]

Figure 2 · Diagram

Reading it: a matrix multiply is nothing but this box, repeated for every (row, column) pair. The shape rule falls out: the row of A and the column of B must be the same length, or there is nothing to pair up. Almost all the computation in a neural network is this one operation, which is why GPUs, built to do thousands of multiply-adds at once, are the hardware of AI.

In code: transpose turns columns into rows; matmul transposes B once, then fills each cell with one dot of a row of A and a column of B.

Chapter 6

e and log: growth, and its undo button

Everyday picture Money in an account that compounds continuously at 100% a year grows by a factor of e ≈ 2.718 in one year. is that growth run for years. The natural logarithm (often just in ML papers) answers the reverse question: how many years of that growth turn 1 into ?

You see Say Example
or "e to the x" , , ,
or "log of y" , ,

Two facts carry most of machine learning:

  1. is always positive and grows fast, which is why softmax uses it.
  2. . Logs turn multiplying into adding. The probability of a whole sentence is a product of thousands of small numbers, which would underflow to 0 on a computer. Its log is a sum of manageable negative numbers. This is why models work with "log-probabilities", and why the standard training loss is (see primer.ml.losses).

Figure 3 · Chart

−2 0 2 4 6 8 −4 −2 0 2 4 6 8 e and its undo button, ln ln 0.9 = -0.11 ln 0.01 = -4.61 e x   ( a l w a y s   >   0 ,   g r o w s   f a s t ) l n x e   ( u n d o e s   ) x mirror line y = x

e to the x stays above zero and climbs steeply; ln x mirrors it across y = x and plunges near 0, so ln 0.01 = -4.61 while ln 0.9 = -0.11

Reading it: the blue curve is always above zero and climbs steeply; every step of 1 to the right multiplies its height by 2.718. The red curve is the same curve mirrored across the dashed diagonal, because it undoes . It only exists for positive inputs, and it dives towards −∞ as its input approaches 0. That dive is the "confidently wrong" penalty in cross-entropy: , far more than .

The second fact above, happening in a computer's own numbers.

Chapter 7

softmax and argmax

softmax turns a list of scores into shares that are positive and sum to

  1. It is covered step by step, with its symbols decoded, in primer.ml.attention:
Level 3: the formula and its symbols

With the numbers: softmax(2.0, 1.0, 0.5) = (0.63, 0.23, 0.14).

Level 3: in Python
import math
z = [2.0, 1.0, 0.5]
# e^(z_i) for each score
exps = [math.exp(z_i) for z_i in z]
# Σ_j e^(z_j)
total = sum(exps)
# each share of the total
[round(e / total, 2) for e in exps]  # → [0.63, 0.23, 0.14]
# argmax: the position of the largest
max(range(len(z)), key=lambda i: z[i])  # → 0

argmax is simpler: it answers "which position holds the largest value?", not "what is the largest value?". argmax(0.1, 7.0, 3.0) = position 2 (index 1 in Python). "Greedy decoding" in a language model is picking the argmax token every step.

In code: softmax subtracts the largest score before exponentiating, so the exponentials cannot overflow; argmax walks the list and remembers the position of the biggest value.

Chapter 8

Mean, variance, standard deviation: the middle and the spread

Everyday picture Two classes both average 70% on a test. In one, everyone scored 68 to 72; in the other, scores ran from 30 to 100. The mean is the same; the spread is not.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
(mu) the mean: the average
(sigma squared) the variance: the average squared distance from the mean; also written
(sigma) the standard deviation: the typical distance from the mean, in the data's own units
"add them up and divide by how many": an average

In words: "the mean is the average; the variance is the average squared distance from the mean; the standard deviation is its square root."

With the numbers: for (2, 4, 4, 4, 5, 5, 7, 9): μ = 40 / 8 = 5; the squared distances are (9, 1, 1, 1, 0, 0, 4, 16), which sum to 32, so σ² = 32 / 8 = 4 and σ = 2.

Level 3: in Python
import math
x = [2, 4, 4, 4, 5, 5, 7, 9]
n = len(x)
# μ = (1/n) Σ x_i
mu = sum(x) / n
mu  # → 5.0
# the squared distances from μ
[(x_i - mu) ** 2 for x_i in x]  # → [9.0, 1.0, 1.0, 1.0, 0.0, 0.0, 4.0, 16.0]
# σ² = their average
variance = sum((x_i - mu) ** 2 for x_i in x) / n
# σ is its square root
variance, math.sqrt(variance)  # → (4.0, 2.0)

Spread matters constantly in neural networks. If numbers flowing through a network spread out layer after layer, training blows up; normalization layers and careful initialization exist to hold σ near 1 (see primer.ml.deep_nets and the √d_k in primer.ml.attention).

Figure 4 · Chart

−10.0 −7.5 −5.0 −2.5 0.0 2.5 5.0 7.5 10.0 value 0 100 200 300 400 500 how many samples Same mean (μ = 0), different spread σ = 1 σ = 3

Two histograms share the mean 0, but the sigma 1 samples pile up tall and narrow while the sigma 3 samples spread about three times as wide

Reading it: both histograms are centred on the same mean (the dashed line), but the blue one is tall and narrow (σ = 1) while the red one is low and wide (σ = 3). The standard deviation is roughly how far from the dashed line a typical sample lands. About two thirds of samples fall within one σ of the mean.

In code: mean, variance and std are the three formulas above, each a short loop built on summation.

Chapter 9

Derivatives and gradients: which way is downhill?

Feel the slope before the formula for it.

Everyday picture You're on a hillside in thick fog and want to reach the valley. You can't see it, but you can feel the slope under your feet. Take a small step in the steepest downhill direction, feel again, and repeat. That is how every neural network is trained (gradient descent).

The derivative of a function at a point is its slope there: how much the output changes per tiny nudge of the input.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
a function: put in, get a number out
or the derivative: the slope of at ("dee f dee x")
a tiny nudge, like 0.00001
"approximately equal": exact as shrinks to 0

In words: "nudge the input a hair up and a hair down, and see how much the output changes per unit of nudge."

With the numbers: for at : (3.00001² − 2.99999²) / 0.00002 = 6. The slope of at 3 is 6 (the rule is ).

Level 3: in Python
def f(x):
    return x ** 2
h = 0.00001
# rise over run across a tiny step
round((f(3 + h) - f(3 - h)) / (2 * h), 6)  # → 6.0

With many inputs, nudge each one separately. The slope in each direction is a partial derivative, written (the curly ∂ just means "only this input moves, the others stay fixed"). Collect them into a vector and you have the gradient, written ("nabla f" or "grad f"). It points in the steepest uphill direction, so training steps the opposite way:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
(theta) all the model's adjustable numbers (its parameters or weights)
the loss: one number measuring how wrong the model is
the gradient of the loss: for each weight, how much the loss rises if that weight is nudged up
(eta) the learning rate: how big a step to take, e.g. 0.001

In words: "move every weight a small step in the direction that lowers the loss the fastest."

With the numbers: for the bowl at (1, 2), the gradient is (2, 4), pointing uphill away from the bottom at (0, 0). With η = 0.1, the step goes to (1 − 0.2, 2 − 0.4) = (0.8, 1.6), closer to the bottom.

Level 3: in Python
# (x, y)
theta = [1.0, 2.0]
# ∇L for L = x² + y²: each slope is 2 times the value
gradient = [2 * theta[0], 2 * theta[1]]
gradient  # → [2.0, 4.0]
eta = 0.1
# θ_new = θ − η ∇L
[t - eta * g for t, g in zip(theta, gradient)]  # → [0.8, 1.6]

Figure 5 · Chart

−2 −1 0 1 2 x −2 −1 0 1 2 y L o s s   :   a r r o w s   p o i n t   d o w n h i l l x y 2 2 + gradient descent from (1, 2)

Arrows on circular contours point straight at the centre, longest on the steep rim; descent from (1, 2) takes big steps first, then ever smaller ones

Reading it: the rings are contour lines, as on a hiking map: every point on a ring has the same loss, and the bottom of the bowl is the centre. The arrows show the negative gradient at several spots. They always cross the rings at right angles, pointing straight downhill, and they are longer where the slope is steeper. The dotted path is gradient descent from (1, 2): big steps on the steep outer slope, shrinking steps as the ground flattens near the bottom.

The chain rule says how slopes combine when functions are chained: if , then . Multiply the slopes along the chain. Backpropagation is the chain rule applied backwards through every layer of a network, reusing work as it goes (see primer.ml.neural_net).

In code: derivative measures a slope by nudging the input a hair up and a hair down; gradient does that for each input in turn and collects the slopes into one list.

Chapter 10

Probability notation

You see Say Meaning
or "probability of A" a number from 0 (never) to 1 (certain)
"probability of A given B" the chance of A once you know B happened
"probability of word t given the words before it" what a language model computes, one word at a time
"expected value of x" the average you'd get over many tries
"x is drawn from a normal distribution with mean 0 and variance 1" rng.standard_normal()

Chapter 11

Big-O: how cost grows

, "order n squared", describes how cost grows as the input grows, ignoring constant factors. If doubles, an cost doubles and an cost quadruples. Attention is in sequence length, which is why long context is expensive (primer.ml.attention).

Chapter 12

Greek letters you'll meet

Letter Name Usually means
α alpha a mixing weight, or a learning rate
β beta momentum coefficients in optimizers (β₁, β₂); a strength knob in DPO
γ, β gamma, beta the learned scale and shift in normalization layers
δ delta a small change, or an error signal in backprop
ε epsilon a tiny number added to avoid dividing by zero
η eta the learning rate
θ theta all the model's parameters
λ lambda the strength of a penalty (weight decay)
μ mu a mean
σ sigma a standard deviation, or the sigmoid function
τ tau a temperature (sharpness of softmax)
∇ nabla the gradient
∂ "partial" a partial derivative

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Read aloud and evaluate it.Think it through, then reveal

"The sum, for i from 1 to 3, of i squared": 1 + 4 + 9 = 14.

Question 2What is (2, −1, 3) · (1, 4, 0), and what does its sign tell you?Think it through, then reveal

2 − 4 + 0 = −2. It's negative, so the vectors point somewhat apart.

Question 3A is 4 × 3 and B is 3 × 5. What shape is AB? What about BA?Think it through, then reveal

AB is 4 × 5. BA is not defined: B's 5 columns don't match A's 4 rows.

Question 4Why do language models add log-probabilities instead of multiplying probabilities?Think it through, then reveal

A product of thousands of numbers below 1 underflows to 0 in floating point. log(a·b) = log a + log b, so the product becomes a sum of moderate negative numbers.

Question 5What does ∇L tell you, and which way does training move?Think it through, then reveal

For every weight, how fast the loss rises as that weight increases. Training moves each weight the opposite way, scaled by the learning rate η.

Researcher's shelf

Further reading

  • 3Blue1Brown, Essence of Linear Algebra (vectors, matrices, dot products, visually): https://www.3blue1brown.com/topics/linear-algebra
  • 3Blue1Brown, Essence of Calculus (derivatives and the chain rule): https://www.3blue1brown.com/topics/calculus
  • Khan Academy, Linear algebra: https://www.khanacademy.org/math/linear-algebra
  • Deisenroth, Faisal and Ong, Mathematics for Machine Learning (free book): https://mml-book.github.io/
  • NumPy, absolute basics for beginners: https://numpy.org/doc/stable/user/absolute_beginners.html

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.