rumblr Work in progressWIP

● The AI Primer · Lesson 0 · Before you begin

Math notation, from zero

You'll be able to explain Every symbol in an ML formula, as a short loop

Free lesson 26 min5 figures and diagrams12 interactive
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Σ is a loop that adds; Π is a loop that multiplies.
  2. A dot product multiplies matching entries and adds them up; it measures how much two vectors agree.
  3. A matrix multiply is a grid of dot products, and it is most of the work a neural network does.
  4. log undoes e, and it turns products into sums, which is why models work with log-probabilities.
  5. The gradient points uphill; training steps the other way.

Level 1

The practitioner's guide

In one sentence

The notation of AI is shorthand for short loops (Σ adds a list up, Π multiplies it, a dot product multiplies matching entries and adds, ∇ lists the slopes), and the same few dozen symbols fill the equations of every paper, the tables of every model card and the parameter lists of every API.

When you need it

You need it the moment a decision hinges on something written in symbols: a model card that says "70B parameters, 4k context, 2.0T tokens" (Llama 2's Table 1), an API reference with temperature, top_p and max_tokens, a training library's betas=(0.9, 0.999), or a paper whose whole claim is one equation. You don't need to derive anything; you need to read. The tell: you skip the equation, read the sentence after it, and the sentence says "see Equation 1". You don't need this lesson to use a chat product, and you don't need proofs, derivations or the appendix of any paper to use its result. One number from this lesson shows what reading buys: ten agent steps that each succeed 95% of the time all succeed with probability 0.95¹⁰ ≈ 0.60 (the demo's first table). Anyone who can read Π sees why long chains of steps fail before building one.

Your options

Five ways to handle a formula when you meet one, from the cheapest to the most certain:

Option What it does What it guarantees What it costs Where it lives
Read the prose, skip the formula Trusts the author's sentence about what the equation says Nothing; the prose usually points back at the equation for the part that matters Free Your reading
Decode the symbols Looks each letter up (Σ, η, θ, ‖x‖, ∂) and reads the whole line aloud as a sentence You know what is added, multiplied or divided, and over what Minutes, with a symbols table or the Greek-letter table in Level 2 Your reading
Check the shapes Follows the sizes through the line: a 4 × 3 matrix times a 3 × 5 matrix is 4 × 5, n tokens by d dimensions stays n by d Catches most misreadings, because a formula whose shapes don't line up cannot run One line of arithmetic per formula The paper's margin, or the shape comments in code
Evaluate it on three numbers Puts a tiny example through the formula by hand A number you can compare with the paper's own table Ten minutes Paper and pencil
Write it as a loop and run it Translates Σ into for, a dot product into multiply-then-add, and checks the result against NumPy The definition itself, executable; every function in this lesson is built and tested that way An hour the first time, minutes after A notebook

How to choose

Match the effort to what the notation decides.

  • A model card or a config file (parameter count, layers, context length): decode the names; no formula is involved. The size words are this lesson's shapes: in Hugging Face's LlamaConfig, hidden_size 4096 is the vector width d, num_hidden_layers 32 is the number of blocks, vocab_size 32000 is V, and rms_norm_eps 1e-6 is the ε that stops a division by zero.
  • An API parameter (temperature, top_p, top_k, max_tokens): read the one formula behind it once (softmax, with the scores divided by the temperature; primer.ml.big_picture walks it), then follow the vendor's advice. Claude's Messages API documents temperature from 0.0 to 1.0, default 1.0, closer to 0.0 for analytical and multiple-choice work and closer to 1.0 for creative work.
  • A training recipe (η, β₁, β₂, ε, λ, warmup steps, a clipping norm): decode the Greek and copy the values, because these are settings, not derivations. Attention Is All You Need trains with Adam at β₁ = 0.9, β₂ = 0.98, ε = 10⁻⁹ and 4000 warmup steps; Llama 2 with AdamW at β₁ = 0.9, β₂ = 0.95, ε = 10⁻⁵, weight decay 0.1 and gradient clipping 1.0. primer.ml.optimizers explains each knob.
  • A paper's central equation (a new loss, a new attention variant): evaluate it on three numbers, and write the loop if you will implement it.
  • What you can safely skip: derivations and convergence proofs (appendices), and the notation of the theory (expectations 𝔼, distributions 𝒩) until you reproduce a result. What you cannot skip: shapes, Σ, softmax, log, ∇ and the Greek letters that name hyperparameters.
  • Whatever you pick, read every formula aloud as a sentence before deciding it is beyond you. Every formula in this primer has a symbols table and an "In words" line for exactly that.

What it costs

Learning the vocabulary costs an afternoon: the whole of this lesson is a couple of dozen symbols, and python -m primer.notation runs every one of them in under a second. Misreading costs more. A log-probability is a natural logarithm, so an API that reports a token's logprob as −4.61 is saying 1%, and −0.11 is saying 90% (the lesson's e and log section); read it as base 10 and every confidence you compute is wrong. Attention's cost is O(n²) in sequence length (Table 1 of the transformer paper gives O(n²·d) per layer), so doubling the context quadruples that part of the work, which is the arithmetic behind long-context pricing. And the scaling laws are written in this notation: Kaplan et al. (2020) found that loss falls as a power law in model size, dataset size and compute, and Hoffmann et al. (2022, Chinchilla) that for every doubling of model size the training tokens should double too, which is how a 70-billion-parameter model came to beat a 280-billion one. A reader who cannot follow N, D and a power law cannot check a vendor's claim about either.

What breaks

  • Counting from 1 or from 0. Mathematics writes for the first entry; Python writes x[0]. A position formula copied from a paper into code is off by one until you check which convention it uses.
  • log means ln. In ML papers and API responses, log is the natural logarithm. −ln(0.01) = 4.61 and −ln(0.9) = 0.11; that gap is the "confidently wrong" penalty in every training loss.
  • One letter, several meanings. β is the momentum coefficient in Adam (β₁, β₂), the learned shift in a normalization layer and the strength knob in DPO; σ is a standard deviation or the sigmoid. The symbols table wins over memory every time.
  • Shapes that don't line up. A is 4 × 3 and B is 3 × 5: AB is 4 × 5 and BA does not exist. When a formula's shapes fail, you have misread a transpose, and the code will fail the same way.
  • Temperature 0 read as determinism. argmax picks the largest score, but Claude's API reference says results are not fully deterministic even at temperature 0.0, and the pipeline lesson explains why serving hardware makes that so.
  • Products of probabilities. The probability of a sentence is a product of thousands of numbers below 1, which underflows to 0 in floating point. That is why models add log-probabilities instead, and why a "score" in a log is negative.

In the wild

The transformer paper's Equation 1, softmax(QKᵀ/√d_k)V, packs a matrix multiply, a transpose, a square root and a softmax into one line, and its Table 1 is the big-O comparison of layer types. Model cards and configs carry the shapes: LlamaConfig (hidden_size 4096, intermediate_size 11008, 32 layers, 32 heads, vocab 32000, initializer_range 0.02), Llama 2's Table 1 (7B to 70B parameters, 2.0T tokens, learning rates 3.0 × 10⁻⁴ and 1.5 × 10⁻⁴), the Llama 3 abstract (a dense transformer with 405B parameters and a 128K-token context). APIs carry the sampling symbols: Claude's Messages API (temperature, max_tokens, stop_sequences; models released after Claude Opus 4.6 accept only the default temperature of 1.0 and no top_k), Hugging Face's GenerationConfig (do_sample, otherwise greedy; temperature 1.0, top_k 50, top_p 1.0, max_new_tokens, repetition_penalty). Training libraries carry the Greek: PyTorch's AdamW(lr=0.001, betas=(0.9, 0.999), eps=1e-08, weight_decay=0.01), Hugging Face's TrainingArguments (learning_rate 5e-5, adam_beta1 0.9, adam_beta2 0.999, adam_epsilon 1e-8, max_grad_norm 1.0). Every value above is quoted from the paper or the reference page named beside it; the papers are linked from the lessons that build on them.

Go deeper

Level 2 builds each symbol as the loop it stands for: Σ and Π, the dot product, ‖x‖, matrix multiply and transpose, e and log, softmax and argmax, mean and spread, derivatives and the gradient, then probability notation, big-O and the Greek alphabet, every one with numbers you can check by hand and a figure to read. If you only needed to read a model card or an API reference, you are done.

Level 2

How it works, from scratch

Machine learning papers look impenetrable mostly because of notation: Greek letters, big sigmas, little superscript Ts. Almost every symbol is shorthand for a short loop you could write in a few lines of Python. This lesson writes each one out as that loop, so that when a formula appears elsewhere in this primer you can read it aloud.

Every function below is deliberately written the slow, obvious way, with plain Python loops, so the code is the definition. The tests check each one against NumPy, which does the same thing fast.

Chapter 1

Lists and tables of numbers: vectors and matrices

Everyday picture A vector is a list of numbers, like a shopping receipt: (apples 3, bread 1, milk 2). A matrix is a table of numbers, like a spreadsheet with rows and columns.

Tiny example is a vector with 3 entries. Its entries are written with a small subscript: , , . A matrix with 2 rows and 3 columns has shape , and means "row 2, column 3".

You see Say Python
(lowercase, sometimes bold x) "the vector x" x = [3, 1, 2]
"x sub i": the i-th entry x[i - 1] (maths counts from 1, Python from 0)
(uppercase) "the matrix A" A = [[1, 2, 3], [4, 5, 6]]
or row i, column j A[i - 1][j - 1]
"the set of all lists of d real numbers" any list of d floats
"x is a list of 768 numbers" len(x) == 768
shape: n rows, d columns np.zeros((n, d))

In practice, a word's embedding is a vector, a batch of embeddings is a matrix (one row per word), and a model's weights are mostly matrices.

Chapter 2

Σ (capital sigma): add them all up

Run the loop yourself first; the symbol will read as what it is.

Everyday picture Totting up a receipt.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
"sum": add up everything that follows
(below) start a counter called at 1
(above) stop after the counter reaches
the thing being added on each step
"and so on, following the same pattern"

In words: "for i from 1 to n, add up x sub i."

With the numbers: . In code, Σ is a for loop with a running total (summation below).

Level 3: in Python
x = [1, 2, 3, 4]
total = 0
# Σ: visit each x_i, from i = 1 to n ...
for x_i in x:
    # ... and add it to a running total
    total += x_i
total  # → 10
# Python's built-in sum is the same loop
sum(x)  # → 10

Π (capital pi) is the same idea with multiplication: . It shows up whenever independent chances combine. For example, ten steps that each succeed 95% of the time all succeed with probability . This single fact explains why long chains of AI agent steps fail so often (primer.agents.planning).

In Python:

x = [1, 2, 3, 4]
product = 1
# Π: the same loop, multiplying instead of adding
for x_i in x:
    product *= x_i
product  # → 24
# ten steps that each succeed 95% of the time
round(0.95 ** 10, 2)  # → 0.6

In code: product is Π written the same way: a loop with a running product that starts at 1.

Chapter 3

The dot product: how much two lists agree

Everyday picture A recipe needs 2 eggs, 3 cups of flour and 1 cup of sugar, and eggs cost $1, flour $0.50 and sugar $2 per unit. The total cost is 2·1 + 3·0.5 + 1·2 = $5.50: multiply matching items, then add. That is a dot product.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
two vectors with the same number of entries,
the -th entries multiplied together
"dot": the whole multiply-then-add operation

In words: "multiply the lists position by position and add up the products."

With the numbers: .

Level 3: in Python
a = [1, 2]
b = [3, 0.5]
# Σ over i of a_i times b_i
sum(a_i * b_i for a_i, b_i in zip(a, b))  # → 4.0

Geometrically, the dot product is large when two arrows point the same way, zero when they are at right angles, and negative when they point apart. That is why it is the standard similarity score for attention and for embeddings.

Figure 1 · Diagram

Reading it: each pair of matching positions meets in a multiply box, first entries with first entries and second with second. Every product then flows into one addition. A dot product is always this shape: many multiplications feeding one sum, however long the lists get.

In code: dot pairs up matching entries, multiplies them and adds the products with summation, refusing lists of different lengths.

Chapter 4

‖x‖: the length of a vector

Everyday picture Walk 3 blocks east and 4 blocks north. As the crow flies, you are 5 blocks from where you started (Pythagoras).

Level 3: the formula and its symbols

Symbols

Symbol Meaning
the norm (length) of , also written
the -th entry squared ()
square root

In words: "square every entry, add them up, and take the square root."

With the numbers: and .

Level 3: in Python
import math
def length(x):
    # √ of Σ x_i²
    return math.sqrt(sum(x_i ** 2 for x_i in x))
length([3, 4])  # → 5.0
length([1, 2, 2])  # → 3.0

Dividing a vector by its length gives a unit vector of length 1 that points the same way. Embedding systems do this constantly, because then the dot product measures only direction (see primer.ml.embeddings.similarity).

In code: norm is the square root of a vector's dot product with itself.

Chapter 5

Matrix multiply and transpose

Transpose, written ("A transpose"), flips a table on its diagonal so rows become columns: a matrix becomes .

Matrix multiply, written with nothing in between, is a whole grid of dot products. The cell in row , column of the answer is row of dotted with column of .

Level 3: the formula and its symbols

Symbols

Symbol Meaning
an matrix
an matrix; its row count must equal 's column count,
the answer's cell at row , column ; the answer is
the counter that walks along row of and down column of together

In words: "to fill cell (i, j), walk across row i of A and down column j of B at the same pace, multiplying and adding."

With the numbers: $\begin{pmatrix}1&2\3&4\end{pmatrix}\begin{pmatrix}5&6\7&8\end{pmatrix} = \begin{pmatrix}1\cdot5+2\cdot7 & 1\cdot6+2\cdot8\ 3\cdot5+4\cdot7 & 3\cdot6+4\cdot8\end{pmatrix} = \begin{pmatrix}19&22\43&50\end{pmatrix}$.

Level 3: in Python
A = [[1, 2], [3, 4]]
B = [[5, 6], [7, 8]]
# A has m columns, B has m rows: they must match
m = len(B)
# (AB)_ij = Σ_k A_ik B_kj
AB = [[sum(A[i][k] * B[k][j] for k in range(m))
       for j in range(len(B[0]))]
      for i in range(len(A))]
AB  # → [[19, 22], [43, 50]]
# the transpose: rows become columns
[list(column) for column in zip(*A)]  # → [[1, 3], [2, 4]]

Figure 2 · Diagram

Reading it: a matrix multiply is nothing but this box, repeated for every (row, column) pair. The shape rule falls out: the row of A and the column of B must be the same length, or there is nothing to pair up. Almost all the computation in a neural network is this one operation, which is why GPUs, built to do thousands of multiply-adds at once, are the hardware of AI.

In code: transpose turns columns into rows; matmul transposes B once, then fills each cell with one dot of a row of A and a column of B.

Chapter 6

e and log: growth, and its undo button

Everyday picture Money in an account that compounds continuously at 100% a year grows by a factor of e ≈ 2.718 in one year. is that growth run for years. The natural logarithm (often just in ML papers) answers the reverse question: how many years of that growth turn 1 into ?

You see Say Example
or "e to the x" , , ,
or "log of y" , ,

Two facts carry most of machine learning:

  1. is always positive and grows fast, which is why softmax uses it.
  2. . Logs turn multiplying into adding. The probability of a whole sentence is a product of thousands of small numbers, which would underflow to 0 on a computer. Its log is a sum of manageable negative numbers. This is why models work with "log-probabilities", and why the standard training loss is (see primer.ml.losses).

Figure 3 · Drawn from the lesson's code

−2 0 2 4 6 8 −4 −2 0 2 4 6 8 e and its undo button, ln ln 0.9 = -0.11 ln 0.01 = -4.61 e x   ( a l w a y s   >   0 ,   g r o w s   f a s t ) l n x e   ( u n d o e s   ) x mirror line y = x

e to the x stays above zero and climbs steeply; ln x mirrors it across y = x and plunges near 0, so ln 0.01 = -4.61 while ln 0.9 = -0.11

Reading it: the blue curve is always above zero and climbs steeply; every step of 1 to the right multiplies its height by 2.718. The red curve is the same curve mirrored across the dashed diagonal, because it undoes . It only exists for positive inputs, and it dives towards −∞ as its input approaches 0. That dive is the "confidently wrong" penalty in cross-entropy: , far more than .

The second fact above, happening in a computer's own numbers.

Chapter 7

softmax and argmax

softmax turns a list of scores into shares that are positive and sum to

  1. It is covered step by step, with its symbols decoded, in primer.ml.attention:
Level 3: the formula and its symbols

With the numbers: softmax(2.0, 1.0, 0.5) = (0.63, 0.23, 0.14).

Level 3: in Python
import math
z = [2.0, 1.0, 0.5]
# e^(z_i) for each score
exps = [math.exp(z_i) for z_i in z]
# Σ_j e^(z_j)
total = sum(exps)
# each share of the total
[round(e / total, 2) for e in exps]  # → [0.63, 0.23, 0.14]
# argmax: the position of the largest
max(range(len(z)), key=lambda i: z[i])  # → 0

argmax is simpler: it answers "which position holds the largest value?", not "what is the largest value?". argmax(0.1, 7.0, 3.0) = position 2 (index 1 in Python). "Greedy decoding" in a language model is picking the argmax token every step.

In code: softmax subtracts the largest score before exponentiating, so the exponentials cannot overflow; argmax walks the list and remembers the position of the biggest value.

Chapter 8

Mean, variance, standard deviation: the middle and the spread

Everyday picture Two classes both average 70% on a test. In one, everyone scored 68 to 72; in the other, scores ran from 30 to 100. The mean is the same; the spread is not.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
(mu) the mean: the average
(sigma squared) the variance: the average squared distance from the mean; also written
(sigma) the standard deviation: the typical distance from the mean, in the data's own units
"add them up and divide by how many": an average

In words: "the mean is the average; the variance is the average squared distance from the mean; the standard deviation is its square root."

With the numbers: for (2, 4, 4, 4, 5, 5, 7, 9): μ = 40 / 8 = 5; the squared distances are (9, 1, 1, 1, 0, 0, 4, 16), which sum to 32, so σ² = 32 / 8 = 4 and σ = 2.

Level 3: in Python
import math
x = [2, 4, 4, 4, 5, 5, 7, 9]
n = len(x)
# μ = (1/n) Σ x_i
mu = sum(x) / n
mu  # → 5.0
# the squared distances from μ
[(x_i - mu) ** 2 for x_i in x]  # → [9.0, 1.0, 1.0, 1.0, 0.0, 0.0, 4.0, 16.0]
# σ² = their average
variance = sum((x_i - mu) ** 2 for x_i in x) / n
# σ is its square root
variance, math.sqrt(variance)  # → (4.0, 2.0)

Spread matters constantly in neural networks. If numbers flowing through a network spread out layer after layer, training blows up; normalization layers and careful initialization exist to hold σ near 1 (see primer.ml.deep_nets and the √d_k in primer.ml.attention).

Figure 4 · Drawn from the lesson's code

−10.0 −7.5 −5.0 −2.5 0.0 2.5 5.0 7.5 10.0 value 0 100 200 300 400 500 how many samples Same mean (μ = 0), different spread σ = 1 σ = 3

Two histograms share the mean 0, but the sigma 1 samples pile up tall and narrow while the sigma 3 samples spread about three times as wide

Reading it: both histograms are centred on the same mean (the dashed line), but the blue one is tall and narrow (σ = 1) while the red one is low and wide (σ = 3). The standard deviation is roughly how far from the dashed line a typical sample lands. About two thirds of samples fall within one σ of the mean.

In code: mean, variance and std are the three formulas above, each a short loop built on summation.

Chapter 9

Derivatives and gradients: which way is downhill?

Feel the slope before the formula for it.

Everyday picture You're on a hillside in thick fog and want to reach the valley. You can't see it, but you can feel the slope under your feet. Take a small step in the steepest downhill direction, feel again, and repeat. That is how every neural network is trained (gradient descent).

The derivative of a function at a point is its slope there: how much the output changes per tiny nudge of the input.

Level 3: the formula and its symbols

Symbols

Symbol Meaning
a function: put in, get a number out
or the derivative: the slope of at ("dee f dee x")
a tiny nudge, like 0.00001
"approximately equal": exact as shrinks to 0

In words: "nudge the input a hair up and a hair down, and see how much the output changes per unit of nudge."

With the numbers: for at : (3.00001² − 2.99999²) / 0.00002 = 6. The slope of at 3 is 6 (the rule is ).

Level 3: in Python
def f(x):
    return x ** 2
h = 0.00001
# rise over run across a tiny step
round((f(3 + h) - f(3 - h)) / (2 * h), 6)  # → 6.0

With many inputs, nudge each one separately. The slope in each direction is a partial derivative, written (the curly ∂ just means "only this input moves, the others stay fixed"). Collect them into a vector and you have the gradient, written ("nabla f" or "grad f"). It points in the steepest uphill direction, so training steps the opposite way:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
(theta) all the model's adjustable numbers (its parameters or weights)
the loss: one number measuring how wrong the model is
the gradient of the loss: for each weight, how much the loss rises if that weight is nudged up
(eta) the learning rate: how big a step to take, e.g. 0.001

In words: "move every weight a small step in the direction that lowers the loss the fastest."

With the numbers: for the bowl at (1, 2), the gradient is (2, 4), pointing uphill away from the bottom at (0, 0). With η = 0.1, the step goes to (1 − 0.2, 2 − 0.4) = (0.8, 1.6), closer to the bottom.

Level 3: in Python
# (x, y)
theta = [1.0, 2.0]
# ∇L for L = x² + y²: each slope is 2 times the value
gradient = [2 * theta[0], 2 * theta[1]]
gradient  # → [2.0, 4.0]
eta = 0.1
# θ_new = θ − η ∇L
[t - eta * g for t, g in zip(theta, gradient)]  # → [0.8, 1.6]

Figure 5 · Drawn from the lesson's code

−2 −1 0 1 2 x −2 −1 0 1 2 y L o s s   :   a r r o w s   p o i n t   d o w n h i l l x y 2 2 + gradient descent from (1, 2)

Arrows on circular contours point straight at the centre, longest on the steep rim; descent from (1, 2) takes big steps first, then ever smaller ones

Reading it: the rings are contour lines, as on a hiking map: every point on a ring has the same loss, and the bottom of the bowl is the centre. The arrows show the negative gradient at several spots. They always cross the rings at right angles, pointing straight downhill, and they are longer where the slope is steeper. The dotted path is gradient descent from (1, 2): big steps on the steep outer slope, shrinking steps as the ground flattens near the bottom.

The chain rule says how slopes combine when functions are chained: if , then . Multiply the slopes along the chain. Backpropagation is the chain rule applied backwards through every layer of a network, reusing work as it goes (see primer.ml.neural_net).

In code: derivative measures a slope by nudging the input a hair up and a hair down; gradient does that for each input in turn and collects the slopes into one list.

Chapter 10

Probability notation

You see Say Meaning
or "probability of A" a number from 0 (never) to 1 (certain)
"probability of A given B" the chance of A once you know B happened
"probability of word t given the words before it" what a language model computes, one word at a time
"expected value of x" the average you'd get over many tries
"x is drawn from a normal distribution with mean 0 and variance 1" rng.standard_normal()

Chapter 11

Big-O: how cost grows

, "order n squared", describes how cost grows as the input grows, ignoring constant factors. If doubles, an cost doubles and an cost quadruples. Attention is in sequence length, which is why long context is expensive (primer.ml.attention).

Chapter 12

Greek letters you'll meet

Letter Name Usually means
α alpha a mixing weight, or a learning rate
β beta momentum coefficients in optimizers (β₁, β₂); a strength knob in DPO
γ, β gamma, beta the learned scale and shift in normalization layers
δ delta a small change, or an error signal in backprop
ε epsilon a tiny number added to avoid dividing by zero
η eta the learning rate
θ theta all the model's parameters
λ lambda the strength of a penalty (weight decay)
μ mu a mean
σ sigma a standard deviation, or the sigmoid function
τ tau a temperature (sharpness of softmax)
∇ nabla the gradient
∂ "partial" a partial derivative

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Read aloud and evaluate it.Think it through, then reveal

"The sum, for i from 1 to 3, of i squared": 1 + 4 + 9 = 14.

Question 2What is (2, −1, 3) · (1, 4, 0), and what does its sign tell you?Think it through, then reveal

2 − 4 + 0 = −2. It's negative, so the vectors point somewhat apart.

Question 3A is 4 × 3 and B is 3 × 5. What shape is AB? What about BA?Think it through, then reveal

AB is 4 × 5. BA is not defined: B's 5 columns don't match A's 4 rows.

Question 4Why do language models add log-probabilities instead of multiplying probabilities?Think it through, then reveal

A product of thousands of numbers below 1 underflows to 0 in floating point. log(a·b) = log a + log b, so the product becomes a sum of moderate negative numbers.

Question 5What does ∇L tell you, and which way does training move?Think it through, then reveal

For every weight, how fast the loss rises as that weight increases. Training moves each weight the opposite way, scaled by the learning rate η.

Researcher's shelf

Further reading

  • 3Blue1Brown, Essence of Linear Algebra (vectors, matrices, dot products, visually): https://www.3blue1brown.com/topics/linear-algebra
  • 3Blue1Brown, Essence of Calculus (derivatives and the chain rule): https://www.3blue1brown.com/topics/calculus
  • Khan Academy, Linear algebra: https://www.khanacademy.org/math/linear-algebra
  • Deisenroth, Faisal and Ong, Mathematics for Machine Learning (free book): https://mml-book.github.io/
  • NumPy, absolute basics for beginners: https://numpy.org/doc/stable/user/absolute_beginners.html

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.