B Binance · The world's largest crypto exchangeBinance Sign up → AD OKX OKX · A leading global crypto exchangeOKX Sign up → AD
📐 Math Basics · Lesson 10 / 10

The Math for Understanding AI

AI's calculations come down to repeating two things: "multiply lists of numbers together and add them up" and "adjust a little at a time in the direction that reduces the error." Work through small examples by hand, and you'll see that giant models run on the same pattern.

⏱ About 20 min ✍️ 4 practice questions Updated 2026-10-08
🎯 By the end of this lesson you can
  • Calculate the dot product and length of vectors, and explain what cosine similarity means
  • Calculate the product of a matrix and a vector as weighted sums
  • Calculate how wrong predictions are using mean squared error
  • Calculate one step of gradient descent and explain what goes wrong when the learning rate is too large

1.Vectors: Anything as a List of Numbers

AI handles text, photos, and even tastes by turning them into lists of numbers. Such a list of numbers is called a vector. If you express movie tastes as preferences for (action, romance, comedy), you can write Jisoo as [5, 1, 3], Minho as [4, 2, 3], and Seoyeon as [1, 5, 2]. A language model turns each word or word piece (token) into a vector of hundreds to thousands of numbers, but the principle is the same as in this three-number example.

Given two vectors, you can calculate "how similar they are." The distance between two points from Lesson 6 is one way; the ways used more often in AI are the dot product and cosine similarity.

2.Dot Products and Cosine Similarity

The dot product multiplies the numbers in the same positions of two vectors and adds them all up. The more two people like the same items, the more large numbers get multiplied together, and the bigger the dot product. However, the dot product automatically grows for vectors whose numbers are large overall, so cosine similarity divides by the vectors' lengths (magnitudes) to compare direction alone.

Cosine similarity takes values from −1 to 1. Close to 1 means the same direction (similar tastes), 0 means unrelated directions (at right angles), and −1 means exact opposites. In search and recommendations, "finding documents with similar meaning" usually means calculating this value and picking the largest ones.

Dot product a · b = a₁b₁ + a₂b₂ + a₃b₃
Length |a| = √(a₁² + a₂² + a₃²)
Cosine similarity = (a · b) ÷ (|a| × |b|)
ExampleFind the cosine similarity between Jisoo [5, 1, 3] and Minho [4, 2, 3], and between Jisoo and Seoyeon [1, 5, 2], and decide whose taste is closer to Jisoo's.
  1. Step 1: Dot products: Jisoo·Minho = 5×4 + 1×2 + 3×3 = 20 + 2 + 9 = 31. Jisoo·Seoyeon = 5×1 + 1×5 + 3×2 = 5 + 5 + 6 = 16.
  2. Step 2: Lengths: |Jisoo| = √(25 + 1 + 9) = √35 ≈ 5.92, |Minho| = √(16 + 4 + 9) = √29 ≈ 5.39, |Seoyeon| = √(1 + 25 + 4) = √30 ≈ 5.48.
  3. Step 3: Jisoo·Minho similarity = 31 ÷ √(35 × 29) = 31 ÷ √1015 ≈ 31 ÷ 31.86 ≈ 0.97.
  4. Step 4: Jisoo·Seoyeon similarity = 16 ÷ √(35 × 30) = 16 ÷ √1050 ≈ 16 ÷ 32.40 ≈ 0.49.
  5. Check: Both similarities lie between −1 and 1. Jisoo and Minho both like action the most, while Seoyeon likes romance the most, so the result matches intuition too.
AnswerJisoo·Minho ≈ 0.97, Jisoo·Seoyeon ≈ 0.49. Minho's taste is closer

3.Matrix Multiplication: Many Weighted Sums at Once

What a single neuron in a neural network does is simple. It multiplies each input by a weight and adds them up (a weighted sum), adds a bias, and then transforms the result once by a fixed rule. One of the most common rules is to turn negative values into 0 and leave positive values as they are. With inputs [2, 3], weights [0.5, −1], and a bias of 1, you get 0.5×2 + (−1)×3 + 1 = −1, which is negative, so the neuron outputs 0.

With several neurons, you have to compute several weighted sums. A matrix is the way to write them all at once. Each row of the matrix holds one neuron's weights, and multiplying a matrix by a vector means "taking the dot product of each row with the input vector." Most of the computation in a giant AI model is this kind of matrix multiplication, which is why specialized chips that do matrix multiplication fast have become so important.

[[1, 2], [3, −1]] × [4, 1]
= [1×4 + 2×1, 3×4 + (−1)×1]
= [6, 11]
ExampleThe first row of a weight matrix is [1, 2], the second row is [3, −1], and the input is [4, 1]. Find the weighted sums of the two neurons (no bias).
  1. Step 1: First neuron = dot product of the first row and the input: 1×4 + 2×1 = 4 + 2 = 6.
  2. Step 2: Second neuron = dot product of the second row and the input: 3×4 + (−1)×1 = 12 − 1 = 11.
  3. Step 3: The resulting vector is [6, 11].
  4. Check: Recalculate column by column. Input 4 × first column [1, 3] = [4, 12], input 1 × second column [2, −1] = [2, −1], and adding them gives [6, 11], the same.
Answer[6, 11]

4.Loss Functions: How Wrong, in a Number

For a model to learn, it first has to measure "how wrong it is" with a single number. The function that produces this number is called a loss function. A common one for models that predict numbers is mean squared error (MSE): the average of the squared differences (errors) between predictions and correct answers. Its form is almost the same as the variance from Lesson 8.

There are two reasons for squaring. One is to count errors as "wrong" whether they are positive or negative; the other is to penalize big mistakes more heavily. An error of 2 becomes 4, and an error of 4 becomes 16.

MSE = (error₁² + error₂² + … + errorₙ²) ÷ n
Error = prediction − correct answer
ExampleFind the mean squared error when the predictions are [3, 5, 8] and the correct answers are [2, 5, 10].
  1. Step 1: Errors: 3 − 2 = 1, 5 − 5 = 0, 8 − 10 = −2.
  2. Step 2: Squares: 1, 0, 4. The sum is 5.
  3. Step 3: Average: 5 ÷ 3 ≈ 1.67.
  4. Check: The third value, with the biggest error (−2), accounts for 4 of the total of 5. This fits the property of squaring: big mistakes dominate the loss.
Answer5/3 ≈ 1.67

5.Slopes and Gradient Descent

Which way should the weights be adjusted to reduce the loss? Picture someone climbing down a foggy mountain to the valley. They can't see far, but they can feel the slope under their feet. They take a step downhill, feel the slope again, and take another step. "The slope at a single point" from Lesson 4 is exactly this slope underfoot, and descending this way is called gradient descent.

As the simplest example, take a loss of L(w) = (w − 3)². The loss is smallest, 0, when w = 3. The slope of this function is 2(w − 3). If w is less than 3, the slope is negative, so w should be increased; if it is greater, the slope is positive, so w should be decreased. So you update with "new w = w − learning rate × slope." The learning rate is the size of one step.

A real model may have billions of weights, but it does the same thing. For each weight, it calculates "how much the loss changes if this weight changes a little" (the slope, or gradient), and repeats small adjustments in the direction that reduces the loss countless times.

L(w) = (w − 3)², slope = 2(w − 3)
New w = w − learning rate × slope
ExampleFor L(w) = (w − 3)², start at w = 0 with a learning rate of 0.25 and perform gradient descent three times.
  1. Step 1: w = 0: slope 2(0 − 3) = −6, new w = 0 − 0.25 × (−6) = 1.5. Loss 9 → (1.5 − 3)² = 2.25.
  2. Step 2: w = 1.5: slope 2(1.5 − 3) = −3, new w = 1.5 − 0.25 × (−3) = 2.25. Loss (2.25 − 3)² = 0.5625.
  3. Step 3: w = 2.25: slope 2(2.25 − 3) = −1.5, new w = 2.25 − 0.25 × (−1.5) = 2.625. Loss (2.625 − 3)² = 0.140625.
  4. Check: The remaining distance to the target 3 is halved each time: 3 → 1.5 → 0.75 → 0.375. The loss becomes 1/4 each time, falling 9 → 2.25 → 0.5625 → 0.140625.
Answerw approaches 3 through 1.5, 2.25, and 2.625, and the loss falls to 0.140625
ExampleFor the same function, what happens if you take one step from w = 0 with a large learning rate of 1.1?
  1. Step 1: Slope −6, new w = 0 − 1.1 × (−6) = 6.6.
  2. Step 2: Distance from the target 3: it started at 3 and is now |6.6 − 3| = 3.6, so it has actually moved farther away. The loss has also grown from 9 to 12.96.
  3. Step 3: The next step: slope 2(6.6 − 3) = 7.2, new w = 6.6 − 7.92 = −1.32, a distance of 4.32. It overshoots to the other side and drifts farther and farther away.
  4. Check: 3.6² = 12.96, and the distance is multiplied by 1.2 each time (3 → 3.6 → 4.32), so it diverges.
AnswerIt overshoots the valley, bouncing to the other side and drifting farther away (divergence). If the learning rate is too large, learning fails
If the learning rate is too small, the descent takes a long time; if it is too large, you overshoot the valley and bounce upward. In real training, too, the learning rate is one of the first values people tune.

📌 Key points

  • AI handles words, photos, and tastes by turning them into lists of numbers (vectors)
  • The dot product multiplies matching positions and adds them up; cosine similarity is similarity of direction, found by dividing by the lengths
  • Matrix × vector = a dot product for each row = the weighted sums of many neurons computed at once
  • A loss function measures how wrong a model is with a single number; mean squared error is the mean of the squared errors
  • Gradient descent: new weight = weight − learning rate × slope, repeated adjustments in the direction that reduces the loss
  • If the learning rate is too large, training diverges; if it is too small, training is slow

✍️ Practice questions

Answer first, then open "Answer and explanation".

Q1. What is the dot product of the vectors [1, 2, 3] and [4, 5, 6]?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② 32

1×4 + 2×5 + 3×6 = 4 + 10 + 18 = 32. [4, 10, 18] is what you get if you multiply but don't add; the result of a dot product is a single number.

Q2. What does it mean when the cosine similarity of two vectors is 0?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ③ The two vectors are at right angles (unrelated directions)

A cosine similarity of 0 means the dot product is 0, that is, the two directions are at right angles. The same direction gives 1, and opposite directions give −1.

Q3. Find the mean squared error when the errors of two predictions are 2 and −2.

Answer and explanation
Answer 4

Squared, they are 4 and 4, and the average is (4 + 4) ÷ 2 = 4. You square them because simply averaging the errors gives 0, which would wrongly look like "not wrong at all."

Q4. For L(w) = (w − 3)², with w = 5 and a learning rate of 0.1, find w after one step of gradient descent.

Answer and explanation
Answer 4.6

Slope 2(5 − 3) = 4, new w = 5 − 0.1 × 4 = 4.6. It has moved closer to the target 3, and the loss fell from 4 to (4.6 − 3)² = 2.56.

🤖 Try asking AI like this

Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.

When you want to check AI concepts yourself with small numbers

Explain cosine similarity using three vectors with 3 numbers each. Calculate the dot products, lengths, and similarities step by step, but before giving the final answers, show me only the intermediate results first so I can try calculating them myself.

When you want to follow gradient descent by hand

I want to practice gradient descent with the loss function L(w) = (w − 3)². Varying the starting value and learning rate, show me three cases (converges well, too slow, diverges), each for five steps in a table. Include columns for w, slope, new w, and loss in each table, and I'll calculate the first two rows myself to check.

When verifying a math explanation an AI gave you

For each formula in the explanation you just gave, plug in a tiny numerical example and show the calculation. If anywhere the explanation and the calculated result disagree, point it out.
References
  • Standard high school math textbook content (vectors, matrices, basics of differentiation)
  • Standard content of high school "Mathematics for Artificial Intelligence" courses

Reached every goal above? Mark the lesson complete.

📐 Math Basics

  1. 1Numbers and Operations: Fractions and Decimals Revisited
  2. 2Ratios and Rates: Reading Percentages Correctly
  3. 3Equations: A Balance for Finding Unknown Numbers
  4. 4Functions and Graphs: An Eye for Change
  5. 5Exponents and Logarithms: A World That Grows by Multiplying
  6. 6Geometry Basics: Area and Pythagoras
  7. 7Probability: Putting Numbers on Uncertainty
  8. 8Statistics: Mean, Median, and Variance
  9. 9Reading Data: The Traps in Graphs
  10. 10The Math for Understanding AI
📚 Worth reading
🧠What Generative AI Does and Where It Fails→ ✍️How to Write a Good Prompt→ 🔍Checking AI Answers→ 📚Using AI for Study and Work Without Plagiarism→
← Foundations for the AI Era