1.Vectors: Anything as a List of Numbers
AI handles text, photos, and even tastes by turning them into lists of numbers. Such a list of numbers is called a vector. If you express movie tastes as preferences for (action, romance, comedy), you can write Jisoo as [5, 1, 3], Minho as [4, 2, 3], and Seoyeon as [1, 5, 2]. A language model turns each word or word piece (token) into a vector of hundreds to thousands of numbers, but the principle is the same as in this three-number example.
Given two vectors, you can calculate "how similar they are." The distance between two points from Lesson 6 is one way; the ways used more often in AI are the dot product and cosine similarity.
2.Dot Products and Cosine Similarity
The dot product multiplies the numbers in the same positions of two vectors and adds them all up. The more two people like the same items, the more large numbers get multiplied together, and the bigger the dot product. However, the dot product automatically grows for vectors whose numbers are large overall, so cosine similarity divides by the vectors' lengths (magnitudes) to compare direction alone.
Cosine similarity takes values from −1 to 1. Close to 1 means the same direction (similar tastes), 0 means unrelated directions (at right angles), and −1 means exact opposites. In search and recommendations, "finding documents with similar meaning" usually means calculating this value and picking the largest ones.
- Step 1: Dot products: Jisoo·Minho = 5×4 + 1×2 + 3×3 = 20 + 2 + 9 = 31. Jisoo·Seoyeon = 5×1 + 1×5 + 3×2 = 5 + 5 + 6 = 16.
- Step 2: Lengths: |Jisoo| = √(25 + 1 + 9) = √35 ≈ 5.92, |Minho| = √(16 + 4 + 9) = √29 ≈ 5.39, |Seoyeon| = √(1 + 25 + 4) = √30 ≈ 5.48.
- Step 3: Jisoo·Minho similarity = 31 ÷ √(35 × 29) = 31 ÷ √1015 ≈ 31 ÷ 31.86 ≈ 0.97.
- Step 4: Jisoo·Seoyeon similarity = 16 ÷ √(35 × 30) = 16 ÷ √1050 ≈ 16 ÷ 32.40 ≈ 0.49.
- Check: Both similarities lie between −1 and 1. Jisoo and Minho both like action the most, while Seoyeon likes romance the most, so the result matches intuition too.
3.Matrix Multiplication: Many Weighted Sums at Once
What a single neuron in a neural network does is simple. It multiplies each input by a weight and adds them up (a weighted sum), adds a bias, and then transforms the result once by a fixed rule. One of the most common rules is to turn negative values into 0 and leave positive values as they are. With inputs [2, 3], weights [0.5, −1], and a bias of 1, you get 0.5×2 + (−1)×3 + 1 = −1, which is negative, so the neuron outputs 0.
With several neurons, you have to compute several weighted sums. A matrix is the way to write them all at once. Each row of the matrix holds one neuron's weights, and multiplying a matrix by a vector means "taking the dot product of each row with the input vector." Most of the computation in a giant AI model is this kind of matrix multiplication, which is why specialized chips that do matrix multiplication fast have become so important.
- Step 1: First neuron = dot product of the first row and the input: 1×4 + 2×1 = 4 + 2 = 6.
- Step 2: Second neuron = dot product of the second row and the input: 3×4 + (−1)×1 = 12 − 1 = 11.
- Step 3: The resulting vector is [6, 11].
- Check: Recalculate column by column. Input 4 × first column [1, 3] = [4, 12], input 1 × second column [2, −1] = [2, −1], and adding them gives [6, 11], the same.
4.Loss Functions: How Wrong, in a Number
For a model to learn, it first has to measure "how wrong it is" with a single number. The function that produces this number is called a loss function. A common one for models that predict numbers is mean squared error (MSE): the average of the squared differences (errors) between predictions and correct answers. Its form is almost the same as the variance from Lesson 8.
There are two reasons for squaring. One is to count errors as "wrong" whether they are positive or negative; the other is to penalize big mistakes more heavily. An error of 2 becomes 4, and an error of 4 becomes 16.
- Step 1: Errors: 3 − 2 = 1, 5 − 5 = 0, 8 − 10 = −2.
- Step 2: Squares: 1, 0, 4. The sum is 5.
- Step 3: Average: 5 ÷ 3 ≈ 1.67.
- Check: The third value, with the biggest error (−2), accounts for 4 of the total of 5. This fits the property of squaring: big mistakes dominate the loss.
5.Slopes and Gradient Descent
Which way should the weights be adjusted to reduce the loss? Picture someone climbing down a foggy mountain to the valley. They can't see far, but they can feel the slope under their feet. They take a step downhill, feel the slope again, and take another step. "The slope at a single point" from Lesson 4 is exactly this slope underfoot, and descending this way is called gradient descent.
As the simplest example, take a loss of L(w) = (w − 3)². The loss is smallest, 0, when w = 3. The slope of this function is 2(w − 3). If w is less than 3, the slope is negative, so w should be increased; if it is greater, the slope is positive, so w should be decreased. So you update with "new w = w − learning rate × slope." The learning rate is the size of one step.
A real model may have billions of weights, but it does the same thing. For each weight, it calculates "how much the loss changes if this weight changes a little" (the slope, or gradient), and repeats small adjustments in the direction that reduces the loss countless times.
- Step 1: w = 0: slope 2(0 − 3) = −6, new w = 0 − 0.25 × (−6) = 1.5. Loss 9 → (1.5 − 3)² = 2.25.
- Step 2: w = 1.5: slope 2(1.5 − 3) = −3, new w = 1.5 − 0.25 × (−3) = 2.25. Loss (2.25 − 3)² = 0.5625.
- Step 3: w = 2.25: slope 2(2.25 − 3) = −1.5, new w = 2.25 − 0.25 × (−1.5) = 2.625. Loss (2.625 − 3)² = 0.140625.
- Check: The remaining distance to the target 3 is halved each time: 3 → 1.5 → 0.75 → 0.375. The loss becomes 1/4 each time, falling 9 → 2.25 → 0.5625 → 0.140625.
- Step 1: Slope −6, new w = 0 − 1.1 × (−6) = 6.6.
- Step 2: Distance from the target 3: it started at 3 and is now |6.6 − 3| = 3.6, so it has actually moved farther away. The loss has also grown from 9 to 12.96.
- Step 3: The next step: slope 2(6.6 − 3) = 7.2, new w = 6.6 − 7.92 = −1.32, a distance of 4.32. It overshoots to the other side and drifts farther and farther away.
- Check: 3.6² = 12.96, and the distance is multiplied by 1.2 each time (3 → 3.6 → 4.32), so it diverges.
📌 Key points
- AI handles words, photos, and tastes by turning them into lists of numbers (vectors)
- The dot product multiplies matching positions and adds them up; cosine similarity is similarity of direction, found by dividing by the lengths
- Matrix × vector = a dot product for each row = the weighted sums of many neurons computed at once
- A loss function measures how wrong a model is with a single number; mean squared error is the mean of the squared errors
- Gradient descent: new weight = weight − learning rate × slope, repeated adjustments in the direction that reduces the loss
- If the learning rate is too large, training diverges; if it is too small, training is slow
🤖 Try asking AI like this
Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.
When you want to check AI concepts yourself with small numbers
Explain cosine similarity using three vectors with 3 numbers each. Calculate the dot products, lengths, and similarities step by step, but before giving the final answers, show me only the intermediate results first so I can try calculating them myself.
When you want to follow gradient descent by hand
I want to practice gradient descent with the loss function L(w) = (w − 3)². Varying the starting value and learning rate, show me three cases (converges well, too slow, diverges), each for five steps in a table. Include columns for w, slope, new w, and loss in each table, and I'll calculate the first two rows myself to check.
When verifying a math explanation an AI gave you
For each formula in the explanation you just gave, plug in a tiny numerical example and show the calculation. If anywhere the explanation and the calculated result disagree, point it out.
- Standard high school math textbook content (vectors, matrices, basics of differentiation)
- Standard content of high school "Mathematics for Artificial Intelligence" courses
Reached every goal above? Mark the lesson complete.
Storage is unavailable in this browser, so this lasts only for this page.📐 Math Basics
- 1Numbers and Operations: Fractions and Decimals Revisited
- 2Ratios and Rates: Reading Percentages Correctly
- 3Equations: A Balance for Finding Unknown Numbers
- 4Functions and Graphs: An Eye for Change
- 5Exponents and Logarithms: A World That Grows by Multiplying
- 6Geometry Basics: Area and Pythagoras
- 7Probability: Putting Numbers on Uncertainty
- 8Statistics: Mean, Median, and Variance
- 9Reading Data: The Traps in Graphs
- 10The Math for Understanding AI