1.The four steps of machine learning
In Lesson 1 we said machine learning is the approach in which a machine finds patterns in example data. If you break this process down a little further, almost every machine learning project goes through the same four steps: collect data, decide what shape of rule to look for (the model), adjust the model to fit the data (training), and check how well it does on data it has never seen (evaluation).
Take, for example, the problem of predicting an apartment's price from its size. A table of sizes and prices for apartments that have already sold is the data. A straight-line rule, "price = size times some number plus some other number," is the model. Trying different values for the number to multiply by and the number to add, and finding the values that make the gap from actual prices smallest, is training. Finally, measuring how accurate the predictions are on sales records that were not used in training is evaluation.
| Step | What it does | In house-price prediction |
|---|---|---|
| Data | Collect inputs paired with correct answers | Size (input) and actual sale price (correct answer) |
| Model | Decide the shape of the rule to look for | A straight line of the form price = a × size + b |
| Training | Adjust the model's numbers to make the error smaller | Change a and b to shrink the gap between prediction and reality |
| Evaluation | Measure performance on data it has never seen | Check how accurate predictions are on sales not used in training |
2.Data: examples with and without correct answers
Learning from data where every input has a correct answer (a label) attached is called supervised learning. It learns from examples that people have labeled, such as "this email is spam," "this photo is a cat" or "this house is worth 5 price units." Classification (is it spam or not?) and numeric prediction (how much is it? — also called regression) are typical supervised learning problems.
Giving the machine only data, with no correct answers, and having it group similar things or find structure is unsupervised learning. An online store grouping customers with similar tastes based only on their purchase histories is an example. Since nobody defined what the right answer is, a person ultimately has to judge whether the groups the machine found are meaningful.
Either way, the quality of the data determines the quality of the result. Examples with wrong labels, skewed examples and too few examples all make the model worse to the same degree. That is why people often say, "garbage in, garbage out."
- Supervised learning: learns from labeled examples — spam classification, house-price prediction, photo classification
- Unsupervised learning: finds structure without labels — grouping customers, clustering similar documents
- Reinforcement learning: learns to increase the rewards it receives for its actions — games, robot control
3.Training: reducing the error little by little
Training is a loop: measure "how wrong" the model is as a number, then nudge the numbers inside the model in the direction that makes that number smaller. The measure of how wrong it is is called the loss (error). The farther a prediction is from the correct answer, the bigger the loss.
The machine starts with numbers chosen at random. It makes a prediction, measures the loss, moves the numbers slightly in the direction that reduces the loss, and predicts again. Repeating this thousands to millions of times brings it close to the point where the loss stops going down. The method for calculating which way to move the numbers to reduce the loss is the gradient (gradient descent), covered in Math Basics, Lesson 10.
- Current prediction: 0.1 × 30 = 3 price units
- That is 3 price units below the actual price of 6, so the loss is large.
- The prediction needs to go up, so a is moved a little in the direction of getting bigger. For example, with a = 0.15 the prediction is 4.5 price units, closer to the actual price.
- Looking at this one sale alone, a = 0.2 makes the prediction exactly 6 price units (0.2 × 30 = 6). Real training looks at many sales at once and finds the a with the smallest total loss.
4.Evaluation: taking a test with new questions
You must not measure performance with the data used for training. That is like taking a test on questions you have already seen the answers to, so the score comes out higher than it really is. That is why the data is split from the start. Most of it is used for training, and a portion is set aside and used only for evaluation after training is finished.
The evaluation measure depends on the problem. For classification problems you look at how many it got right (accuracy), but accuracy alone can easily fool you. If only 5 out of 100 emails are spam, even a useless model that answers "all legitimate" has 95% accuracy. So you also look at "how many of the spam emails did it catch?" and "of the emails it called spam, how many really were spam?"
- It got all 95 legitimate emails right and all 5 spam emails wrong.
- Accuracy = number correct ÷ total = 95 ÷ 100 = 95%
- Share of spam caught = spam caught ÷ total spam = 0 ÷ 5 = 0%
- The accuracy is high, but as a spam filter it is completely useless.
5.Overfitting: the student who memorized past exams
Some models are nearly perfect on the training data but terrible on data they have never seen. This is called overfitting. It is like a student who memorized the answers to past exam questions and cannot solve a new question where only the numbers have changed. The model has memorized not only the patterns but also the random noise.
The opposite, where the model is so simple that it cannot even fit the training data well, is called underfitting. A good model sits in between: it does fairly well on the training data and about as well on data it has never seen. When the gap between the training score and the evaluation score grows large, suspect overfitting.
This principle matters for people who use AI, too. Just because an AI did brilliantly on some examples does not guarantee it will do brilliantly on every similar task. Testing it yourself on new examples that fit your own work is the most accurate evaluation.
📌 Key points
- Machine learning consists of four steps: data → model → training → evaluation
- Supervised learning learns from labeled examples; unsupervised learning finds structure without labels
- Training is a loop of measuring the loss (how wrong it is) and adjusting numbers in the direction that reduces it
- Always evaluate on data not used for training — looking at accuracy alone can fool you
- Overfitting is when a model memorizes the training data and becomes weak on new data
🤖 Try asking AI like this
Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.
When you want a feel for whether machine learning could be used in your work
I'm working on [description of the task or problem]. If this problem were solved with machine learning, suggest what the input data, the correct answers (labels) and the evaluation measure should each be. Also point out any biases that could creep in when collecting the data.
When you want to check whether you've really understood a concept
Explain overfitting using studying for an exam as an analogy. Then give me three questions that ask me to tell overfitting from underfitting, and grade my answers after I reply.
- High school "Fundamentals of Artificial Intelligence" course content (Korean national curriculum, Ministry of Education)
- Standard explanations in introductory machine learning textbooks (supervised learning, unsupervised learning, overfitting)
Reached every goal above? Mark the lesson complete.
Storage is unavailable in this browser, so this lasts only for this page.🤖 AI Basics
- 1What Is AI? — Rules and Learning
- 2Machine Learning Basics — Data, Model, Training, Evaluation
- 3Neural Networks, Intuitively — Small Calculations Add Up to Judgment
- 4How Do Large Language Models Produce Text?
- 5Prompt Basics — Saying Exactly What You Want
- 6Advanced Prompting — Roles, Examples, Steps, Format
- 7Hallucination and Verification — How to Doubt What Sounds Plausible
- 8Privacy, Copyright and Ethics — Using AI Responsibly
- 9Working with AI — Delegate, Check, Connect
- 10Learning Strategies for the AI Era — The Basics Shape Your Questions