B Binance · The world's largest crypto exchangeBinance Sign up → AD OKX OKX · A leading global crypto exchangeOKX Sign up → AD
🤖 AI Basics · Lesson 2 / 10

Machine Learning Basics — Data, Model, Training, Evaluation

Machine learning means collecting data, choosing a model, training it to reduce its error, and evaluating it on data it has never seen. The final evaluation matters most.

⏱ About 15 min ✍️ 4 practice questions Updated 2026-10-08
🎯 By the end of this lesson you can
  • Explain the four steps of machine learning (data, model, training, evaluation) in order
  • Give an example of the difference between supervised and unsupervised learning
  • Explain why training data and evaluation data are kept separate
  • Say what overfitting is and why it is bad

1.The four steps of machine learning

In Lesson 1 we said machine learning is the approach in which a machine finds patterns in example data. If you break this process down a little further, almost every machine learning project goes through the same four steps: collect data, decide what shape of rule to look for (the model), adjust the model to fit the data (training), and check how well it does on data it has never seen (evaluation).

Take, for example, the problem of predicting an apartment's price from its size. A table of sizes and prices for apartments that have already sold is the data. A straight-line rule, "price = size times some number plus some other number," is the model. Trying different values for the number to multiply by and the number to add, and finding the values that make the gap from actual prices smallest, is training. Finally, measuring how accurate the predictions are on sales records that were not used in training is evaluation.

The four steps, seen through house-price prediction
StepWhat it doesIn house-price prediction
DataCollect inputs paired with correct answersSize (input) and actual sale price (correct answer)
ModelDecide the shape of the rule to look forA straight line of the form price = a × size + b
TrainingAdjust the model's numbers to make the error smallerChange a and b to shrink the gap between prediction and reality
EvaluationMeasure performance on data it has never seenCheck how accurate predictions are on sales not used in training

2.Data: examples with and without correct answers

Learning from data where every input has a correct answer (a label) attached is called supervised learning. It learns from examples that people have labeled, such as "this email is spam," "this photo is a cat" or "this house is worth 5 price units." Classification (is it spam or not?) and numeric prediction (how much is it? — also called regression) are typical supervised learning problems.

Giving the machine only data, with no correct answers, and having it group similar things or find structure is unsupervised learning. An online store grouping customers with similar tastes based only on their purchase histories is an example. Since nobody defined what the right answer is, a person ultimately has to judge whether the groups the machine found are meaningful.

Either way, the quality of the data determines the quality of the result. Examples with wrong labels, skewed examples and too few examples all make the model worse to the same degree. That is why people often say, "garbage in, garbage out."

  • Supervised learning: learns from labeled examples — spam classification, house-price prediction, photo classification
  • Unsupervised learning: finds structure without labels — grouping customers, clustering similar documents
  • Reinforcement learning: learns to increase the rewards it receives for its actions — games, robot control

3.Training: reducing the error little by little

Training is a loop: measure "how wrong" the model is as a number, then nudge the numbers inside the model in the direction that makes that number smaller. The measure of how wrong it is is called the loss (error). The farther a prediction is from the correct answer, the bigger the loss.

The machine starts with numbers chosen at random. It makes a prediction, measures the loss, moves the numbers slightly in the direction that reduces the loss, and predicts again. Repeating this thousands to millions of times brings it close to the point where the loss stops going down. The method for calculating which way to move the numbers to reduce the loss is the gradient (gradient descent), covered in Math Basics, Lesson 10.

ExampleThe model is "price = a × size," and an apartment of 30 area units actually sold for 6 price units. If training starts at a = 0.1 (price units per area unit), which way should training move a?
  1. Current prediction: 0.1 × 30 = 3 price units
  2. That is 3 price units below the actual price of 6, so the loss is large.
  3. The prediction needs to go up, so a is moved a little in the direction of getting bigger. For example, with a = 0.15 the prediction is 4.5 price units, closer to the actual price.
  4. Looking at this one sale alone, a = 0.2 makes the prediction exactly 6 price units (0.2 × 30 = 6). Real training looks at many sales at once and finds the a with the smallest total loss.
AnswerMove a in the direction of getting bigger (from 0.1 toward 0.2).

4.Evaluation: taking a test with new questions

You must not measure performance with the data used for training. That is like taking a test on questions you have already seen the answers to, so the score comes out higher than it really is. That is why the data is split from the start. Most of it is used for training, and a portion is set aside and used only for evaluation after training is finished.

The evaluation measure depends on the problem. For classification problems you look at how many it got right (accuracy), but accuracy alone can easily fool you. If only 5 out of 100 emails are spam, even a useless model that answers "all legitimate" has 95% accuracy. So you also look at "how many of the spam emails did it catch?" and "of the emails it called spam, how many really were spam?"

ExampleOut of 100 emails, 5 are spam. A model classified every email as "legitimate." What are its accuracy and the share of spam it caught?
  1. It got all 95 legitimate emails right and all 5 spam emails wrong.
  2. Accuracy = number correct ÷ total = 95 ÷ 100 = 95%
  3. Share of spam caught = spam caught ÷ total spam = 0 ÷ 5 = 0%
  4. The accuracy is high, but as a spam filter it is completely useless.
AnswerAccuracy 95%, spam detection rate 0%

5.Overfitting: the student who memorized past exams

Some models are nearly perfect on the training data but terrible on data they have never seen. This is called overfitting. It is like a student who memorized the answers to past exam questions and cannot solve a new question where only the numbers have changed. The model has memorized not only the patterns but also the random noise.

The opposite, where the model is so simple that it cannot even fit the training data well, is called underfitting. A good model sits in between: it does fairly well on the training data and about as well on data it has never seen. When the gap between the training score and the evaluation score grows large, suspect overfitting.

This principle matters for people who use AI, too. Just because an AI did brilliantly on some examples does not guarantee it will do brilliantly on every similar task. Testing it yourself on new examples that fit your own work is the most accurate evaluation.

The most common way to reduce overfitting is to collect more, and more varied, data. Making the model simpler or stopping training at the right point are also used.

📌 Key points

  • Machine learning consists of four steps: data → model → training → evaluation
  • Supervised learning learns from labeled examples; unsupervised learning finds structure without labels
  • Training is a loop of measuring the loss (how wrong it is) and adjusting numbers in the direction that reduces it
  • Always evaluate on data not used for training — looking at accuracy alone can fool you
  • Overfitting is when a model memorizes the training data and becomes weak on new data

✍️ Practice questions

Answer first, then open "Answer and explanation".

Q1. Grouping similar customers using only their purchase histories, with no correct answers, is which kind of learning?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② Unsupervised learning

It finds structure in the data without correct answers (labels), so it is unsupervised learning.

Q2. What is the main reason you shouldn't use the training data as-is to evaluate a model?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② It has already seen that data, so performance looks better than it really is

It is like taking a test on questions you have already seen the answers to, so you cannot tell the real performance on data it has never seen.

Q3. A model has 99% accuracy on the training data but 60% on the evaluation data. What state is it most likely in?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② Overfitting

This is classic overfitting: the model fits the training data too closely, and its performance drops sharply on new data.

Q4. Out of 1,000 patients, 10 have a disease. What is the accuracy (in %) of a model that classifies everyone as "healthy"?

Answer and explanation
Answer 99%

It gets 990 people right, so 990 ÷ 1000 = 99%. But it does not find a single one of the 10 patients, so it is a useless model. For problems about finding rare cases, you must not judge by accuracy alone.

🤖 Try asking AI like this

Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.

When you want a feel for whether machine learning could be used in your work

I'm working on [description of the task or problem]. If this problem were solved with machine learning, suggest what the input data, the correct answers (labels) and the evaluation measure should each be. Also point out any biases that could creep in when collecting the data.

When you want to check whether you've really understood a concept

Explain overfitting using studying for an exam as an analogy. Then give me three questions that ask me to tell overfitting from underfitting, and grade my answers after I reply.
References
  • High school "Fundamentals of Artificial Intelligence" course content (Korean national curriculum, Ministry of Education)
  • Standard explanations in introductory machine learning textbooks (supervised learning, unsupervised learning, overfitting)

Reached every goal above? Mark the lesson complete.

🤖 AI Basics

  1. 1What Is AI? — Rules and Learning
  2. 2Machine Learning Basics — Data, Model, Training, Evaluation
  3. 3Neural Networks, Intuitively — Small Calculations Add Up to Judgment
  4. 4How Do Large Language Models Produce Text?
  5. 5Prompt Basics — Saying Exactly What You Want
  6. 6Advanced Prompting — Roles, Examples, Steps, Format
  7. 7Hallucination and Verification — How to Doubt What Sounds Plausible
  8. 8Privacy, Copyright and Ethics — Using AI Responsibly
  9. 9Working with AI — Delegate, Check, Connect
  10. 10Learning Strategies for the AI Era — The Basics Shape Your Questions
📚 Worth reading
🧠What Generative AI Does and Where It Fails→ ✍️How to Write a Good Prompt→ 🔍Checking AI Answers→ 📚Using AI for Study and Work Without Plagiarism→
← Foundations for the AI Era