1.Text is split into pieces called tokens
A language model does not read text letter by letter as it is. First it cuts the text into pieces called tokens. A token can be a whole word, part of a word, a few characters or a punctuation mark. Common words tend to become a single token, while rare words tend to be split into several tokens.
Next, each token is turned into a list of numbers (a vector). Through training, these lists of numbers are adjusted so that tokens with similar meanings have similar values. So inside the model, "puppy" and "dog" are close together, while "puppy" and "tax" are far apart. This is exactly the idea of vector similarity covered in Math Basics, Lesson 10.
Tokens matter in practice, too. Many AI services count both how much text they can read at once and how much you use them in tokens. Even for the same content, the number of tokens varies with the language and the wording.
- Each model splits tokens differently, so the following is just an example for illustration.
- Example: "The" " weather" " is" " really" " nice" " today"
- Common words like "weather" and "today" stay whole, while rarer words or word endings may be split off into separate pieces.
- The model turns each of these pieces into a list of numbers and then calculates with them.
2.The core principle: next-token prediction
In one sentence, what a language model does is "look at the tokens so far and calculate the probabilities of the token that comes next." After "The sky is so," tokens such as "clear," "blue" or "cloudy" might come next, each with a different probability. The model calculates this probability distribution, picks one of them, and attaches it to the end of the text.
Then it takes the text, now including the token it just added, as its new input and predicts the next token. It repeats this until the answer is finished. Even a long answer is ultimately the result of hundreds or thousands of these one-piece-at-a-time choices in a row.
To predict the next token well, a model needs to know not only grammar but also facts, logic, tone and context. To get what follows "The capital of South Korea is" right, it needs to know geography; to get the next line of a math solution right, it needs to know the rules of calculation. In the process of training on this prediction with an enormous amount of text, the model picks up a surprising amount of knowledge and patterns.
3.Following how one short answer gets made
Let's follow the principle all the way through once. Suppose a user types, "Describe a rainy day in one sentence." The model cuts this request into tokens and then chooses the tokens that follow, one at a time, starting with the first. All the probabilities below are made-up numbers for illustration.
There are also options for how to choose. If it always picks the token with the highest probability, the same question always gets the same answer, but the wording tends to be flat. If it sometimes picks less likely tokens, in proportion to their probability, the wording becomes more varied, but the answer differs a little each time. Many services have a setting that controls how much randomness there is.
This also clears up a common misconception. The model does not write a finished answer in its head first and then show it to you. The answer is built as the tokens chosen earlier keep narrowing down the choices that follow. So if the first few tokens head in the wrong direction, the model can carry on with that flow and plausibly add wrong content later on.
- It cuts the whole input into tokens and turns each one into a list of numbers.
- It calculates the probabilities for the first token. Example: "Outside" 0.30 · "Raindrops" 0.25 · "Gray" 0.15 · … and picks "Raindrops."
- It takes the text with "Raindrops" added as its new input and calculates the next token. Example: " tap" 0.60 · " fall" 0.20 · … and picks " tap."
- It repeats the same process, writing on until it reaches "Raindrops tap on the windowpane, and the streets slowly turn gray and wet."
- When a special token meaning "the answer is finished" becomes the most likely next token, it picks that token and stops generating.
4.Pretraining and additional training for conversation
Large language models are usually built in two stages. In the first stage, pretraining, the model learns the patterns of language and knowledge by repeating next-token prediction over a vast amount of text such as books, web pages and code. "Large" means that the amount of text it learned from and the number of weights inside the model are both very large.
A model that has only been pretrained is good at continuing text but clumsy at answering questions helpfully. So in the second stage it gets additional training on examples of "questions and good answers," and it is refined using evaluations in which people picked the better of several answers. After this process it becomes a model that follows instructions, holds conversations and refuses dangerous requests.
There is one important consequence. The model's knowledge stops at the point when its training data was collected. It does not know about events after that unless a search feature is added or the user provides the material. Keep this in mind whenever a question needs up-to-date information.
5.The context window: how much text it can see at once
In a conversation, a language model takes the whole conversation so far, plus any material you pasted in, as its input. The maximum number of tokens it can take in at once is called the context window. Basically, the model does not "remember" the conversation; each time, it rereads the entire conversation before it and produces the next answer.
So when a conversation gets very long and exceeds the context window, the earlier parts may be cut off or given less weight. Rather than stating an important condition once at the start of a conversation and leaving it at that, it is safer to remind the model briefly when needed. And starting a new topic in a new conversation gives cleaner results.
6.Strengths and weaknesses that follow from the principle
Because it picks the next token based on probability, the answer to the same question can vary slightly. This property is a strength when you want several ideas and a weakness when you need exactly the same result.
Also, the model picks "plausible next words"; it does not check facts. Content that appears often in its training data is usually accurate, but for content that is rare, recent or full of numbers, it can naturally produce sentences that are plausible but wrong. This is the hallucination we cover in Lesson 7.
On the other hand, it is very strong at tasks where "language patterns" are the core: changing the format of text, summarizing, explaining things more simply and laying out multiple perspectives. Once you understand the principle, you can tell which tasks to hand over and which to check.
| Property | Used well | Watch out for |
|---|---|---|
| It chooses probabilistically | You get several ideas | The answer can differ for the same question |
| It follows plausibility | Natural sentences, format conversion | It states wrong facts just as naturally |
| Its knowledge stops at training time | It knows widely known facts well | It may not know recent information |
| It sees only the text inside the context window | Paste in material and it answers based on that material | The early parts of a very long conversation may be given less weight |
📌 Key points
- A language model splits text into tokens and turns each token into a list of numbers (a vector) to calculate with
- An answer is made by repeating "calculate next-token probabilities → pick one" until the end
- Pretraining teaches it language and knowledge; additional training turns it into a conversational model that follows instructions
- The model's knowledge stops at training time, and it answers looking only at the text inside its context window
- Plausibility and fact can differ — check numbers, recent information and rare facts
🤖 Try asking AI like this
Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.
When you ask a question that needs up-to-date information
Before you answer the question below, first tell me whether this information might have changed since your training cutoff. For anything that might have changed, also tell me where I should check for the latest information. Question: [your question]
When you want an answer based on a long document
Answer using only the material below. Don't guess at anything that isn't in the material; write "Not in the material" instead. At the end of each sentence in your answer, add the number of the paragraph you based it on, in parentheses. [paste the material] Question: [your question]
- Standard explanations in introductory natural language processing textbooks (tokenization, language models, context)
- High school "Fundamentals of Artificial Intelligence" course content (Korean national curriculum, Ministry of Education)
Reached every goal above? Mark the lesson complete.
Storage is unavailable in this browser, so this lasts only for this page.🤖 AI Basics
- 1What Is AI? — Rules and Learning
- 2Machine Learning Basics — Data, Model, Training, Evaluation
- 3Neural Networks, Intuitively — Small Calculations Add Up to Judgment
- 4How Do Large Language Models Produce Text?
- 5Prompt Basics — Saying Exactly What You Want
- 6Advanced Prompting — Roles, Examples, Steps, Format
- 7Hallucination and Verification — How to Doubt What Sounds Plausible
- 8Privacy, Copyright and Ethics — Using AI Responsibly
- 9Working with AI — Delegate, Check, Connect
- 10Learning Strategies for the AI Era — The Basics Shape Your Questions