B Binance · The world's largest crypto exchangeBinance Sign up → AD OKX OKX · A leading global crypto exchangeOKX Sign up → AD
💻 Coding Basics · Lesson 7 / 10

Strings and Text — Cutting, Finding and Replacing Characters

A string is a value made of characters in a row, and JavaScript comes with built-in methods for cutting, finding, splitting and replacing. The key ideas are that methods never change the original string but return a new one, and that `length` is not always the same as the number of characters you see.

⏱ About 20 min ✍️ 5 practice questions 🧪 4 code exercises Updated 2026-10-09
🎯 By the end of this lesson you can
  • Take part of a string with length, indexes and slice
  • Use indexOf, includes and startsWith to find whether and where some text appears
  • Split, tidy and change text with split, join, trim and replaceAll
  • Explain why length can differ from the visible character count for accented letters and emoji, and read a simple regular expression

1.A string is a row of characters

As we saw in Lesson 2, a string is a value wrapped in quotes. Double quotes "apple" and single quotes 'apple' make exactly the same string, and a template string (template literal), wrapped in backticks, is used when you want to drop a variable's value into the text. Check what a template string looks like in the code below. Whichever quotes you use, sticking to one style within a program makes it easier to read.

Because a string is a row of characters in order, you can take out one character at a time by its number (index), just like the arrays in Lesson 6. The first character is number 0, and s.length is the length counted in characters. So the last character is s[s.length - 1], and nowadays you can also count from the end with s.at(-1). Asking for a number that does not exist gives undefined, not an error — again just like arrays.

There is one big difference from arrays. Once a string is created, the characters inside it cannot be changed. If you try to fix one character with s[0] = "z", code running in strict mode — as this playground does — throws a TypeError. That is why every method in this lesson leaves the original string alone and returns a 'new string'. To use the result, you must store it in a variable.

Indexes start at 0; a missing number gives undefined
const name = "Alex";
const fruit = 'apple';
const msg = `Hi ${name}, ${fruit} x3`;
console.log(msg);
console.log(msg.length);
console.log(msg[0], msg.at(-1), msg[99]);
Output
Hi Alex, apple x3
17
H 3 undefined
Methods return a new string — if you don't store it, nothing changes
let s = "  hello  ";
s.trim();
console.log("[" + s + "]");
s = s.trim();
console.log("[" + s + "]");
Output
[  hello  ]
[hello]
If you called a method and the result looks unchanged, first check whether you stored the returned value in a variable. It is the most common beginner mistake.

2.The character-count trap — accents and emoji

length is not 'the number of characters you see'. It is the number of units JavaScript uses to store the string (UTF-16 code units). Luckily, plain English letters, digits, spaces and punctuation are one unit each, and so are common accented letters such as é, so "café".length is 4 and "hello".length is 5, just as they look. Most everyday Chinese, Japanese and Korean characters are also one unit each.

Emoji are different. Many emoji are stored as two units, so "😀".length is 2. Emoji built by joining several pieces into one picture, like a hand with a skin tone, can come out as 4 or more. If you spread a string with [...s], it is split into code points, so a single 😀 counts as 1, but a joined emoji still counts as however many pieces it has. To count exactly by the shapes you see on screen, you need a separate tool such as Intl.Segmenter.

Accented letters are not always safe either. The same é can arrive in a decomposed form (NFD): a plain e followed by a separate combining accent mark. Then its length is 2. File names copied from some programs or systems come in this form, so two words that look identical can give false with ===. Korean text has the same issue: one syllable can be stored as three separate pieces. In these cases, bring both sides to one form with s.normalize("NFC") before comparing.

One more thing: character count and byte count are different. In UTF-8, the encoding most files and networks use, an English letter is 1 byte, é is 2 bytes, most Chinese, Japanese and Korean characters are 3 bytes and many emoji are 4 bytes. If an input field is limited in bytes, say 'up to 90 bytes', you must not count with length. In short: for ordinary text, counting with length is fine, but once emoji, decomposed letters or byte limits are involved, decide first what exactly you are counting.

Actual results from running this in the playground
console.log("café".length, "hello".length);
console.log("😀".length, [..."😀"].length);
console.log("👍🏽".length, [..."👍🏽"].length);
Output
4 5
2 1
4 2
Decomposed letters and UTF-8 byte counts
const a = "é";
const b = a.normalize("NFD");
console.log(a.length, b.length);
console.log(a === b);
console.log(a === b.normalize("NFC"));
const bytes = new TextEncoder().encode("café");
console.log(bytes.length);
Output
1 2
false
true
5
The 'character count' depends on what you count (values checked in this playground)
String`length``[...s].length`UTF-8 bytes
"café"445
"abc"333
"😀"214
"👍🏽"428

3.Cutting and finding — `slice`, `indexOf`, `includes`

s.slice(start, end) cuts from the start number up to 'just before' the end number and returns it as a new string. Leave out the end and it cuts to the last character; give a negative number and it counts from the end. So s.slice(0, 3) is the first three characters and s.slice(-4) is the last four. The rule that the end number is not included is the same feeling as writing i < n in the loops of Lesson 4, and it means the length is simply 'end − start'.

For finding, three methods do most of the work. s.indexOf("text") returns the position where the text first appears, or -1 if it is not there. If you don't need the position and only want to know whether it is there, s.includes("text") returns true/false, which drops straight into the conditions of Lesson 3. startsWith and endsWith check whether the string begins or ends with that text — for example, whether a file name ends with ".csv".

Using the result of indexOf directly as a condition is a trap. If the text is at the very start, the result is 0, and as we saw in Lesson 3, 0 is falsy, so if (s.indexOf("x")) is false exactly when it is at the front. When you only want to know whether it is there, use includes; if you use indexOf, compare with !== -1. These methods are also case-sensitive, so "Apple".includes("apple") is false. To ignore case, convert both sides with toLowerCase() before searching.

The end number is not included
const phone = "010-1234-5678";
console.log(phone.slice(0, 3));
console.log(phone.slice(-4));
console.log(phone.slice(4, 8));
Output
010
5678
1234
const file = "2026_sales.csv";
console.log(file.indexOf("_"));
console.log(file.indexOf("x"));
console.log(file.includes("sales"));
console.log(file.endsWith(".csv"));
console.log("Apple".includes("apple"));
const t = "Apple".toLowerCase();
console.log(t.includes("apple"));
Output
4
-1
true
true
false
true
ExampleHow would you take only the user name before the @ in the sample address "[email protected]"?
const email = "[email protected]";
const at = email.indexOf("@");
if (at === -1) {
  console.log("Not an email address");
} else {
  console.log(email.slice(0, at));
}
Output
alex
  1. Step 1 Find the position of @. indexOf("@") is 4 (a-l-e-x are positions 0–3).
  2. Step 2 Cut from position 0 up to 'just before' position 4. That is slice(0, 4), and the length is 4 − 0 = 4 characters.
  3. Step 3 If an address without @ comes in, indexOf gives -1, and slice(0, -1) returns the whole thing minus its last character — a wrong value. So add a condition that checks for -1 first.
Answeremail.slice(0, email.indexOf("@")) — but filter out the case with no @ first.

4.Split, join and tidy — `split`, `join`, `trim`

Half of text processing is split and join. s.split(",") cuts the string at each comma and returns an array, and arr.join("-") does the opposite, joining an array into one string. Once a string becomes an array, you can use everything from Lesson 6 — length, map, filter — so 'split → work on the array → join again' becomes the basic flow. The automation examples in Lesson 9 are built on this flow too.

split has some edges worth knowing. Splitting the empty string "" gives not an empty array but an array holding one empty string, [ '' ]. Two commas in a row leave an empty string between them. So if you count words with split(" ").length, an empty input counts as 1 and places with doubled spaces inflate the count. That is why you should always test edge cases as well.

trim() removes spaces, tabs and line breaks from both ends (spaces in the middle stay). Values typed by people often arrive with extra spaces at the ends, so it is a good habit to pass them through trim() before comparing or saving. toUpperCase() and toLowerCase() change the case of letters; characters that have no case, such as digits, punctuation or Chinese characters, stay as they are. padStart(3, "0") fills the front with "0" until the length is 3, turning "7" into "007". It is handy for showing numbers or times with a fixed number of digits.

Split gives an array, join gives a string
const line = "Alex,apple,3";
const cols = line.split(",");
console.log(cols);
console.log(cols.length, cols[1]);
console.log(cols.join(" / "));
Output
[ 'Alex', 'apple', '3' ]
3 apple
Alex / apple / 3
Splitting an empty string gives [ '' ]
console.log("".split(","));
console.log("a,,b".split(","));
console.log("  Hi  ".trim().toUpperCase());
console.log("7".padStart(3, "0"));
console.log("no.42 ok".toUpperCase());
Output
[ '' ]
[ 'a', '', 'b' ]
HI
007
NO.42 OK
  • split("") (splitting on the empty string) cuts the string into single units. Emoji can break into pieces, so when you split into characters, [...s] is safer.
  • split, trim and toUpperCase all return new values. The original string stays the same.
  • To split into lines, use split("\n"). Files made on Windows may have one extra invisible character at the end of each line, so it is safer to trim() each line after splitting.

5.Replacing — `replace` and `replaceAll`

s.replace("find", "new") replaces only the first match. To replace every match, use replaceAll. This difference often leads to 'why did only the first one change?', so whenever you write replacing code, test it with an input where the same text appears two or more times. Both methods leave the original string untouched and return a new, changed string.

Replacing also works as 'deleting'. If the replacement is the empty string "", the matched parts disappear. For example, "010-1234-5678".replaceAll("-", "") is "01012345678". Going the other way, to add hyphens to a digits-only phone number, it is cleaner to cut it with slice and put it together with join. Whatever you do, the original stays as it is and a new value is made, so it is easy to put the before and after side by side and compare.

replace changes only the first one
const s = "apple-pear-apple";
console.log(s.replace("apple", "plum"));
console.log(s.replaceAll("apple", "plum"));
console.log(s);
Output
plum-pear-apple
plum-pear-plum
apple-pear-apple
Delete, then join again
const raw = "010-1234-5678";
const digits = raw.replaceAll("-", "");
console.log(digits);
const parts = [
  digits.slice(0, 3),
  digits.slice(3, 7),
  digits.slice(7),
];
console.log(parts.join("-"));
Output
01012345678
010-1234-5678
When the text before and after replacing is long, it is hard to compare by eye. Paste both into a text diff checker and only the changed parts are highlighted, so you can quickly see whether anything you didn't intend was changed.

6.A taste of regular expressions — finding by shape

So far we have searched for 'exactly this text'. When you want to search by shape — 'a run of digits', 'a place where several spaces follow each other' — you use a regular expression (regex). In JavaScript you write the pattern between two / characters, and a g after it means 'all matches, not just the first'. Regular expressions are a small language of their own, so here we only look at a few common pieces, enough to read them.

\d is one digit, \s is one whitespace character (space, tab or line break), and a + after something means 'one or more in a row'. So /\d+/g means 'every run of digits' and /\s+/g means 'every run of whitespace'. match returns the found pieces as an array — and note that when nothing is found it returns null, not an empty array. test only tells you whether something was found, as true/false.

Regular expressions are powerful but hard to read, and a small mistake makes them match the wrong places or miss the right ones. There are two rules of thumb. First, if includes, split or replaceAll can do the job, use them, and reach for a regex only when you really need one. Second, whether a regex came from an AI or you wrote it yourself, check it in a regex tester with several examples that should match and several that should not before you use it.

match gives null when nothing is found
const memo = "ordered 3, returned 12";
console.log(memo.match(/\d+/g));
console.log("no digits".match(/\d+/g));
console.log(/\d/.test(memo));
Output
[ '3', '12' ]
null
true
Squash runs of spaces into one, hide the digits
const messy = "  apple   pear  plum ";
const clean = messy.trim().replace(/\s+/g, " ");
console.log("[" + clean + "]");
const masked = "010-1234-5678".replace(/\d/g, "*");
console.log(masked);
Output
[apple pear plum]
***-****-****
Regex pieces you will see often
PieceMeaningExample
\done digit/\d+/g → every run of digits
\sone whitespace character (including tabs and line breaks)/\s+/g → every run of whitespace
+the thing before it, one or more times/a+/ → a, aa, aaa
^, $start and end of the string/^\d+$/ → only when it is all digits
gfind all (without it, only the first)s.replace(/\d/g, "*")

📌 Key points

  • Strings can't be changed, so every method returns a new string — you must store the result in a variable
  • slice(start, end) does not include the end, and indexOf gives -1 when nothing is found — use includes when you only need to know whether it is there
  • 'Split (split) → work on the array → join (join)' is the basic flow of text processing, and splitting an empty string gives [ '' ]
  • replace changes only the first match, replaceAll changes them all
  • length counts UTF-16 units — a plain or accented letter is 1, but an emoji is 2 or more, a decomposed é is 2, and the UTF-8 byte count is different again

✍️ Practice questions

Answer first, then open "Answer and explanation".

Q1. What does const s = "abc"; s.toUpperCase(); console.log(s); print?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ① abc

toUpperCase() only returns a new string "ABC"; it does not change s. You need to store the result, as in s = s.toUpperCase().

Q2. What is the value of "010-1234-5678".slice(-4)?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ① "5678"

Negative numbers count from the end. It runs from -4 to the end, so it is the last four characters, "5678".

Q3. When does if (s.indexOf("a")) { ... } behave differently from what you intended?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ① When "a" is at the very start

At the very start the result is 0, and 0 is falsy, so the condition is false. Conversely, the -1 you get when it is missing is truthy, so the condition is true. Use includes or compare with !== -1.

Q4. What is the result of "apple,pear,apple".replace("apple", "plum")?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ① "plum,pear,apple"

When you give replace a string, it changes only the first match. To change them all, use replaceAll.

Q5. "hi".length is 2, and "😀".length is also 2. Explain why in one sentence.

Answer and explanation
Answer length counts UTF-16 code units, not visible characters; each plain letter is one unit, while this emoji is stored as two units.

[..."😀"].length is 1. Emoji made by joining several pieces can still give 2 or more even when spread like this, so to count by visible shapes you need a tool such as Intl.Segmenter.

🧪 Code lab

Build functions that cut, count and tidy text yourself. Edge-case tests, such as the empty string, are graded too.

Your code runs only inside an isolated sandbox in this browser and is never sent to a server. It has no network access and is stopped after 2 seconds. Edited code is saved only in this browser. Ctrl+Enter (⌘+Enter) runs it; Tab inserts two spaces (press Esc, then Tab, to move on).

JavaScript is off, so the code can't run here, but you can still read each task, its starter code, the automatic checks and a sample solution.

1Masking a name

Write a function maskName(name) that takes a name name and returns a string that keeps only the first character and turns every other character into *. Example: "Sam" → "S**". A one-character name stays as it is, and an empty string returns an empty string.

Automatic checks
  • maskName("Sam")expected "S**"
  • maskName("Alex")expected "A***"
  • maskName("J")expected "J"
  • maskName("")expected ""
💡 Hint

For an empty string, name.length - 1 is -1, and "*".repeat(-1) throws a RangeError. Guard against it with Math.max(…, 0) or a condition.

Show a sample solution
function maskName(name) {
  const first = name.slice(0, 1);
  const rest = name.length - 1;
  return first + "*".repeat(Math.max(rest, 0));
}

2How many times does a word appear?

Write a function countWord(text, word) that counts how many times the word word appears in the sentence text and returns the number. It is case-sensitive. Example: countWord("apple pear apple", "apple") → 2. If text is an empty string, the answer is 0. Count it even when it is inside another word (the "apple" in "pineapple" counts once). You may assume word is not an empty string.

Automatic checks
  • countWord("apple pear apple", "apple")expected 2
  • countWord("apple pear apple", "plum")expected 0
  • countWord("a-a-a-a", "a")expected 4
  • countWord("Apple apple", "apple")expected 1
  • countWord("", "apple")expected 0
💡 Hint

Even when "a-a".split("a") leaves empty pieces at both ends, as in [ '', '-', '' ], the number of pieces is still 'times it appears + 1'.

Show a sample solution
function countWord(text, word) {
  return text.split(word).length - 1;
}

3Tidying up spaces

Write a function cleanSpaces(s) that removes the spaces at both ends of the string s and shrinks every run of several spaces in the middle to a single space. Example: " apple pear " → "apple pear". A string with only spaces becomes an empty string.

Automatic checks
  • cleanSpaces(" apple pear ")expected "apple pear"
  • cleanSpaces("a b c")expected "a b c"
  • cleanSpaces("already clean")expected "already clean"
  • cleanSpaces(" ")expected ""
  • cleanSpaces("")expected ""
💡 Hint

Splitting with split(/\s+/) and joining with join(" ") also works. Even then, trim() first, or you get empty pieces at both ends.

Show a sample solution
function cleanSpaces(s) {
  return s.trim().replace(/\s+/g, " ");
}

4Turning a title into a URL name

Write a function toSlug(title) that turns an English title title into a name that works well in a web address (a slug). Rules: remove spaces at both ends → make everything lowercase → turn each run of spaces into a single hyphen -. Example: " Hello World " → "hello-world".

Automatic checks
  • toSlug(" Hello World ")expected "hello-world"
  • toSlug("My First Post")expected "my-first-post"
  • toSlug("JS")expected "js"
  • toSlug("a b c d")expected "a-b-c-d"
  • toSlug("")expected ""
💡 Hint

In the starter code, replace(" ", "-") changes only the first space. You can chain methods with dots (.), and the order matters.

Show a sample solution
function toSlug(title) {
  return title
    .trim()
    .toLowerCase()
    .replace(/\s+/g, "-");
}

🤖 Try asking AI like this

Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.

When you want code that cleans up text

I want a JavaScript function that cleans up a string. Example inputs: [2–3 example inputs]. What I want: [the expected result for each]. Unless a regular expression is really needed, please use basic methods like `split`, `replaceAll` and `trim`, and tell me what happens with an empty string and with input that is only spaces.

When you don't understand a regex an AI gave you

Break the following regular expression into pieces and explain what each one means: [regex]. Give me 3 examples that match it and 3 examples that look like they should match but don't. What I actually want to find is [the shape I'm looking for] — point out any cases where this regex behaves differently from that.

When a character count comes out different from what you expected

In JavaScript, the `length` of [example string] comes out as [value I got]. The number of characters I see on screen is [visible count]. Explain why they differ in terms of UTF-16, code points and bytes, and tell me which method to use when what I need to count is [characters / bytes].

🧰 Related tools

Tools for checking and tidying the JSON, regular expressions and text from this lesson. Don't paste real personal data or passwords.

References
  • MDN Web Docs — String (JavaScript reference)
  • MDN Web Docs — Regular expressions guide
  • MDN Web Docs — String.prototype.normalize(), Intl.Segmenter
  • ECMAScript Language Specification (ECMA-262) — String objects

Reached every goal above? Mark the lesson complete.

💻 Coding Basics

  1. 1What Is a Program? — Breaking Work into Sequence, Choice, and Repetition
  2. 2Values, Variables, and Types — Putting Name Tags on Values
  3. 3Conditionals — Taking Different Paths Depending on the Situation
  4. 4Loops — Doing the Same Thing Many Times, Exactly
  5. 5Functions — Splitting Work into Small Named Machines
  6. 6Arrays and Objects — Storing and Handling Data
  7. 7Strings and Text — Cutting, Finding and Replacing Characters
  8. 8Debugging and Reading Errors — Turning Red Text into Clues
  9. 9Simple Automation — Spreadsheet Math and Text Cleanup in Code
  10. 10Reading and Verifying AI-Written Code — Tests, Security, Licenses, Privacy
📚 Worth reading
📊Using AI for Spreadsheets, and Checking the Formulas→ 💻AI Coding Assistants: Security and Licence Risks→ 💬Practising Conversation in a Foreign Language with AI→ ⚙️Finding Work Tasks Worth Automating with AI→
← Foundations for the AI Era