Learn JS Series (#5) - Strings: Template Literals, Unicode, and the Methods You Actually Use

Words
2595
Reading
12 min
Listen
Play
1M

Learn JS Series (#5) - Strings: Template Literals, Unicode, and the Methods You Actually Use

js-banner.png

What will I learn

  • You will learn template literals in depth, including multi-line strings and expression interpolation;
  • why strings are immutable, and what that means for every method you call;
  • the everyday string methods worth memorizing, and how they always return new strings;
  • how JavaScript handles Unicode, and why .length sometimes lies about "characters";
  • practical patterns for searching, slicing, splitting, and cleaning up text;
  • how JavaScript's string model compares to Python, Rust, Go and C, so the design choices make sense.

Requirements

  • A working modern computer running macOS, Windows or Ubuntu;
  • An installed Node.js (20+) distribution, or just a modern browser console;
  • Episodes 1-4 read, so operators and types are familiar.

Difficulty

  • Beginner

Curriculum (of the Learn JS Series):

Learn JS Series (#5) - Strings: Template Literals, Unicode, and the Methods You Actually Use

Solutions to Episode 4 Exercises

As always, we open with worked solutions to last episode's three exercises. Run them, then compare against your own attempts -- the small gaps are exactly where the understanding sharpens.

Exercise 1 - a leap-year-candidate test with %:

console.log(2026 % 4 === 0); // false - 2026 divided by 4 leaves remainder 2
console.log(2027 % 4 === 0); // false - 2027 leaves remainder 3
console.log(2028 % 4 === 0); // true  - 2028 divides evenly by 4

The insight: 2026 % 4 is 2, which is not 0, so 2026 is not a leap-year candidate. The full leap-year rule also drags in 100 and 400 (a year divisible by 100 is NOT a leap year unless it is also divisible by 400), a nice puzzle to revisit once we have proper control flow.

Exercise 2 - four expressions, predicted and then checked:

console.log("3" + 4);   // "34"  - + with a string concatenates
console.log(3 + 4);     // 7     - both numbers, real addition
console.log("3" * 4);   // 12    - * has no string meaning, so "3" becomes 3
console.log(3 === "3"); // false - strict equality, different types

The insight: + is the only arithmetic operator that also means "join text", so * is forced to treat "3" as the number 3, while + pulls the number 4 toward the string. That single split personality of + is the seed of today's whole topic.

Exercise 3 - a nickname with a fallback:

function displayName(nick) {
  return nick || "anonymous";
}
console.log(displayName("scipio")); // "scipio"
console.log(displayName(""));        // "anonymous"

The insight: with ||, the empty string "" is falsy, so it triggers the fallback. Had you used ?? in stead, "" would be kept (it is not null or undefined) and an empty name would slip straight through. Which one is correct depends entirely on whether an empty string should count as "no name". Choosing on purpose is the skill.

Right. Last episode we made values do things with operators. Today the value under the microscope is text, and you will handle text in nearly every program you ever write -- reading input, formatting output, parsing files, building messages. So let us get properly comfortable with strings.

Template literals, properly

We met backtick strings briefly a couple of episodes ago. Now we use them in full. Their two superpowers are interpolation (dropping values into text) and effortless multi-line support.

Interpolation first. Inside a backtick string, ${...} is a hole you pour a value into:

const user = "scipio";
const episodes = 5;
const summary = `${user} has published ${episodes} episodes so far.`;
console.log(summary); // "scipio has published 5 episodes so far."

The crucial detail is that anything between the braces is a full expression, not merely a variable name. Remember from last episode that an expression is any code that produces a value, so you can compute right there inside the text:

const price = 20;
const qty = 3;
console.log(`Total: ${price * qty} EUR`);                     // "Total: 60 EUR"
console.log(`Status: ${qty > 0 ? "in stock" : "sold out"}`);  // "Status: in stock"
console.log(`Name in caps: ${user.toUpperCase()}`);           // "Name in caps: SCIPIO"

That last line even calls a method inside the placeholder. The template literal does not care what the expression is -- it evaluates it, converts the result to a string, and drops it in.

And multi-line text just works. Newlines inside the backticks are preserved literally, no \n gymnastics required:

const letter = `Dear reader,

Thanks for following along.
- scipio`;
console.log(letter);

Compare that to the old way, where you would have written "Dear reader,\n\nThanks..." and manually counted your newline escapes. The backtick version is what you actually see, which is far easier to get right. Template literals are the modern default for building strings; reach for them over + concatenation almost every time. They are clearer, they handle numbers without you thinking about it, and they never accidentally do math the way + can.

Strings are immutable

Here is a single fact that governs every string method in the language: strings cannot be changed. Once a string exists, it is frozen. Every method that appears to modify a string actually leaves the original untouched and hands you back a brand new string:

const original = "hello";
const shouted = original.toUpperCase();
console.log(shouted);  // "HELLO" - a new string
console.log(original); // "hello" - the original is untouched

You also cannot poke a single character in place. Indexing a string is read-only:

const word = "cat";
// word[0] = "b"; // does nothing in sloppy mode, THROWS in strict mode
console.log(word); // "cat" - unchanged either way

Why does the language work this way? Immutability makes strings safe to share freely -- if you pass a string to a function, you never have to worry that the function secretly rewrote it under you. The trade is that every "edit" allocates a new string, but engines are heavily optimised for exactly this, so in practice you rarely pay a noticeable price.

The practical consequence is a rule you must burn into muscle memory: always capture the return value of a string method. Forgetting is the single most common beginner string bug. This does nothing useful:

let messy = "  scipio  ";
messy.trim();          // computes "scipio", then throws it away!
console.log(messy);    // "  scipio  " - still has the spaces

Whereas this is what you meant:

let messy2 = "  scipio  ";
messy2 = messy2.trim(); // capture the new string back into the variable
console.log(messy2);    // "scipio"

Same method, completely different outcome, and the only difference is whether you kept the result. Keep this in the back of your head for the rest of the episode: every method below returns something, and ignoring that return value is almost always a mistake.

The methods you actually use

There are dozens of string methods in the standard library, but a surprisingly small set covers the overwhelming majority of real work. Let us tour the ones that earn their keep.

Getting length and reaching individual characters:

const s = "javascript";
console.log(s.length);      // 10
console.log(s[0]);          // "j" - bracket index access
console.log(s.at(-1));      // "t" - at() understands negative indices, unlike []
console.log(s.charAt(2));   // "v" - the older way to index

Note at(-1): the modern at method lets you count from the end with a negative index, which is a lot nicer than writing s[s.length - 1].

Changing case and trimming whitespace, which you will lean on constantly when cleaning up user input:

console.log("  Hello  ".trim());        // "Hello"     - both ends
console.log("  Hello  ".trimStart());   // "Hello  "   - left only
console.log("loud".toUpperCase());      // "LOUD"
console.log("QUIET".toLowerCase());     // "quiet"

Searching within a string, to ask "is this in there, and where":

const phrase = "learn javascript today";
console.log(phrase.includes("script")); // true  - is the substring present?
console.log(phrase.startsWith("learn")); // true
console.log(phrase.endsWith("today"));   // true
console.log(phrase.indexOf("java"));     // 6     - position, or -1 if absent
console.log(phrase.indexOf("python"));   // -1    - not found

The -1 from indexOf is a classic gotcha: it is truthy-ish math but means "not found". These days includes is the readable way to ask a yes/no question, and you reach for indexOf only when you actually need the position.

Extracting a piece with slice. You pass a start index and an optional end index, where the end is not included:

const text = "javascript";
console.log(text.slice(0, 4));  // "java"   - indices 0,1,2,3 (4 is excluded)
console.log(text.slice(4));     // "script" - from index 4 to the end
console.log(text.slice(-6));    // "script" - negative counts back from the end

That "end is excluded" convention feels odd at first, but it has a lovely property: slice(0, 4) gives you exactly 4 characters, and slice(a, b) gives you b - a characters. Once that clicks, off-by-one slicing bugs mostly disappear.

Replacing and repeating:

console.log("a-b-c".replace("-", " "));    // "a b c" - replace ONLY replaces the first match
console.log("a-b-c".replaceAll("-", " ")); // "a b c" - replaceAll gets every match
console.log("na".repeat(4) + " batman");   // "nananana batman" - four na's

Watch that first one: plain .replace with a string argument swaps only the first occurrence, which surprises people who expect it to change all of them. When you mean "every match", use replaceAll. (There is also a regular-expression form of replace, which we will meet much later.)

And the pairing you will use endlessly -- split turns a string into an array, and join turns an array back into a string:

const csv = "scipio,7,amsterdam";
const parts = csv.split(",");
console.log(parts);              // ["scipio", "7", "amsterdam"]
console.log(parts.length);       // 3
console.log(parts.join(" | "));  // "scipio | 7 | amsterdam"

split/join are the bread and butter of text processing: split a line into fields, do some work on the array (which we will cover soon when we reach arrays properly), then join it back into a string for output.

A small worked example: cleaning a messy field

Let us chain a few of those together the way you actually would in real code. Say a form hands you a username that is padded with spaces, inconsistently cased, and uses underscores where you want spaces. You want to normalise it in one clean pass:

function normalizeName(raw) {
  return raw
    .trim()            // strip the outer whitespace
    .toLowerCase()     // fold to a consistent case
    .replaceAll("_", " "); // underscores become spaces
}

console.log(normalizeName("  Scipio_NL  ")); // "scipio nl"
console.log(normalizeName("ADA_Lovelace"));  // "ada lovelace"

Because every method returns a new string, you can chain them: the result of trim() flows into toLowerCase(), whose result flows into replaceAll. This "pipeline" reads top to bottom like a recipe. And notice we never mutated raw -- immutability is what makes chaining safe, since each step just produces the next string without disturbing anything before it.

Unicode, and why .length can lie

Now a topic most beginner tutorials skip entirely, and it bites people badly in production: a JavaScript string is a sequence of UTF-16 code units, not "characters" in the intuitive human sense. For plain English (and most Latin-script text) this distinction never matters, and .length behaves exactly as you expect:

console.log("cafe".length);   // 4 - fine
console.log("rocket".length); // 6 - fine
console.log("scipio".length); // 6 - fine

The trouble starts with characters that live outside the original 16-bit range, which includes most emoji and many less common symbols. Those are stored as a surrogate pair -- two UTF-16 code units glued together to represent one character. So .length counts them as 2, and indexing by number can slice one clean in half. Watch (I am writing the rocket emoji with its code-point escape \u{1F680} so this stays pure ASCII in the source, but it is a single rocket on screen):

const rocket = "\u{1F680}";     // one rocket emoji
console.log(rocket.length);     // 2  - it is TWO UTF-16 code units!
console.log(rocket[0]);         // a broken half of the character (a lone surrogate)
console.log("a\u{1F680}b".length); // 4, not 3 - the emoji counts as two

So "a<rocket>b".length reports 4 even though a human sees three characters. When you genuinely need to walk real user-perceived characters, use for...of or the spread operator ..., both of which iterate by code point and keep surrogate pairs intact:

const text = "a\u{1F680}b";
for (const ch of text) {
  console.log(ch); // "a", then the WHOLE rocket, then "b" - three iterations
}
console.log([...text].length); // 3 - spread splits into code points, not code units

You do not need to master Unicode today -- it is a genuinely deep rabbit hole (there are even single "characters" that are themselves built from several code points). Just carry the warning with you: .length measures UTF-16 code units, which is not always what a human calls a character. When you are counting emoji, validating a tweet-length limit, or reversing a string, reach for for...of or spread. Much later, when we get into how the engine stores strings under the hood, this will all snap into place.

Converting between numbers and strings

You will constantly move values back and forth between text and numbers -- reading "42" from a file and needing the number 42, or having a number and needing to build a message from it. Going from number to string is easy: String() or a template literal:

const n = 42;
console.log(String(n)); // "42"
console.log(`${n}`);    // "42" - template literal converts automatically
console.log(n.toString()); // "42" - the method form
console.log((255).toString(16)); // "ff" - toString can even change the base!

Going the other way, from string to number, you have Number(), parseInt, and parseFloat, and they differ in an important way:

console.log(Number("3.14"));       // 3.14
console.log(Number("nope"));       // NaN - any garbage yields NaN
console.log(parseInt("42px", 10)); // 42  - reads leading digits, ignores the "px"
console.log(parseFloat("3.5kg"));  // 3.5 - same idea, but keeps the decimal

The distinction: Number is strict -- if the whole string is not a clean number, you get NaN (that "not a number" value we met last episode, which is never equal to itself). parseInt and parseFloat are lenient -- they read as many valid leading digits as they can and cheerfully ignore trailing junk like "px" or "kg". Which behaviour you want depends on the situation: strict Number is safer when you expect clean input and want to catch mistakes, while parseInt is handy when you are pulling a number off the front of something like a CSS value.

One firm habit: always pass the base 10 to parseInt. That second argument (the "radix") tells it to read the digits as ordinary base-10, and passing it explicitly sidesteps a legacy footgun where some strings were once interpreted as octal. It costs you three characters and saves you a genuinely nasty class of bug.

How this compares to Python, Rust, Go and C

Many of you arrived from the Learn Python Series, so a glance sideways sharpens the picture -- every language draws the "what is a string, really" line a little differently, and JavaScript's choices look less arbitrary once you see the alternatives.

Python agrees with JavaScript on the big things: its strings are immutable too, and f-strings (f"{user} has {n} episodes") are the direct cousin of template literals. But Python quietly wins on Unicode. A Python 3 string is a sequence of code points, not UTF-16 units, so len() counts what a human would call characters, and the surrogate-pair surprise simply does not exist:

# Python: len counts code points, so an emoji is 1, not 2
print(len("cafe"))     # 4
print(len("a\U0001F680b"))  # 3  - the rocket counts as ONE
name = f"{user} rules" # f-string, the same idea as a template literal

Rust goes the strict, no-surprises route. A Rust String is UTF-8 bytes, and .len() returns the byte count, not the character count. Crucially, you cannot index a string with an integer at all -- it will not even compile, precisely because "the 3rd byte" is rarely "the 3rd character". Rust forces you to say what you mean with .chars() or .bytes():

// Rust: len() is BYTES, and integer indexing is a compile error
let s = String::from("cafe");
println!("{}", s.len());        // 4 here, but would be more for multibyte text
// let c = s[0];                // COMPILE ERROR: strings are not integer-indexable
let first = s.chars().next();   // the explicit, correct way to get a character

Go is similar in spirit to Rust: strings are UTF-8 byte sequences, len() gives bytes, and indexing gives you a single byte. But ranging over a string with for range decodes proper code points (Go calls them "runes"), so the language nudges you toward the correct iteration without forbidding the byte view. C sits at the far opposite end from Python: a C string is just a null-terminated array of bytes, strlen counts bytes up to the terminator, and the language has no built-in notion of Unicode whatsoever -- correctness there is entirely on you.

So where does JavaScript land? Right in an awkward historical middle. It chose UTF-16 code units back in the 1990s, when Unicode still fit in 16 bits and that seemed like plenty. When Unicode outgrew 16 bits, surrogate pairs were bolted on, and JavaScript's .length inherited the quirk forever. That is the whole reason .length can "lie": it is faithfully reporting code units, we just tend to think in characters. Knowing that this is a frozen historical accident (not some deep design intent) is exactly what makes for...of and spread feel like the sensible tools they are, rather than mysterious workarounds.

Putting it together

Let us close with a small program that leans on several of today's tools at once, so you can watch them cooperate on something slightly more real than a one-liner. We will take a raw "log line", clean it up, pull it apart, and build a friendly summary:

function summarizeLine(raw) {
  const clean = raw.trim();                 // strip stray whitespace
  const parts = clean.split(",");           // "user,action,count" -> three fields
  const user = parts[0].toUpperCase();      // normalise the name loudly
  const action = parts[1];
  const count = Number(parts[2]);           // strict parse of the count field
  const plural = count === 1 ? "time" : "times"; // ternary from last episode
  return `${user} did "${action}" ${count} ${plural}.`;
}

console.log(summarizeLine("  scipio,publish,5 ")); // 'SCIPIO did "publish" 5 times.'
console.log(summarizeLine("ada,login,1"));         // 'ADA did "login" 1 time.'

Trace the first call. The input has leading and trailing spaces, so trim cleans it. split(",") breaks it into ["scipio", "publish", "5"]. We uppercase the first field, keep the second as is, and use Number to turn "5" into a real number so our count === 1 comparison actually compares numbers (had we left it as the string "5", "5" === 1 would be false and we would always get the plural). Finally a template literal stitches the whole summary together, singular-versus-plural handled by the ternary. Nothing here is exotic -- it is just today's pieces, chosen deliberately and snapped together. That deliberate choosing is the entire craft.

Try it yourself

  1. Take the string " Scipio_NL ". In one chain of methods, trim the whitespace, lowercase it, and replace the underscore with a space, producing "scipio nl". Print it, and explain in a comment why you had to capture the final return value in stead of expecting the original string to change.
  2. Given const line = "2026-08-18", split it on "-" and print the year, month, and day on three separate lines, each built with its own template literal (for example "Year: 2026").
  3. Write a function initials(fullName) that takes "scipio the great" and returns "S.T.G.". Hint: split on spaces, take the first character of each part, uppercase it, and join the pieces with dots (do not forget the trailing dot). Then test it with your own name.

So what did we actually cover?

  • Template literals (backticks) give you ${...} interpolation of any expression and painless multi-line strings; prefer them over + concatenation almost always.
  • Strings are immutable: every "modifying" method returns a NEW string and leaves the original alone, so you must always capture the return value.
  • The everyday methods worth knowing cold: length/at, trim, toUpperCase/toLowerCase, includes/startsWith/endsWith/indexOf, slice, replace/replaceAll, and the split/join pairing.
  • JavaScript strings are UTF-16 code units, so .length can miscount emoji and some scripts (surrogate pairs count as 2); iterate real characters with for...of or the spread operator.
  • Convert number-to-string with String()/template literals/toString, and string-to-number with Number (strict) or parseInt/parseFloat (lenient) -- and always pass base 10 to parseInt.
  • Python counts code points, Rust and Go count UTF-8 bytes (but iterate code points cleanly), and C punts on Unicode entirely; JavaScript's UTF-16 model is a frozen 1990s accident, which is why .length surprises you.

Next episode we confront the single most infamous quirk in the whole language head-on: numbers, the IEEE 754 floating-point format they are built on, exactly why 0.1 + 0.2 does not equal 0.3, and how to write money and math code that behaves anyway ;-)

Thanks for reading, and see you in the next one.

scipio@scipio

Learn JS Series (#5) - Strings: Template Literals, Unicode, and the... | Ecency