← Back to Core Python
Lesson 5 · Python Foundations

Strings

Quotes, indexing, slicing, escapes, raw strings, and immutability.

Beginner35 min

What you will be able to do

  • Create strings with single, double, and triple quotes, and know when each is right
  • Access any character by positive or negative index
  • Extract a portion of a string with slicing, including a step and a reverse
  • Explain what immutability means for a string, and what it forces you to do
  • Use escape sequences, and reach for a raw string when the backslashes pile up
  • Join, repeat, measure, and search strings with +, *, len(), and in

The idea, in plain English

A string is a sequence of characters - letters, digits, spaces, symbols, anything. That one sentence is worth holding onto, because "sequence" is what makes indexing, slicing, len(), and `in` all work the way they do. The same ideas will come back unchanged when you meet lists.

Every character has a position, and Python counts positions from zero. The first character of "Python" is at index 0 and the last is at index 5, which means the last index of any string is len(s) - 1. Nearly every off-by-one error you will write starts by forgetting that.

Strings are immutable: once a string object exists, its characters can never change. You cannot assign to s[0], and every method that looks like it edits a string actually builds a new one and hands it back. If you do not store what comes back, nothing has happened.

This lesson covers what a string is and how to take it apart. The next one covers the methods for working with text - splitting, joining, replacing, searching - in depth.

Worked example: Pulling "Python" out of "I am learning Python" with one slice.

Indexing, forwards and backwards

Every character sits at two indexes at once: one counting forward from zero, one counting back from the end. For "Python", P is both 0 and -6, and n is both 5 and -1.

Negative indexes exist so you do not have to compute len(s) - 1 every time you want the end of something. s[-1] is always the last character, whatever the length.

The two index systems for "Python"
PForward index 0, backward index -6. s[0] and s[-6] are the same character.
yForward 1, backward -5.
tForward 2, backward -4.
hForward 3, backward -3.
oForward 4, backward -2.
nForward 5, backward -1. The last character is always s[-1].
s[6]IndexError. Valid indexes are 0 to 5 and -6 to -1; there is no index 6.

Tip: len(s) is 6 but the last index is 5. Reading "length six, last index five" out loud a few times saves a lot of IndexErrors later.

Slicing: start, stop, step

s[start:stop] takes everything from start up to but not including stop. That exclusive end feels arbitrary for about a day, then starts paying off: s[:3] and s[3:] split a string cleanly with no overlap and no gap, and the length of s[a:b] is simply b - a.

Every part is optional. Leave out the start and it begins at the beginning; leave out the stop and it runs to the end; add a third number and it becomes a step. A negative step walks backwards, which is where the famous reverse comes from.

Slicing never raises IndexError. Ask for s[0:100] on a six-character string and you get the whole string rather than an error - which is convenient, and occasionally hides a bug.

Slicing "Python"
[0:3]"Pyt" - from 0, stopping before 3. Three characters, because 3 - 0 is 3.
[:3]"Pyt" - the start defaults to the beginning.
[2:]"thon" - the stop defaults to the end.
[-3:]"hon" - the last three characters, whatever the length.
[:-2]"Pyth" - everything except the last two.
[::2]"Pto" - every second character, from the start.
[::-1]"nohtyP" - a negative step walks backwards, so this reverses the string.
[0:100]"Python" - out-of-range slices are clipped rather than raising.

Immutable means you get a new string back

A string can never be edited in place. s[0] = "J" raises TypeError, and there is no method anywhere that changes a string - they all return a new one.

That produces the single most common beginner bug in Python: calling a method and throwing the result away. name.upper() on its own computes an uppercase string and immediately discards it, leaving name exactly as it was. You have to write name = name.upper().

It also means you build a changed string rather than editing one: "J" + name[1:] turns "Python" into "Jython" by making a third string out of two pieces.

Watch out: If a string method seems to have done nothing, you almost certainly forgot to assign its result. Check that line before checking anything else.

Syntax and examples

Three ways to quote, and why all three exist
single = 'Python' double = "Python" # identical; pick one and be consistent # The quote you did not use is free to appear inside message = "I'm learning Python" quoted = 'He said "Hello".' # Triple quotes span lines, keeping the line breaks email = """Hello John, Your registration was successful. Thank you"""
Indexing
language = "Python" print(language[0]) # P first print(language[5]) # n last, the hard way print(language[-1]) # n last, the useful way print(language[-2]) # o print(len(language)) # 6 length is 6, last index is 5 print(language[6]) # IndexError: string index out of range
Slicing
word = "Python" print(word[0:3]) # Pyt start included, stop excluded print(word[:3]) # Pyt print(word[2:]) # thon print(word[-3:]) # hon print(word[:-2]) # Pyth print(word[::2]) # Pto every second character print(word[::-1]) # nohtyP reversed # A slice splits with no overlap and no gap print(word[:3] + word[3:] == word) # True
Immutability, and the mistake it causes
name = "Python" # name[0] = "J" # TypeError: 'str' object does not support item assignment name = "J" + name[1:] # build a new string instead print(name) # Jython # The classic: greeting = "chandu" greeting.upper() # computes "CHANDU" and throws it away print(greeting) # chandu <- unchanged greeting = greeting.upper() print(greeting) # CHANDU
Joining, repeating, measuring, searching
first, last = "Chandu", "Prasad" print(first + " " + last) # Chandu Prasad print("-" * 20) # -------------------- print(len("Hello World")) # 11 the space counts message = "Welcome to Python" print("Python" in message) # True print("Java" not in message) # True # Numbers need converting first age = 30 # print("Age: " + age) # TypeError print("Age: " + str(age)) # Age: 30
Escapes and raw strings
print("Name:\tChandu") # a tab between them print("Line one\nLine two") # two lines print("He said \"Hello\".") # escaped double quotes # Backslashes must themselves be escaped... path = "C:\\Users\\Chandu" print(path) # C:\Users\Chandu # ...unless the string is raw path = r"C:\Users\Chandu\Documents" print(path) # C:\Users\Chandu\Documents pattern = r"\d+" # regex patterns are almost always raw

Tip: A string in single quotes and the same string in double quotes are the same object to Python - the quotes are only how you typed it. Pick one style for a project; most codebases use double.

Escape sequences

A backslash inside a string starts an escape sequence. These are the ones you will actually meet.

\n

New line. The most used of all of them.

"Line one\nLine two"
\t

Tab. Handy for lining up console output.

"Name:\tChandu"
\\

A literal backslash - you need two to write one.

"C:\\Users"
\'

A single quote inside a single-quoted string.

'I\'m here'
\"

A double quote inside a double-quoted string.

"He said \"Hi\""
r"..."

A raw string: backslashes lose their special meaning. For Windows paths and regex.

r"C:\Users"

The patterns worth memorising

Six shapes that cover most everyday string work. The methods themselves get a lesson of their own next.

Last character

Always s[-1], regardless of length.

s[-1]
Reverse

A slice with a step of -1.

s[::-1]
Split in two

No overlap, no gap, because the stop is exclusive.

s[:n], s[n:]
Drop the last n

A negative stop counts back from the end.

s[:-n]
Contains?

The `in` operator, not a method.

"Py" in s
Length

A built-in function, not a method - len(s), never s.len().

len(s)

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Both index systems

“w = "Python"; print(w[0], w[-6], w[5], w[-1])”

Slices reassemble

“w = "Python"; print(w[:3] + w[3:] == w)”

Out of range is fine

“print("Python"[0:100], "Python"[100:])”

A palindrome check

“s = "racecar"; print(s == s[::-1])”

What usually goes wrong

Counting from one

The first character is at index 0, so the last is at len(s) - 1. Reaching for s[len(s)] is the classic off-by-one, and it raises rather than returning an empty string.

✗ first = word[1]        # that is the second character
✓ first = word[0]
Throwing away the result of a method

Strings are immutable, so every method returns a new string and leaves the original alone. If you do not assign the result, nothing changes and no error tells you.

✗ name.strip()
name.upper()
✓ name = name.strip().upper()
Trying to edit a character

Item assignment is not supported on strings. Build a new string from slices instead.

✗ text[0] = "J"      # TypeError
✓ text = "J" + text[1:]
Expecting the slice end to be included

s[0:3] gives three characters, not four. The stop index is where slicing stops, not the last one it takes.

✗ first_three = word[0:2]     # only two
✓ first_three = word[0:3]
Adding a number to a string

Python will not guess whether you meant to join text or to add. Convert explicitly, or use an f-string, which converts for you.

✗ "Age: " + age        # TypeError
✓ f"Age: {age}"
Forgetting that comparison is case sensitive

"Python" == "python" is False. Normalise both sides before comparing anything a user typed.

✗ if username == "chandu":
✓ if username.lower() == "chandu":
Expecting strip() to remove characters from the middle

strip() only trims from the two ends. To remove every occurrence of something, replace() is the tool.

✗ "9876-543-210".strip("-")     # unchanged
✓ "9876-543-210".replace("-", "")

Best practices

  • Use s[-1] rather than s[len(s) - 1]. It is shorter and cannot go wrong.
  • Normalise user input early: text = text.strip().lower(), once, at the boundary.
  • Use raw strings for Windows paths and regular expressions, so backslashes stay as typed.
  • Use `in` for a simple contains check rather than find() != -1 - it says what you mean.
  • Prefer f-strings over + for building messages, especially with numbers in them.
  • Remember that a raw string still cannot end in a single backslash - r"C:\" is a syntax error.

Practice

Write these yourself before opening anything. Getting them wrong first is most of how this sticks.

1.

Given language = "Python", print the first character and the last character, using a negative index for the last.

Show solution
language = "Python" print(language[0]) # P print(language[-1]) # n
2.

Print the length of "Welcome to Python". Count on paper first - remember the spaces.

Show solution
message = "Welcome to Python" print(len(message)) # 17
3.

From message = "I am learning Python", extract just the word "Python" with a slice.

Show hint

It is the last six characters, so a negative start is easiest.

Show solution
message = "I am learning Python" print(message[-6:]) # Python print(message[14:]) # Python - same result, counted forwards
4.

Reverse the word "Developer" using slicing.

Show solution
word = "Developer" print(word[::-1]) # repoleveD
5.

Turn " USER@EXAMPLE.COM " into "user@example.com".

Show hint

Two methods, chained. Remember to store the result.

Show solution
email = " USER@EXAMPLE.COM " email = email.strip().lower() print(email) # user@example.com print(f"[{email}]") # brackets prove the spaces are gone
6.

Check whether "Python" appears in "Learn Python programming from beginner to advanced", in a way that also works if it were written "python".

Show solution
description = "Learn Python programming from beginner to advanced" print("python" in description.lower()) # True
7.

Write a check that says whether a word is a palindrome - the same forwards and backwards. Test it with "racecar" and "python".

Show hint

You already know how to reverse a string.

Show solution
for word in ("racecar", "python"): if word == word[::-1]: print(word, "is a palindrome") else: print(word, "is not a palindrome")
Coding challenge

Username processor

Take a raw username as typed, clean it up, and decide whether it is acceptable. Then run it against the awkward cases at the bottom and see whether your rules hold.

It should
  • Trim leading and trailing spaces, and lowercase the result
  • Accept only letters and digits - no spaces, no symbols
  • Require a length between 3 and 20 characters after cleaning
  • Print the cleaned username and whether it is valid
  • Say which rule failed, rather than just "invalid"
Start here
username = " Chandu123 " # clean it, then check it # Try these too: # "ab" too short # "chandu@123" symbol # "this_is_a_very_long_username" underscore, and length # "PythonDeveloper" fine
Show one solution
One solution
for raw in [" Chandu123 ", "ab", "chandu@123", "this_is_a_very_long_username", "PythonDeveloper"]: username = raw.strip().lower() # Saying which rule failed makes the output useful to a user. if not username.isalnum(): reason = "must be letters and digits only" elif len(username) < 3: reason = "too short (minimum 3)" elif len(username) > 20: reason = "too long (maximum 20)" else: reason = None status = "valid" if reason is None else f"rejected - {reason}" print(f"{raw!r:34} -> {username!r:32} {status}") # ' Chandu123 ' -> 'chandu123' valid # 'ab' -> 'ab' rejected - too short (minimum 3) # 'chandu@123' -> 'chandu@123' rejected - must be letters and digits only # 'this_is_a_very_long_username' -> 'this_is_a_very_long_username' rejected - must be letters and digits only # 'PythonDeveloper' -> 'pythondeveloper' valid

Key points

  • A string is an immutable sequence of characters, of type str.
  • Single and double quotes are interchangeable; triple quotes span lines.
  • Indexing starts at 0, so the last index is len(s) - 1.
  • Negative indexes count from the end: s[-1] is always the last character.
  • s[start:stop] includes start and excludes stop, so len(s[a:b]) is b - a.
  • Every part of a slice is optional, and a third number is the step.
  • s[::-1] reverses a string, because a negative step walks backwards.
  • Slicing clips out-of-range values instead of raising; indexing raises IndexError.
  • Strings cannot be changed in place - s[0] = "J" is a TypeError.
  • Every string method returns a new string, so the result has to be assigned.
  • + joins, * repeats, len() measures, and `in` tests for a substring.
  • Raw strings turn off escape processing, which is what Windows paths and regexes want.

Quick check before you move on

For text = "Python", what do text[0] and text[-1] print?
P and n. Index 0 is the first character; -1 is always the last.
What does text[1:4] give for "Python"?
"yth". It starts at index 1 and stops before index 4, so three characters - indexes 1, 2 and 3.
What does text[::-1] do, and why?
It reverses the string. The start and stop are omitted so it covers everything, and a step of -1 walks backwards through it.
After name = "chandu" and name.upper(), what does print(name) show?
"chandu", unchanged. upper() built a new string and nothing stored it. You need name = name.upper().

Interview questions

Why are Python strings immutable?

It makes them safe to share - two names can point at one string with no risk of one changing it under the other, which also makes them safe as dictionary keys and across threads. It lets CPython cache and intern short strings, and it means a string’s hash can be computed once and reused. The cost is that building a string in a loop with += creates a new object each time, which is why join() exists.

Why does slicing not raise IndexError when indexing does?

Indexing asks for one specific character, so an index that does not exist has no sensible answer. A slice asks for a range, and an empty or clipped range is a perfectly meaningful answer - s[100:] is the empty string. It makes slicing safe to use without bounds checks, which is usually convenient and occasionally hides a mistake.

How would you reverse a string, and what are the alternatives?

s[::-1] is the idiomatic way and the fastest, because it is one C-level slice. "".join(reversed(s)) is more explicit and slightly slower. A manual loop with += is the worst of the three, since every step builds a whole new string.

What is the difference between lower() and casefold()?

Both normalise case, but casefold is more aggressive and designed specifically for caseless matching across Unicode - German "straße" casefolds to "strasse", where lower() leaves the ß alone. For comparing user input in any language, casefold is the safer choice.

Why is building a string with += in a loop discouraged?

Because strings are immutable, each += allocates a new string and copies everything so far, making the loop quadratic. Collect the pieces in a list and call "".join(parts) once, which allocates a single result. CPython optimises some simple cases, but the pattern is not something to rely on.

What does it mean that a string is a sequence?

It supports the sequence protocol - indexing, slicing, len(), iteration, `in`, and concatenation with +. That is why everything you learn here transfers directly to lists and tuples; the difference is that a list is mutable and a string is not.

Quiz

  1. 1.

    What is a string in Python?

  2. 2.

    What is the last valid index of a string of length 6?

  3. 3.

    Why does s[0:3] return three characters rather than four?

  4. 4.

    What happens with "Python"[0:100]?

  5. 5.

    Why does text[0] = "J" fail?

  6. 6.

    When would you use a raw string?

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...