← Back to Core Python
Lesson 6 · Python Foundations

String Methods

split, join, replace, find, strip, and the family around them.

Beginner40 min

What you will be able to do

  • Use the case, whitespace, replace, split/join, search, and validation families
  • Explain why strip("abc") does not remove the substring "abc"
  • Choose between find() and index(), and between split() and partition()
  • Apply the split-process-join pattern to real text
  • Chain methods, and know when a chain has got too long
  • Normalise user input the way a backend actually needs it

The idea, in plain English

A method is a function attached to an object. Every string carries around forty of them, and you will use perhaps a dozen regularly. They divide neatly into six families: change the case, trim whitespace, replace text, break apart and put back together, search, and validate.

The rule from the last lesson governs all of them: strings are immutable, so no method ever changes a string. Each one builds a new string and returns it. If you do not assign the result, you have computed something and thrown it away.

That immutability is also what makes chaining work. text.strip().lower().replace(" ", "-") is four strings, each built from the last, and it reads left to right in the order things happen.

The methods themselves are easy. What is worth your attention is the handful that do something slightly different from what their name suggests - strip, split, find - because those are where the bugs come from.

Worked example: Turning " Python, BACKEND , FastAPI " into clean lowercase tags.

strip() takes characters, not a substring

strip() with no argument removes whitespace from both ends, which is what most people want and the reason it exists. Give it an argument and behaviour changes in a way the name does not hint at.

"abcPythonabc".strip("abc") gives "Python" - which looks like it removed the substring. It did not. It removed any of the characters a, b, or c from both ends, repeatedly, until it hit something else. "cbaPythonbac".strip("abc") also gives "Python", and "abcPython".strip("abcP") gives "ython".

When you genuinely want to remove a known prefix or suffix, Python 3.9 added the methods that do exactly that and nothing else: removeprefix() and removesuffix().

Trimming, precisely
strip()Removes whitespace from both ends. The everyday case.
lstrip() / rstrip()The same, but only the left or only the right end.
strip("#")Removes every # from both ends. "###Python###" becomes "Python".
strip("abc")Removes any of a, b, c from both ends - as a set of characters, not as a word.
removeprefix("http://")Removes that exact prefix once, if present. Leaves the string alone otherwise.
removesuffix(".csv")Removes that exact suffix once. The right tool for file extensions.
replace(" ", "")Removes from everywhere, not just the ends. Different job entirely.

Watch out: The classic bug: "report.csv".strip(".csv") gives "report" by luck, but "scv.csv".strip(".csv") gives an empty string. Use removesuffix(".csv").

split() with no argument is not split(" ")

These look interchangeable and are not. split() with no argument splits on any run of whitespace and discards empty pieces, so "a b".split() gives ["a", "b"]. split(" ") splits on each single space individually, so "a b".split(" ") gives ["a", "", "", "b"].

Given real user input - which has double spaces, tabs, and a trailing newline in it - the bare split() is almost always what you want. It is also why the text-cleaning pattern below never needs a separate step to collapse repeated spaces.

When the separator is meaningful, like a comma in a CSV field, pass it explicitly. There you do want the empty pieces, because an empty field is data.

Breaking a string apart
split()Splits on any run of whitespace, drops empties. The right default for prose.
split(",")Splits on each comma, keeps empties. "a,,b" gives ["a", "", "b"].
split(",", 2)At most two splits, so at most three pieces. The rest stays intact.
rsplit("/", 1)Splits from the right. Perfect for "everything, then the last part".
partition("@")Always three pieces: before, the separator, after. Even when not found.
rpartition("/")The same, searching from the right.
splitlines()Splits on line endings, handling \n and \r\n both.

split, process, join

This three-step shape solves most text problems you will meet, and once you recognise it you will use it constantly: break the text into pieces, do something to each piece, put them back together.

join() is called on the separator rather than on the list, which surprises everyone once: " ".join(words), not words.join(" "). The reason is that join has to work on any iterable of strings - a list, a tuple, a generator - and putting the method on str keeps one implementation instead of one per container type.

The pieces must all be strings. ",".join([1, 2, 3]) raises TypeError; convert first with a generator expression.

Tip: Building a string in a loop with += copies everything each time. Collect the pieces and join once - it is both faster and easier to read.

find() or index()?

They do the same search and differ only in how they report failure. find() returns -1; index() raises ValueError. Neither is better - they suit different situations.

Use find() when not finding it is an ordinary outcome you plan to handle. Use index() when its absence means something has already gone wrong and you want the program to stop rather than continue with -1 quietly flowing into a slice.

And if you only want to know whether it is there at all, use neither - `in` is clearer than comparing a position against -1.

Searching
"x" in sTrue or False. Use this whenever you only need the yes/no.
s.find("x")Index of the first occurrence, or -1. Never raises.
s.rfind("x")Index of the last occurrence, or -1.
s.index("x")Index of the first occurrence, or ValueError. Use when absence is a bug.
s.count("x")How many non-overlapping occurrences. 0 when absent, so no special case needed.
s.startswith(("a", "b"))Accepts a tuple, so several prefixes can be checked at once.
s.endswith((".jpg", ".png"))The same for suffixes - the idiomatic file-type check.

Syntax and examples

Changing case
text = "python PROGRAMMING language" print(text.lower()) # python programming language print(text.upper()) # PYTHON PROGRAMMING LANGUAGE print(text.capitalize()) # Python programming language first char only print(text.title()) # Python Programming Language every word print(text.swapcase()) # PYTHON programming LANGUAGE # casefold() is lower() built for comparing, not displaying print("straße".lower()) # straße print("straße".casefold()) # strasse <- matches "STRASSE"
Trimming, and the trap
print(" Chandu ".strip()) # "Chandu" print("###Python###".strip("#")) # "Python" # strip() takes a SET OF CHARACTERS, not a substring print("abcPythonabc".strip("abc")) # "Python" looks right... print("scv.csv".strip(".csv")) # "" ...but is not # Python 3.9+: say what you actually mean print("report.csv".removesuffix(".csv")) # "report" print("scv.csv".removesuffix(".csv")) # "scv" print("https://x.com".removeprefix("https://")) # "x.com"
replace, with and without a limit
text = "Python is easy. Python is powerful." print(text.replace("Python", "Go")) # replaces both print(text.replace("Python", "Go", 1)) # replaces only the first # Chaining replacements to clean a phone number phone = "+91-9876 543-210" print(phone.replace("-", "").replace(" ", "")) # +919876543210
split() vs split(" ")
messy = "Python is easy" print(messy.split()) # ['Python', 'is', 'easy'] print(messy.split(" ")) # ['Python', '', '', 'is', '', '', '', 'easy'] # With a real separator, the empties are data print("a,,b".split(",")) # ['a', '', 'b'] # maxsplit keeps the remainder in one piece print("Python is a language".split(" ", 2)) # ['Python', 'is', 'a language'] # partition always gives three parts print("chandu@example.com".partition("@")) # ('chandu', '@', 'example.com') print("no-at-sign".partition("@")) # ('no-at-sign', '', '')
join, and the split-process-join pattern
words = ["Python", "is", "powerful"] print(" ".join(words)) # Python is powerful print(", ".join(words)) # Python, is, powerful print("\n".join(words)) # one per line # join is called on the SEPARATOR, and needs strings numbers = [10, 20, 30] print(",".join(str(n) for n in numbers)) # 10,20,30 # The pattern, on messy input text = " PYTHON IS POWERFUL " print(" ".join(text.strip().lower().split())) # python is powerful
Searching and validating
text = "Python programming" print("Python" in text) # True - use this for yes/no print(text.find("Java")) # -1 - never raises print(text.count("m")) # 2 print("photo.JPEG".lower().endswith((".jpg", ".jpeg", ".png"))) # True username = "chandu123" print(username.isalnum()) # True letters and digits only print("chandu_123".isalnum()) # False underscore is neither print("12345".isdigit()) # True print(" ".isspace()) # True "looks blank" check
What this looks like in a backend
# Normalising what a user typed, at the boundary email = " CHANDU@Example.COM " email = email.strip().casefold() # chandu@example.com # Tags arriving as one string from a form tags = " Python, BACKEND , FastAPI, PostgreSQL " tags = [tag.strip().lower() for tag in tags.split(",")] print(tags) # ['python', 'backend', 'fastapi', 'postgresql'] # A URL slug from a title title = " Learn Python String Methods " print("-".join(title.strip().lower().split())) # learn-python-string-methods # Pulling the filename off a path path = "reports/2026/september/report.csv" print(path.rpartition("/")[2]) # report.csv

Tip: A chain of three or four methods reads well. Beyond that, give the intermediate results names - a debugger cannot stop in the middle of a chain, and neither can a reader.

The methods worth knowing, by family

Around forty methods exist. These are the ones that earn their place in everyday code.

Case

lower, upper, capitalize, title, swapcase, casefold. Use casefold for comparing, lower for displaying.

Whitespace

strip, lstrip, rstrip. With no argument they trim whitespace; with one they trim a set of characters.

Prefix / suffix

removeprefix, removesuffix. Python 3.9+, and the correct answer whenever you were tempted to use strip for this.

Replace

replace(old, new) changes every occurrence; a third argument limits how many.

Break apart

split, rsplit, splitlines, partition, rpartition.

Put together

join, called on the separator and given an iterable of strings.

Search

in, find, rfind, index, rindex, count, startswith, endswith.

Validate

isalpha, isdigit, isalnum, isspace, islower, isupper, istitle - see below.

The is* validation methods

All of them return False for an empty string, which is usually what you want and occasionally a surprise.

isalpha()

Letters only. "Python" True, "Python123" False, "" False.

isdigit()

Digits only. Accepts superscripts like "²"; isdecimal() is stricter and isnumeric() looser.

isalnum()

Letters or digits, nothing else. An underscore or a space makes it False.

isspace()

Whitespace only - the test for input that looks blank but is not empty.

islower() / isupper()

True when every cased character is that case. "python123".islower() is True.

istitle()

True when every word starts uppercase and the rest is lower.

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

The split difference

“print("a b".split(), "a b".split(" "))”

The strip trap

“print("scv.csv".strip(".csv"), "scv.csv".removesuffix(".csv"))”

partition always gives 3

“print("no-at-sign".partition("@"))”

Clean in one line

“print(" ".join(" A B ".strip().lower().split()))”

What usually goes wrong

Not storing the result

Still the most common Python mistake, and it produces no error. Every string method returns a new string; the original is never touched.

✗ name.strip()
name.upper()
✓ name = name.strip().upper()
Using strip() to remove a substring

strip() takes a set of characters and removes any of them from both ends, repeatedly. "scv.csv".strip(".csv") returns an empty string. removesuffix() is the method that does what you meant.

✗ name = filename.strip(".csv")
✓ name = filename.removesuffix(".csv")
Writing split(" ") for prose

Real text has double spaces and tabs in it. split(" ") produces empty strings for each extra space; bare split() collapses any run of whitespace.

✗ words = text.split(" ")
✓ words = text.split()
Calling join on the list

join is a method of the separator, not of the sequence. It reads oddly at first and is worth saying aloud once: "join these words with a space".

✗ words.join(" ")
✓ " ".join(words)
Joining things that are not strings

join needs an iterable of strings and will not convert for you. Convert inside the call.

✗ ",".join([1, 2, 3])        # TypeError
✓ ",".join(str(n) for n in [1, 2, 3])
Expecting find() to raise, or index() to return -1

They are mirror images. find() returns -1 and never raises; index() raises ValueError and never returns -1. Using the wrong one means either an unnoticed -1 flowing into a slice, or an unexpected crash.

✗ pos = text.index("x")     # crashes when absent
✓ pos = text.find("x")
if pos != -1:
    ...
Forgetting the parentheses

text.upper without brackets is the method object itself, not a call. Printing it shows something like <built-in method upper> rather than your text.

✗ print(text.upper)
✓ print(text.upper())

Best practices

  • Normalise at the boundary: text = text.strip().casefold(), once, where the data arrives.
  • Use casefold() for comparing and lower() for displaying.
  • Use removeprefix() and removesuffix() for known prefixes and suffixes; keep strip() for whitespace.
  • Prefer `in` over find() != -1 when you only need to know whether something is present.
  • Pass a tuple to startswith() and endswith() rather than writing several or-ed checks.
  • Use bare split() for prose and an explicit separator for structured data.
  • Build strings with join(), not with += in a loop.
  • Reach for pathlib rather than string methods once you are genuinely manipulating filesystem paths.

Practice

Write these yourself before opening anything. Getting them wrong first is most of how this sticks.

1.

Turn " CHANDU PRASAD " into "Chandu Prasad".

Show hint

Trim first, then fix the case. title() capitalises each word.

Show solution
name = " CHANDU PRASAD " name = name.strip().title() print(name) # Chandu Prasad
2.

Normalise " USER@EXAMPLE.COM " to "user@example.com".

Show solution
email = " USER@EXAMPLE.COM " email = email.strip().casefold() print(email) # user@example.com
3.

Count how many times "a" appears in "banana", and how many times "an" does.

Show hint

count() works on substrings as well as single characters.

Show solution
text = "banana" print(text.count("a")) # 3 print(text.count("an")) # 2 - non-overlapping matches
4.

Turn " Python, React , Node.js , PostgreSQL " into a clean lowercase list.

Show solution
skills = " Python, React , Node.js , PostgreSQL " skills = [skill.strip().lower() for skill in skills.split(",")] print(skills) # ['python', 'react', 'node.js', 'postgresql']
5.

Check whether "employee_data.CSV" is a CSV file, in a way that works for any capitalisation.

Show solution
filename = "employee_data.CSV" if filename.lower().endswith(".csv"): print("CSV file") # CSV file
6.

Split "student@example.com" into the part before and the part after the @, using partition().

Show solution
email = "student@example.com" user, _, domain = email.partition("@") print(user) # student print(domain) # example.com
7.

Generate the slug "learn-python-string-methods" from " Learn Python String Methods ".

Show hint

strip, lower, split, then join with a hyphen.

Show solution
title = " Learn Python String Methods " slug = "-".join(title.strip().lower().split()) print(slug) # learn-python-string-methods
Coding challenge

Text cleaner

Write a cleaner that takes whatever a user typed and returns something you would be willing to store. Then run it over the awkward inputs below and check every one.

It should
  • Trim the ends and lowercase the text
  • Collapse runs of spaces and tabs into single spaces
  • Return an empty string for input that is only whitespace
  • Do it without a regular expression - the methods in this lesson are enough
  • Run it over the test inputs and print each result in quotes so the spaces are visible
Start here
def clean(text): ... for raw in [" PYTHON IS POWERFUL ", "Hello\tWorld", " ", "already clean", ""]: print(repr(raw), "->", repr(clean(raw)))
Show one solution
One solution
def clean(text): # split() with no argument splits on any run of whitespace and drops the # empties, so the collapsing happens for free - no regex needed. return " ".join(text.lower().split()) for raw in [" PYTHON IS POWERFUL ", "Hello\tWorld", " ", "already clean", ""]: print(f"{raw!r:34} -> {clean(raw)!r}") # ' PYTHON IS POWERFUL ' -> 'python is powerful' # 'Hello\tWorld' -> 'hello world' # ' ' -> '' # 'already clean' -> 'already clean' # '' -> ''

Key points

  • A string method returns a new string - it never modifies the original.
  • Method chaining runs left to right, each call working on the previous result.
  • strip() with an argument removes a set of characters from the ends, not a substring.
  • removeprefix() and removesuffix() remove an exact prefix or suffix, which is usually what you wanted.
  • split() with no argument splits on any whitespace and drops empties; split(" ") does not.
  • join() is called on the separator and needs every item to already be a string.
  • split, process, join is the shape of most text transformations.
  • find() returns -1 when absent; index() raises ValueError. Use `in` when you only need yes or no.
  • startswith() and endswith() accept a tuple, so several options can be checked at once.
  • The is* methods all return False for an empty string.
  • casefold() is for comparing, lower() is for displaying.
  • The normalisation one-liner worth remembering: text = text.strip().casefold().

Quick check before you move on

What does "Python,React,Node".split() return?
['Python,React,Node'] - a single-item list. With no argument it splits on whitespace, and there is none. You need split(",").
What does " Python ".strip().upper() print?
"PYTHON". strip() runs first and removes the spaces, then upper() runs on its result.
For text = "Python Python", what do find("Python") and rfind("Python") return?
0 and 7. find() reports the first occurrence, rfind() the last.
Why does ",".join([1, 2, 3]) fail?
join() requires every item to be a string and will not convert for you. Use ",".join(str(n) for n in [1, 2, 3]).

Interview questions

Why is join() a method of str rather than of list?

Because it accepts any iterable of strings, not just lists. Putting it on str means one implementation covers lists, tuples, sets, and generators; putting it on the container would mean reimplementing it for each. It also makes the separator explicit at the call site, which is the part that varies.

Someone writes filename.strip(".csv") to remove an extension. What goes wrong?

strip() takes a set of characters, not a substring, and strips any of them from both ends until it hits a character outside the set. "scv.csv" becomes empty, and "csv.report.csv" loses characters from the front too. removesuffix(".csv") removes that exact suffix once, and leaves the string untouched if it is not there.

When would you use casefold() rather than lower()?

For case-insensitive comparison, particularly with non-English text. casefold applies the full Unicode case-folding rules, so German ß folds to ss and "straße" matches "STRASSE"; lower() leaves ß alone and the comparison fails. lower() is for what a person will read; casefold is for what a program will compare.

How would you count word frequency in a paragraph?

Normalise with strip and casefold, split with no argument so runs of whitespace collapse, strip punctuation from each word, then count - collections.Counter does the counting in one line. The subtlety is punctuation: "python." and "python" are different keys until you deal with it.

Why is building a string with += in a loop a problem?

Strings are immutable, so each += allocates a new string and copies everything accumulated so far, making the loop quadratic in the total length. Appending to a list and calling join() once allocates the result a single time. CPython has an optimisation that sometimes hides this, but it is not something to depend on.

What is the difference between split() and partition()?

split() returns a list of however many pieces the separator produces, and drops the separator. partition() always returns exactly three values - before, the separator, after - even when the separator is absent, in which case the last two are empty. Because the shape is fixed, partition unpacks safely into three names without a length check.

Quiz

  1. 1.

    Does upper() change the original string?

  2. 2.

    What is the difference between find() and index()?

  3. 3.

    What does "scv.csv".strip(".csv") return, and why?

  4. 4.

    What is the difference between split() and split(" ")?

  5. 5.

    Why is join() called on the separator rather than the list?

  6. 6.

    How do you check whether a filename is an image, allowing several extensions?

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...