String Methods
split, join, replace, find, strip, and the family around them.
What you will be able to do
- Use the case, whitespace, replace, split/join, search, and validation families
- Explain why strip("abc") does not remove the substring "abc"
- Choose between find() and index(), and between split() and partition()
- Apply the split-process-join pattern to real text
- Chain methods, and know when a chain has got too long
- Normalise user input the way a backend actually needs it
The idea, in plain English
A method is a function attached to an object. Every string carries around forty of them, and you will use perhaps a dozen regularly. They divide neatly into six families: change the case, trim whitespace, replace text, break apart and put back together, search, and validate.
The rule from the last lesson governs all of them: strings are immutable, so no method ever changes a string. Each one builds a new string and returns it. If you do not assign the result, you have computed something and thrown it away.
That immutability is also what makes chaining work. text.strip().lower().replace(" ", "-") is four strings, each built from the last, and it reads left to right in the order things happen.
The methods themselves are easy. What is worth your attention is the handful that do something slightly different from what their name suggests - strip, split, find - because those are where the bugs come from.
Worked example: Turning " Python, BACKEND , FastAPI " into clean lowercase tags.
strip() takes characters, not a substring
strip() with no argument removes whitespace from both ends, which is what most people want and the reason it exists. Give it an argument and behaviour changes in a way the name does not hint at.
"abcPythonabc".strip("abc") gives "Python" - which looks like it removed the substring. It did not. It removed any of the characters a, b, or c from both ends, repeatedly, until it hit something else. "cbaPythonbac".strip("abc") also gives "Python", and "abcPython".strip("abcP") gives "ython".
When you genuinely want to remove a known prefix or suffix, Python 3.9 added the methods that do exactly that and nothing else: removeprefix() and removesuffix().
strip()Removes whitespace from both ends. The everyday case.lstrip() / rstrip()The same, but only the left or only the right end.strip("#")Removes every # from both ends. "###Python###" becomes "Python".strip("abc")Removes any of a, b, c from both ends - as a set of characters, not as a word.removeprefix("http://")Removes that exact prefix once, if present. Leaves the string alone otherwise.removesuffix(".csv")Removes that exact suffix once. The right tool for file extensions.replace(" ", "")Removes from everywhere, not just the ends. Different job entirely.Watch out: The classic bug: "report.csv".strip(".csv") gives "report" by luck, but "scv.csv".strip(".csv") gives an empty string. Use removesuffix(".csv").
split() with no argument is not split(" ")
These look interchangeable and are not. split() with no argument splits on any run of whitespace and discards empty pieces, so "a b".split() gives ["a", "b"]. split(" ") splits on each single space individually, so "a b".split(" ") gives ["a", "", "", "b"].
Given real user input - which has double spaces, tabs, and a trailing newline in it - the bare split() is almost always what you want. It is also why the text-cleaning pattern below never needs a separate step to collapse repeated spaces.
When the separator is meaningful, like a comma in a CSV field, pass it explicitly. There you do want the empty pieces, because an empty field is data.
split()Splits on any run of whitespace, drops empties. The right default for prose.split(",")Splits on each comma, keeps empties. "a,,b" gives ["a", "", "b"].split(",", 2)At most two splits, so at most three pieces. The rest stays intact.rsplit("/", 1)Splits from the right. Perfect for "everything, then the last part".partition("@")Always three pieces: before, the separator, after. Even when not found.rpartition("/")The same, searching from the right.splitlines()Splits on line endings, handling \n and \r\n both.split, process, join
This three-step shape solves most text problems you will meet, and once you recognise it you will use it constantly: break the text into pieces, do something to each piece, put them back together.
join() is called on the separator rather than on the list, which surprises everyone once: " ".join(words), not words.join(" "). The reason is that join has to work on any iterable of strings - a list, a tuple, a generator - and putting the method on str keeps one implementation instead of one per container type.
The pieces must all be strings. ",".join([1, 2, 3]) raises TypeError; convert first with a generator expression.
Tip: Building a string in a loop with += copies everything each time. Collect the pieces and join once - it is both faster and easier to read.
find() or index()?
They do the same search and differ only in how they report failure. find() returns -1; index() raises ValueError. Neither is better - they suit different situations.
Use find() when not finding it is an ordinary outcome you plan to handle. Use index() when its absence means something has already gone wrong and you want the program to stop rather than continue with -1 quietly flowing into a slice.
And if you only want to know whether it is there at all, use neither - `in` is clearer than comparing a position against -1.
"x" in sTrue or False. Use this whenever you only need the yes/no.s.find("x")Index of the first occurrence, or -1. Never raises.s.rfind("x")Index of the last occurrence, or -1.s.index("x")Index of the first occurrence, or ValueError. Use when absence is a bug.s.count("x")How many non-overlapping occurrences. 0 when absent, so no special case needed.s.startswith(("a", "b"))Accepts a tuple, so several prefixes can be checked at once.s.endswith((".jpg", ".png"))The same for suffixes - the idiomatic file-type check.Syntax and examples
text = "python PROGRAMMING language"
print(text.lower()) # python programming language
print(text.upper()) # PYTHON PROGRAMMING LANGUAGE
print(text.capitalize()) # Python programming language first char only
print(text.title()) # Python Programming Language every word
print(text.swapcase()) # PYTHON programming LANGUAGE
# casefold() is lower() built for comparing, not displaying
print("straße".lower()) # straße
print("straße".casefold()) # strasse <- matches "STRASSE"print(" Chandu ".strip()) # "Chandu"
print("###Python###".strip("#")) # "Python"
# strip() takes a SET OF CHARACTERS, not a substring
print("abcPythonabc".strip("abc")) # "Python" looks right...
print("scv.csv".strip(".csv")) # "" ...but is not
# Python 3.9+: say what you actually mean
print("report.csv".removesuffix(".csv")) # "report"
print("scv.csv".removesuffix(".csv")) # "scv"
print("https://x.com".removeprefix("https://")) # "x.com"text = "Python is easy. Python is powerful."
print(text.replace("Python", "Go")) # replaces both
print(text.replace("Python", "Go", 1)) # replaces only the first
# Chaining replacements to clean a phone number
phone = "+91-9876 543-210"
print(phone.replace("-", "").replace(" ", "")) # +919876543210messy = "Python is easy"
print(messy.split()) # ['Python', 'is', 'easy']
print(messy.split(" ")) # ['Python', '', '', 'is', '', '', '', 'easy']
# With a real separator, the empties are data
print("a,,b".split(",")) # ['a', '', 'b']
# maxsplit keeps the remainder in one piece
print("Python is a language".split(" ", 2)) # ['Python', 'is', 'a language']
# partition always gives three parts
print("chandu@example.com".partition("@")) # ('chandu', '@', 'example.com')
print("no-at-sign".partition("@")) # ('no-at-sign', '', '')words = ["Python", "is", "powerful"]
print(" ".join(words)) # Python is powerful
print(", ".join(words)) # Python, is, powerful
print("\n".join(words)) # one per line
# join is called on the SEPARATOR, and needs strings
numbers = [10, 20, 30]
print(",".join(str(n) for n in numbers)) # 10,20,30
# The pattern, on messy input
text = " PYTHON IS POWERFUL "
print(" ".join(text.strip().lower().split())) # python is powerfultext = "Python programming"
print("Python" in text) # True - use this for yes/no
print(text.find("Java")) # -1 - never raises
print(text.count("m")) # 2
print("photo.JPEG".lower().endswith((".jpg", ".jpeg", ".png"))) # True
username = "chandu123"
print(username.isalnum()) # True letters and digits only
print("chandu_123".isalnum()) # False underscore is neither
print("12345".isdigit()) # True
print(" ".isspace()) # True "looks blank" check# Normalising what a user typed, at the boundary
email = " CHANDU@Example.COM "
email = email.strip().casefold() # chandu@example.com
# Tags arriving as one string from a form
tags = " Python, BACKEND , FastAPI, PostgreSQL "
tags = [tag.strip().lower() for tag in tags.split(",")]
print(tags) # ['python', 'backend', 'fastapi', 'postgresql']
# A URL slug from a title
title = " Learn Python String Methods "
print("-".join(title.strip().lower().split())) # learn-python-string-methods
# Pulling the filename off a path
path = "reports/2026/september/report.csv"
print(path.rpartition("/")[2]) # report.csvTip: A chain of three or four methods reads well. Beyond that, give the intermediate results names - a debugger cannot stop in the middle of a chain, and neither can a reader.
The methods worth knowing, by family
Around forty methods exist. These are the ones that earn their place in everyday code.
Caselower, upper, capitalize, title, swapcase, casefold. Use casefold for comparing, lower for displaying.
Whitespacestrip, lstrip, rstrip. With no argument they trim whitespace; with one they trim a set of characters.
Prefix / suffixremoveprefix, removesuffix. Python 3.9+, and the correct answer whenever you were tempted to use strip for this.
Replacereplace(old, new) changes every occurrence; a third argument limits how many.
Break apartsplit, rsplit, splitlines, partition, rpartition.
Put togetherjoin, called on the separator and given an iterable of strings.
Searchin, find, rfind, index, rindex, count, startswith, endswith.
Validateisalpha, isdigit, isalnum, isspace, islower, isupper, istitle - see below.
The is* validation methods
All of them return False for an empty string, which is usually what you want and occasionally a surprise.
isalpha()Letters only. "Python" True, "Python123" False, "" False.
isdigit()Digits only. Accepts superscripts like "²"; isdecimal() is stricter and isnumeric() looser.
isalnum()Letters or digits, nothing else. An underscore or a space makes it False.
isspace()Whitespace only - the test for input that looks blank but is not empty.
islower() / isupper()True when every cased character is that case. "python123".islower() is True.
istitle()True when every word starts uppercase and the rest is lower.
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“print("a b".split(), "a b".split(" "))”
“print("scv.csv".strip(".csv"), "scv.csv".removesuffix(".csv"))”
“print("no-at-sign".partition("@"))”
“print(" ".join(" A B ".strip().lower().split()))”
What usually goes wrong
Still the most common Python mistake, and it produces no error. Every string method returns a new string; the original is never touched.
✗ name.strip()
name.upper()✓ name = name.strip().upper()strip() takes a set of characters and removes any of them from both ends, repeatedly. "scv.csv".strip(".csv") returns an empty string. removesuffix() is the method that does what you meant.
✗ name = filename.strip(".csv")✓ name = filename.removesuffix(".csv")Real text has double spaces and tabs in it. split(" ") produces empty strings for each extra space; bare split() collapses any run of whitespace.
✗ words = text.split(" ")✓ words = text.split()join is a method of the separator, not of the sequence. It reads oddly at first and is worth saying aloud once: "join these words with a space".
✗ words.join(" ")✓ " ".join(words)join needs an iterable of strings and will not convert for you. Convert inside the call.
✗ ",".join([1, 2, 3]) # TypeError✓ ",".join(str(n) for n in [1, 2, 3])They are mirror images. find() returns -1 and never raises; index() raises ValueError and never returns -1. Using the wrong one means either an unnoticed -1 flowing into a slice, or an unexpected crash.
✗ pos = text.index("x") # crashes when absent✓ pos = text.find("x")
if pos != -1:
...text.upper without brackets is the method object itself, not a call. Printing it shows something like <built-in method upper> rather than your text.
✗ print(text.upper)✓ print(text.upper())Best practices
- Normalise at the boundary: text = text.strip().casefold(), once, where the data arrives.
- Use casefold() for comparing and lower() for displaying.
- Use removeprefix() and removesuffix() for known prefixes and suffixes; keep strip() for whitespace.
- Prefer `in` over find() != -1 when you only need to know whether something is present.
- Pass a tuple to startswith() and endswith() rather than writing several or-ed checks.
- Use bare split() for prose and an explicit separator for structured data.
- Build strings with join(), not with += in a loop.
- Reach for pathlib rather than string methods once you are genuinely manipulating filesystem paths.
Practice
Write these yourself before opening anything. Getting them wrong first is most of how this sticks.
Turn " CHANDU PRASAD " into "Chandu Prasad".
Show hintHide hint
Trim first, then fix the case. title() capitalises each word.
Show solutionHide solution
name = " CHANDU PRASAD "
name = name.strip().title()
print(name) # Chandu PrasadNormalise " USER@EXAMPLE.COM " to "user@example.com".
Show solutionHide solution
email = " USER@EXAMPLE.COM "
email = email.strip().casefold()
print(email) # user@example.comCount how many times "a" appears in "banana", and how many times "an" does.
Show hintHide hint
count() works on substrings as well as single characters.
Show solutionHide solution
text = "banana"
print(text.count("a")) # 3
print(text.count("an")) # 2 - non-overlapping matchesTurn " Python, React , Node.js , PostgreSQL " into a clean lowercase list.
Show solutionHide solution
skills = " Python, React , Node.js , PostgreSQL "
skills = [skill.strip().lower() for skill in skills.split(",")]
print(skills) # ['python', 'react', 'node.js', 'postgresql']Check whether "employee_data.CSV" is a CSV file, in a way that works for any capitalisation.
Show solutionHide solution
filename = "employee_data.CSV"
if filename.lower().endswith(".csv"):
print("CSV file") # CSV fileSplit "student@example.com" into the part before and the part after the @, using partition().
Show solutionHide solution
email = "student@example.com"
user, _, domain = email.partition("@")
print(user) # student
print(domain) # example.comGenerate the slug "learn-python-string-methods" from " Learn Python String Methods ".
Show hintHide hint
strip, lower, split, then join with a hyphen.
Show solutionHide solution
title = " Learn Python String Methods "
slug = "-".join(title.strip().lower().split())
print(slug) # learn-python-string-methodsText cleaner
Write a cleaner that takes whatever a user typed and returns something you would be willing to store. Then run it over the awkward inputs below and check every one.
- Trim the ends and lowercase the text
- Collapse runs of spaces and tabs into single spaces
- Return an empty string for input that is only whitespace
- Do it without a regular expression - the methods in this lesson are enough
- Run it over the test inputs and print each result in quotes so the spaces are visible
def clean(text):
...
for raw in [" PYTHON IS POWERFUL ",
"Hello\tWorld",
" ",
"already clean",
""]:
print(repr(raw), "->", repr(clean(raw)))Show one solutionHide solution
def clean(text):
# split() with no argument splits on any run of whitespace and drops the
# empties, so the collapsing happens for free - no regex needed.
return " ".join(text.lower().split())
for raw in [" PYTHON IS POWERFUL ",
"Hello\tWorld",
" ",
"already clean",
""]:
print(f"{raw!r:34} -> {clean(raw)!r}")
# ' PYTHON IS POWERFUL ' -> 'python is powerful'
# 'Hello\tWorld' -> 'hello world'
# ' ' -> ''
# 'already clean' -> 'already clean'
# '' -> ''Key points
- A string method returns a new string - it never modifies the original.
- Method chaining runs left to right, each call working on the previous result.
- strip() with an argument removes a set of characters from the ends, not a substring.
- removeprefix() and removesuffix() remove an exact prefix or suffix, which is usually what you wanted.
- split() with no argument splits on any whitespace and drops empties; split(" ") does not.
- join() is called on the separator and needs every item to already be a string.
- split, process, join is the shape of most text transformations.
- find() returns -1 when absent; index() raises ValueError. Use `in` when you only need yes or no.
- startswith() and endswith() accept a tuple, so several options can be checked at once.
- The is* methods all return False for an empty string.
- casefold() is for comparing, lower() is for displaying.
- The normalisation one-liner worth remembering: text = text.strip().casefold().
Quick check before you move on
Interview questions
Why is join() a method of str rather than of list?
Because it accepts any iterable of strings, not just lists. Putting it on str means one implementation covers lists, tuples, sets, and generators; putting it on the container would mean reimplementing it for each. It also makes the separator explicit at the call site, which is the part that varies.
Someone writes filename.strip(".csv") to remove an extension. What goes wrong?
strip() takes a set of characters, not a substring, and strips any of them from both ends until it hits a character outside the set. "scv.csv" becomes empty, and "csv.report.csv" loses characters from the front too. removesuffix(".csv") removes that exact suffix once, and leaves the string untouched if it is not there.
When would you use casefold() rather than lower()?
For case-insensitive comparison, particularly with non-English text. casefold applies the full Unicode case-folding rules, so German ß folds to ss and "straße" matches "STRASSE"; lower() leaves ß alone and the comparison fails. lower() is for what a person will read; casefold is for what a program will compare.
How would you count word frequency in a paragraph?
Normalise with strip and casefold, split with no argument so runs of whitespace collapse, strip punctuation from each word, then count - collections.Counter does the counting in one line. The subtlety is punctuation: "python." and "python" are different keys until you deal with it.
Why is building a string with += in a loop a problem?
Strings are immutable, so each += allocates a new string and copies everything accumulated so far, making the loop quadratic in the total length. Appending to a list and calling join() once allocates the result a single time. CPython has an optimisation that sometimes hides this, but it is not something to depend on.
What is the difference between split() and partition()?
split() returns a list of however many pieces the separator produces, and drops the separator. partition() always returns exactly three values - before, the separator, after - even when the separator is absent, in which case the last two are empty. Because the shape is fixed, partition unpacks safely into three names without a length check.
Quiz
- 1.
Does upper() change the original string?
- 2.
What is the difference between find() and index()?
- 3.
What does "scv.csv".strip(".csv") return, and why?
- 4.
What is the difference between split() and split(" ")?
- 5.
Why is join() called on the separator rather than the list?
- 6.
How do you check whether a filename is an image, allowing several extensions?
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned