Sets
Unique values, instant membership, and comparing two groups with one operator.
What you will be able to do
- Create sets correctly, including the empty set that {} does not give you
- Explain why a set holds no duplicates and supports no indexing
- Add and remove elements, and choose between remove() and discard()
- Use union, intersection, difference, and symmetric difference as operators or methods
- Test subset, superset, and disjoint relationships
- Say why membership in a set is fast and membership in a list is not
- Remove duplicates while preserving order, using a set as a seen-marker
- Choose between a list, a tuple, and a set from what the data has to do
The idea, in plain English
A set is a collection of unique values with no positions. Adding a value that is already there does nothing, there is no index 0, and the order you see when you print one is not something to rely on. Everything else about sets follows from those two facts.
They exist for three jobs, and each replaces code you would otherwise write by hand. Uniqueness: set(emails) removes duplicates in one call. Membership: `role in allowed_roles` answers instantly no matter how big the set is, where the same question on a list walks the whole thing. Comparison: `required - available` is exactly "what is missing", and `a & b` is exactly "what both have".
That third job is the one worth internalising, because it turns loops into single expressions. Permission checks, skill gaps, comparing two learning paths, finding which tags are shared - all of these are nested loops written with lists and one operator written with sets.
The trade is what you give up. No order, no duplicates, no indexing, and every element must be hashable - so a set can hold tuples but never lists. When order is part of the data, you want a list; when uniqueness and comparison are the point, you want a set.
Worked example: Checking a user against the permissions a feature requires.
{} is a dictionary, and other creation traps
Curly braces were dictionaries before they were sets, so the empty braces went to dictionaries and sets got the leftovers. `items = {}` gives you an empty dict, and type() confirms it. The empty set is written set().
This one is nastier than the missing tuple comma, because an empty dict and an empty set behave alike for a while. Both are falsy, both have len 0, both support `in`. The difference only surfaces when you call .add() and get AttributeError, often far from where the variable was created.
Non-empty is fine: {1, 2, 3} is unambiguously a set, because a dictionary would need colons. And set() accepts any iterable, which is how most sets really get built - set(a_list), set(a_tuple), set("Python") for unique characters.
Every element must be hashable, which in practice means immutable. Numbers, strings, and tuples of hashables are fine. Lists are not: {[1, 2]} raises TypeError: unhashable type: list, which is the same rule you met for dictionary keys in lesson 17, for the same reason.
{1, 2, 3}A set literal. Unambiguous because there are no colons.{}An empty dictionary. Not a set. This is the trap.set()The empty set, and the only way to write one.set([1, 2, 2])From any iterable, dropping duplicates: {1, 2}.set("Python")The unique characters, since a string is iterable.{[1, 2]}TypeError. Lists are mutable, so they are unhashable.{(1, 2)}Fine. A tuple of hashable values is itself hashable.Watch out: print() of an empty set shows set(), not {} - which is Python telling you the same thing.
No order, no index, no duplicates
A set stores its elements by hash, not by position, so there is no first element to ask for. courses[0] raises TypeError. If you genuinely need a position, you wanted a list.
Iteration order is likewise not a promise. It is stable within a single run for a given set, which is exactly enough to fool you into depending on it, and it can differ between runs and between Python versions. When you need predictable output, sort it: `for course in sorted(courses)`. Note that sorted() hands back a list, not a set.
Adding a duplicate is not an error, it is a no-op. courses.add("Python") on a set that already has it changes nothing and raises nothing. That is the whole reason set(values) de-duplicates - it is not a special function, just the normal behaviour applied to each element in turn.
set.pop() removes and returns an arbitrary element, and takes no argument. pop(0) is an error, not a way to get the first item. If you want a particular value gone, name it with remove() or discard().
items[0]TypeError - 'set' object is not subscriptable. There are no positions.items.sort()Does not exist. Use sorted(items), which returns a list.items.pop(0)TypeError. set.pop() takes no argument and removes an arbitrary element.items.add("x") twiceThe second call is a silent no-op, not an error.Relying on print orderStable enough to mislead you, never guaranteed. Sort when it matters.remove() raises, discard() does not
Both take a value and delete it. The difference is what happens when the value was not there: remove() raises KeyError, discard() does nothing at all.
That makes it a genuine design choice rather than a style preference. Use remove() when the element is supposed to exist and its absence means something has gone wrong upstream - the exception is the bug report. Use discard() when removal should be idempotent, which is most cleanup code: permissions.discard("delete_course") should work whether or not the permission was granted.
Note that a set raises KeyError where a list raises ValueError for the same mistake. Sets are keyed collections, like dictionaries, so they report missing keys the way dictionaries do.
On the adding side the same split exists between add() and update(). add(x) puts in one element; update(xs) iterates and adds each. And once again, a string is iterable - letters.update("cd") adds "c" and "d" separately, while letters.add("cd") adds one two-character string.
add(x)One element. Silently does nothing if it is already present.update(xs)Each element of an iterable. Accepts several iterables at once.remove(x)Delete x. KeyError if absent - use it when absence is a bug.discard(x)Delete x if present, otherwise nothing. Use it for idempotent cleanup.pop()Remove and return an arbitrary element. KeyError on an empty set.clear()Remove everything, keeping the set object.Watch out: add([1, 2]) raises TypeError because a list is unhashable. update([1, 2]) works and adds 1 and 2 - two very different outcomes from one bracket.
The four operations, and what each one answers
Union (A | B) is everything from both. Intersection (A & B) is what they share. Difference (A - B) is what A has that B does not. Symmetric difference (A ^ B) is what exactly one of them has.
Each has a method form - union, intersection, difference, symmetric_difference - and the methods accept several arguments and any iterable, while the operators require sets on both sides. So a.union(b, c) and a.union([1, 2]) both work; a | [1, 2] does not.
Difference is the only one with a direction, and it is the one you will reach for most. {1, 2, 3} - {2, 3, 4} is {1}, while the reverse is {4}. In application terms `required - available` reads as "what is missing" and `available - required` as "what is extra" - two different questions, one operator, and getting them the wrong way round is a real bug that raises nothing.
Read them as questions rather than as set theory. Which skills does this candidate lack? required - candidate. Which courses appear in both paths? path_a & path_b. What is the full catalogue across both? path_a | path_b. Written with lists, each of those is a loop.
A | B{1, 2, 3, 4, 5} - union. Everything, each value once. A.union(B).A & B{3} - intersection. Only what both hold. A.intersection(B).A - B{1, 2} - in A, not in B. A.difference(B). Directional.B - A{4, 5} - the other direction, and a different answer.A ^ B{1, 2, 4, 5} - in one or the other, never both. A.symmetric_difference(B).A.union(B, C)The method form takes several arguments; the operator form takes one.Relationships, and why membership is fast
Beyond combining sets you can ask how they relate. A <= B asks whether every element of A is also in B - subset. A >= B is the reverse - superset. A.isdisjoint(B) asks whether they share nothing at all, and is clearer than checking that an intersection is empty.
Subset is the natural way to write a permission check: `if required <= user_permissions:` says "the user has all of these" in one line, with no loop and no counter. The method forms issubset and issuperset read more explicitly and accept any iterable.
The strict forms < and > add "and they are not equal", so A < B means A is contained in B but smaller. A <= A is True; A < A is False. Most application checks want the non-strict version.
All of this rests on hashing. A set computes a hash of each value and stores it in a table, so `x in s` jumps more or less straight to the answer regardless of size - O(1) on average. A list has to compare against each element in turn - O(n). On ten items nobody notices; on a hundred thousand, checked inside a loop, that difference is the reason a page takes seconds.
A <= BSubset - every element of A is in B. A.issubset(B).A < BProper subset - a subset, and not equal.A >= BSuperset - A contains everything in B. A.issuperset(B).A > BProper superset - a superset, and not equal.A.isdisjoint(B)True when they have nothing in common. Clearer than not (A & B).x in AO(1) on average, against O(n) for a list. The reason sets exist.Tip: Converting a list to a set to test membership once is not worth it - building the set is itself O(n). It pays off when you test many times.
Syntax and examples
courses = {"Python", "React", "FastAPI"}
print(len(courses)) # 3
# Duplicates simply collapse
numbers = {10, 20, 10, 30, 20}
print(numbers) # {10, 20, 30}
print(len(numbers)) # 3
# The trap: {} is a dictionary
items = {}
print(type(items)) # <class 'dict'>
items = set()
print(type(items)) # <class 'set'>
print(items) # set() <- not {}
# set() converts any iterable
print(set([1, 2, 2, 3, 4, 4])) # {1, 2, 3, 4}
print(set((10, 20, 20, 30))) # {10, 20, 30}
print(len(set("Python"))) # 6 - the unique characters
# Elements must be hashable
ok = {1, "Python", (1, 2)}
# bad = {[1, 2]} -> TypeError: unhashable type: 'list'courses = {"Python", "React", "AWS"}
# print(courses[0]) -> TypeError: not subscriptable
# courses.sort() -> AttributeError: no attribute 'sort'
# Iteration order is not a promise - sort when output matters
for course in sorted(courses):
print(course)
# AWS
# Python
# React
print(sorted(courses)) # ['AWS', 'Python', 'React'] <- a list
print(type(sorted(courses))) # <class 'list'>
# If you really need a position, convert - but the order is arbitrary
as_list = list(courses)
print(as_list[0]) # whichever one happened to be first
# Adding a duplicate is a no-op, not an error
items = set()
items.add("Python")
items.add("Python")
items.add("React")
print(len(items)) # 2courses = {"Python", "React"}
courses.add("FastAPI") # one element
print(len(courses)) # 3
courses.update(["PostgreSQL", "Docker"]) # each element of an iterable
print(len(courses)) # 5
# A string is iterable, so update splits it
letters = {"a", "b"}
letters.update("cd")
print(sorted(letters)) # ['a', 'b', 'c', 'd']
letters = {"a", "b"}
letters.add("cd")
print(sorted(letters)) # ['a', 'b', 'cd']
# add() will not take a list, but update() will
# courses.add(["Go", "Rust"]) -> TypeError: unhashable type: 'list'
courses.update(["Go", "Rust"]) # adds two elements
# remove() insists the value is there
permissions = {"read", "write"}
permissions.remove("write")
# permissions.remove("delete") -> KeyError: 'delete'
# discard() does not care
permissions.discard("delete") # no error, nothing happens
print(permissions) # {'read'}
# pop() takes no argument and gives back an arbitrary element
items = {"A", "B", "C"}
item = items.pop()
print(item in {"A", "B", "C"}) # True - but which one is not defined
# items.pop(0) -> TypeError
items.clear()
print(items) # set()A = {1, 2, 3}
B = {3, 4, 5}
print(A | B) # {1, 2, 3, 4, 5} union - everything
print(A & B) # {3} intersection - shared
print(A - B) # {1, 2} difference - in A only
print(B - A) # {4, 5} the other direction
print(A ^ B) # {1, 2, 4, 5} symmetric - in exactly one
# Method forms are identical, and take more than one argument
print(A.union(B)) # {1, 2, 3, 4, 5}
print(A.intersection(B)) # {3}
print(A.difference(B)) # {1, 2}
print(A.symmetric_difference(B)) # {1, 2, 4, 5}
frontend = {"React", "Next.js"}
backend = {"Python", "FastAPI"}
devops = {"Docker", "AWS"}
print(len(frontend.union(backend, devops))) # 6
print(len(frontend | backend | devops)) # 6 - operators chain too
team_a = {"Python", "Docker", "AWS", "Git"}
team_b = {"Python", "Docker", "Git"}
team_c = {"Python", "Git"}
print(sorted(team_a.intersection(team_b, team_c))) # ['Git', 'Python']
# The methods accept any iterable; the operators need sets
print(A.union([9, 10])) # {1, 2, 3, 9, 10}
# print(A | [9, 10]) -> TypeError: unsupported operand type(s)user_permissions = {"read", "write", "view_reports"}
required_permissions = {"read", "view_reports"}
# "Does the user have all of these?" - one line, no loop
if required_permissions <= user_permissions:
print("Access: Granted")
else:
print("Access: Denied")
# The two directions of difference answer two different questions
missing = required_permissions - user_permissions
extra = user_permissions - required_permissions
print("Missing:", missing) # Missing: set()
print("Extra:", extra) # Extra: {'write'}
# The same shape, as a skills gap
required = {"Python", "FastAPI", "PostgreSQL", "Docker"}
candidate = {"Python", "Docker"}
print(sorted(required - candidate)) # ['FastAPI', 'PostgreSQL']
# Comparing two learning paths
backend_path = {"Python", "FastAPI", "PostgreSQL", "Docker", "AWS"}
fullstack_path = {"Python", "React", "Node.js", "PostgreSQL", "Docker"}
print(sorted(backend_path & fullstack_path)) # ['Docker', 'PostgreSQL', 'Python']
print(sorted(backend_path - fullstack_path)) # ['AWS', 'FastAPI']
print(sorted(fullstack_path - backend_path)) # ['Node.js', 'React']
print(len(backend_path | fullstack_path)) # 7
print(backend_path <= fullstack_path) # False - not a subset
# Relationships
print({1, 2}.issubset({1, 2, 3})) # True
print({1, 2, 3}.issuperset({1, 2})) # True
print({1, 2} < {1, 2}) # False - proper subset excludes equality
print({1, 2} <= {1, 2}) # True
print({1, 2}.isdisjoint({3, 4})) # Trueemails = [
"ravi@example.com",
"anita@example.com",
"ravi@example.com",
"chandu@example.com",
]
unique_emails = set(emails)
print(len(emails), len(unique_emails)) # 4 3
# Comparing the two lengths is the whole duplicate check
if len(emails) != len(unique_emails):
print("Duplicates found")
user_ids = [101, 102, 103, 101, 104, 102]
print(len(user_ids) != len(set(user_ids))) # True
# set() alone loses the original order
courses = ["Python", "React", "Python", "AWS", "React", "Docker"]
print(sorted(set(courses))) # alphabetical, not original
# Keeping order: a set marks what has been seen, a list keeps the order.
# The set is doing the fast membership test; the list is doing the ordering.
seen = set()
unique = []
for course in courses:
if course in seen:
continue
seen.add(course)
unique.append(course)
print(unique) # ['Python', 'React', 'AWS', 'Docker']
# Set comprehensions build a set directly
numbers = [1, 2, 2, 3, 4, 4]
print({number ** 2 for number in numbers}) # {1, 4, 9, 16}
print({n for n in range(1, 11) if n % 2 == 0}) # {2, 4, 6, 8, 10}
# frozenset is a set that cannot change - and so can be hashed
roles = frozenset({"admin", "teacher"})
print("admin" in roles) # True
# roles.add("manager") -> AttributeError
groups = {frozenset({"read"}): "viewer"} # usable as a dict keyTip: frozenset is the immutable set. Because it cannot change it is hashable, so it can be a dictionary key or an element of another set - which a normal set cannot.
Watch out: Converting to a set to test membership once costs O(n) to build, so it saves nothing. The win comes when you test repeatedly, especially inside a loop.
Set methods
The top group changes the set in place. The bottom group answers a question and returns a new set or a boolean.
add(x)Add one hashable element. A no-op if it is already there.
courses.add("Go")update(xs)Add each element of one or more iterables.
courses.update(["Go", "Rust"])
remove(x)Delete x. KeyError if it is absent.
roles.remove("admin")discard(x)Delete x if present. Never raises.
roles.discard("admin")pop()Remove and return an arbitrary element. Takes no argument.
item = items.pop()
clear()Empty the set, keeping the object.
items.clear()
union(*others)A new set with everything. Operator: |
a.union(b, c)
intersection(*others)A new set with what they all share. Operator: &
a.intersection(b)
difference(*others)In this set, not the others. Operator: -
required.difference(have)
symmetric_difference(other)In one or the other, not both. Operator: ^
a.symmetric_difference(b)
issubset(other)Every element is in other. Operator: <=
required.issubset(have)
issuperset(other)Contains every element of other. Operator: >=
have.issuperset(required)
isdisjoint(other)Nothing in common. No operator form.
a.isdisjoint(b)
List against tuple against set
Three collections, three jobs. The question to ask is what the data has to do, not which is shortest to type.
Ordered and indexableList yes, tuple yes, set no. A set has no positions at all.
MutableList yes, tuple no, set yes. frozenset is the immutable set.
Duplicates allowedList yes, tuple yes, set no - duplicates collapse on the way in.
Membership costList and tuple O(n); set O(1) on average. This is usually the deciding factor.
Set operationsSet only. Union, intersection, and difference have no list equivalent short of loops.
Usable as a dict keyTuple yes if its contents are hashable; frozenset yes; list and set never.
Reach for it whenList: an ordered collection that changes. Tuple: a fixed record. Set: uniqueness, membership, comparison.
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Write items = {} and print its type. Then call items.add("x") and read the error. Fix it with set() and try again.”
“Build a set of five or six strings and print it. Then print sorted() of it. Ask yourself which of the two you would be willing to put in a test assertion.”
“Call remove() with a value that is not in the set, then discard() with the same value. Decide which one you would want in cleanup code that might run twice.”
“With required = {"read", "write"} and have = {"read"}, print both required - have and have - required. Name the question each one answers.”
“Call tags.add("react") and tags.update("react") on two copies of the same set and compare the lengths. Then try add(["a", "b"]) and read the error.”
What usually goes wrong
Empty braces create a dictionary. It stays hidden for a while because both are falsy, both have len 0, and both support `in` - until .add() raises AttributeError somewhere else entirely.
✗ seen = {}
seen.add("Python")
# AttributeError: 'dict' object has no attribute 'add'✓ seen = set()
seen.add("Python")Iteration order is stable enough within one run to look reliable and is not guaranteed between runs or versions. A test that asserts on it will fail eventually.
✗ for course in courses:
print(course) # order not defined✓ for course in sorted(courses):
print(course) # and sorted() returns a listremove() raises KeyError - note KeyError, not the ValueError a list would raise. In cleanup code that may run more than once, that is a crash rather than a no-op.
✗ permissions.remove("delete")
# KeyError: 'delete' if it was never granted✓ permissions.discard("delete") # idempotentEvery element must be hashable and a list is not. The confusing part is that update() with the same argument succeeds - and does something quite different.
✗ items.add([1, 2])
# TypeError: unhashable type: 'list'✓ items.add((1, 2)) # one tuple element
# or
items.update([1, 2]) # two separate elementsset.pop() takes no argument and removes an arbitrary element. There is no positional access to build an index on.
✗ items.pop(0)
# TypeError: set.pop() takes no arguments✓ items.remove(value) # or discard(value) if it may be absentA - B and B - A answer different questions and neither raises. "What is missing" is required - available; "what is extra" is the other way round.
✗ missing = available - required # this is the extras✓ missing = required - available
extra = available - requiredIt de-duplicates but gives no ordering guarantee. Keep a set for the membership test and a list for the order.
✗ unique = list(set(courses)) # original order lost✓ seen = set()
unique = []
for course in courses:
if course not in seen:
seen.add(course)
unique.append(course)Best practices
- Write set() for an empty set, and never {} - that is a dictionary.
- Reach for a set when the question is "is this present?" or "what do these two groups share?".
- Sort before printing or asserting on a set; never depend on its iteration order.
- Use discard() for cleanup that may run twice, and remove() when a missing element means a bug.
- Write permission checks as required <= available rather than looping over each one.
- Name the direction of a difference in the variable: missing = required - available.
- Keep a list alongside a set when you need both order and fast membership.
- Use frozenset when the set must not change or has to be hashable.
- Remember every element must be hashable - tuples are welcome, lists never are.
Practice
Write these yourself before opening anything. Getting them wrong first is most of how this sticks.
Given numbers = [1, 2, 2, 3, 4, 4, 5, 5], build a set of only the unique values and show how many there are.
Show hintHide hint
One call converts the list; sort it if you want predictable output.
Show solutionHide solution
numbers = [1, 2, 2, 3, 4, 4, 5, 5]
unique = set(numbers)
print(sorted(unique)) # [1, 2, 3, 4, 5]
print(len(unique)) # 5
print(len(numbers)) # 8 - so three duplicates were droppedFor student_a = {"Python", "SQL", "Git", "Docker"} and student_b = {"Python", "Git", "AWS", "Docker"}, find the courses in common, those only A has, those only B has, and every unique course.
Show hintHide hint
Four operations, one per question. Watch the direction on the two middle ones.
Show solutionHide solution
student_a = {"Python", "SQL", "Git", "Docker"}
student_b = {"Python", "Git", "AWS", "Docker"}
print(sorted(student_a & student_b)) # ['Docker', 'Git', 'Python']
print(sorted(student_a - student_b)) # ['SQL']
print(sorted(student_b - student_a)) # ['AWS']
print(sorted(student_a | student_b))
# ['AWS', 'Docker', 'Git', 'Python', 'SQL']
# The symmetric difference is the two middles combined
print(sorted(student_a ^ student_b)) # ['AWS', 'SQL']Check whether a user with {"read", "write", "delete"} has all of the required {"read", "write"}.
Show hintHide hint
This is a subset question, and it fits on one line.
Show solutionHide solution
required = {"read", "write"}
user_permissions = {"read", "write", "delete"}
if required <= user_permissions:
print("Access granted")
# The method form reads more explicitly, and does the same thing
print(required.issubset(user_permissions)) # True
print(user_permissions.issuperset(required)) # TrueA role requires {"Python", "FastAPI", "PostgreSQL", "Docker"} and a candidate has {"Python", "Docker"}. Find the missing skills.
Show hintHide hint
Difference, and the order of the two operands decides which question you asked.
Show solutionHide solution
required = {"Python", "FastAPI", "PostgreSQL", "Docker"}
candidate = {"Python", "Docker"}
missing = required - candidate
print(sorted(missing)) # ['FastAPI', 'PostgreSQL']
# The other direction would have answered a different question -
# skills the candidate has that the role does not ask for.
print(sorted(candidate - required)) # []From tags = ["python", "backend", "api", "python", "backend", "docker"], build the set of unique tags and print them alphabetically.
Show hintHide hint
sorted() on a set gives a list back, which is what you want for display.
Show solutionHide solution
tags = ["python", "backend", "api", "python", "backend", "docker"]
unique_tags = set(tags)
print(len(unique_tags)) # 4
for tag in sorted(unique_tags):
print(tag)
# api
# backend
# docker
# pythonFrom courses = ["Python", "React", "Python", "AWS", "React", "Docker"], remove duplicates while keeping the original order. Use a set for the membership test.
Show hintHide hint
Two containers: one set to remember what you have seen, one list to hold the order.
Show solutionHide solution
courses = ["Python", "React", "Python", "AWS", "React", "Docker"]
seen = set()
unique = []
for course in courses:
if course not in seen:
seen.add(course)
unique.append(course)
print(unique) # ['Python', 'React', 'AWS', 'Docker']
# The set does the fast checking, the list does the ordering.
# list(set(courses)) would have been shorter and lost the order.Access review
One user, one feature, and two learning paths. Answer every question with set operations rather than loops, and print the results in a stable order.
- Report whether the user has every required permission, as Granted or Denied.
- Print the missing permissions and the extra permissions separately, and label which is which.
- For the two learning paths, print the shared courses, the ones unique to each, and the full combined catalogue.
- Say whether the backend path is a subset of the full-stack path, and whether the two share nothing at all.
- Detect whether the enrolment list contains duplicate user ids, and list the unique ids in their original order.
user_permissions = {"read", "write", "view_reports"}
required_permissions = {"read", "view_reports", "export"}
backend_path = {"Python", "FastAPI", "PostgreSQL", "Docker", "AWS"}
fullstack_path = {"Python", "React", "Node.js", "PostgreSQL", "Docker"}
enrolments = [101, 102, 103, 101, 104, 102]
# 1. Granted or Denied
# 2. Missing and extra permissions
# 3. Shared, unique to each, and the full catalogue
# 4. Subset and disjoint checks
# 5. Duplicate ids, and the unique ids in original order
Show one solutionHide solution
user_permissions = {"read", "write", "view_reports"}
required_permissions = {"read", "view_reports", "export"}
backend_path = {"Python", "FastAPI", "PostgreSQL", "Docker", "AWS"}
fullstack_path = {"Python", "React", "Node.js", "PostgreSQL", "Docker"}
enrolments = [101, 102, 103, 101, 104, 102]
# 1. Subset is exactly "has all of these"
granted = required_permissions <= user_permissions
print("Access:", "Granted" if granted else "Denied")
# Access: Denied - export is missing
# 2. The two directions answer two different questions
print("Missing:", sorted(required_permissions - user_permissions))
print("Extra:", sorted(user_permissions - required_permissions))
# Missing: ['export']
# Extra: ['write']
# 3. Comparing the paths - sorted() so the output is stable
print("Shared:", sorted(backend_path & fullstack_path))
print("Backend only:", sorted(backend_path - fullstack_path))
print("Fullstack only:", sorted(fullstack_path - backend_path))
print("All:", sorted(backend_path | fullstack_path))
# Shared: ['Docker', 'PostgreSQL', 'Python']
# Backend only: ['AWS', 'FastAPI']
# Fullstack only: ['Node.js', 'React']
# All: ['AWS', 'Docker', 'FastAPI', 'Node.js', 'PostgreSQL', 'Python', 'React']
# 4. Relationships
print("Subset:", backend_path <= fullstack_path) # False
print("Disjoint:", backend_path.isdisjoint(fullstack_path)) # False
# 5. Comparing lengths detects duplicates in one line
unique_ids = set(enrolments)
print("Duplicates:", len(enrolments) != len(unique_ids)) # True
# set() alone would lose the order, so keep a list beside the set
seen = set()
ordered_unique = []
for user_id in enrolments:
if user_id not in seen:
seen.add(user_id)
ordered_unique.append(user_id)
print("Unique in order:", ordered_unique) # [101, 102, 103, 104]Key points
- {} is an empty dictionary. The empty set is set(), and printing one shows set().
- A set holds unique, hashable elements - tuples are allowed, lists never are.
- There is no order and no indexing. items[0] is a TypeError; sort with sorted(), which returns a list.
- Adding a duplicate is a silent no-op, which is why set(values) de-duplicates.
- add() takes one element, update() takes each element of an iterable - and a string is an iterable.
- remove() raises KeyError when the value is absent; discard() does nothing.
- set.pop() takes no argument and removes an arbitrary element.
- A | B union, A & B intersection, A - B difference, A ^ B symmetric difference.
- Difference is directional: required - available is what is missing, the reverse is what is extra.
- A <= B is a subset test and the cleanest way to write "has all of these".
- Membership in a set is O(1) on average against O(n) for a list - that is what sets are for.
- To de-duplicate and keep order, use a set to track what you have seen and a list to hold the order.
- frozenset is the immutable set, and being immutable it is hashable.
Quick check before you move on
Interview questions
What is a set in Python?
A mutable, unordered collection of unique hashable elements, implemented as a hash table. It supports membership testing in O(1) on average and the mathematical set operations - union, intersection, difference, and symmetric difference.
Are sets ordered?
No, and they should be treated as unordered. Iteration order is an artefact of hashing and insertion - stable within a run, but not guaranteed across runs or versions. Sort explicitly when the order matters.
How do you create an empty set, and what is the common mistake?
set(). Writing {} gives an empty dictionary, because dictionaries claimed the brace syntax first. It is easy to miss because both are falsy and both have length 0.
What is the difference between remove() and discard()?
Both delete by value. remove() raises KeyError when the element is absent; discard() does nothing. Use remove() when absence signals a bug and discard() when removal should be idempotent.
What is the difference between add() and update()?
add(x) inserts x as a single element and requires it to be hashable. update(xs) iterates one or more iterables and inserts each element. add(["a"]) raises TypeError while update(["a"]) works - the same argument, two very different results.
Explain union, intersection, difference, and symmetric difference.
Union (|) is every element from both. Intersection (&) is those in both. Difference (-) is those in the first and not the second, and is directional. Symmetric difference (^) is those in exactly one of the two.
What is the difference between the operator and method forms?
They compute the same result, but the methods accept any iterable and more than one argument - a.union(b, c) or a.union([1, 2]) - while the operators require sets on both sides and a | [1, 2] raises TypeError.
Why is membership testing faster on a set than a list?
A set hashes the value and looks in the corresponding bucket, so the cost does not grow with the collection - O(1) on average. A list compares against each element in turn, which is O(n). The difference is what makes sets worth converting to when you test repeatedly.
Can a list be an element of a set, and why?
No. Elements must be hashable, and lists are mutable, so their hash could change after insertion and the set would no longer find them. Tuples of hashable values are fine, as is frozenset.
What is a frozenset and when would you use one?
An immutable set. Because it cannot change it is hashable, so unlike a normal set it can be a dictionary key or an element of another set. It also states that a group of values is fixed - a permission group, for instance.
How would you remove duplicates while preserving order?
Iterate the original, keep a set of values already seen for O(1) checks, and append to a result list only when the value is new. list(set(values)) is shorter but discards the ordering.
How do you decide between a list, a tuple, and a set?
By what the data must do. A list for an ordered collection that changes; a tuple for a fixed record, which can also be a dictionary key; a set when values must be unique, membership is tested often, or you need to compare two groups.
Quiz
- 1.
How do you write an empty set, and why not {}?
- 2.
Which removal method should cleanup code use?
- 3.
How do you check that a user has every required permission?
- 4.
Why can a set hold a tuple but not a list?
- 5.
What does list(set(values)) lose?
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned