Data Structures · Dictionaries and Sets · lesson 9 of 10
Sets
about 12 minutes · free · runs in your browser
Step 1 of 2
Unordered, no duplicates
A set holds unique values with no particular order:
tags = {"python", "beginner", "python"} # {"python", "beginner"}
empty = set() # {} would make an empty dictionary
Two things sets are good at. First, removing duplicates:
unique = list(set(items))
Second, answering "is this in here?" — for a large collection, x in some_set is far
faster than x in some_list, because a set does not have to look through everything.
They also do the operations you remember from school maths:
a | b # union — everything in either
a & b # intersection — only what is in both
a - b # difference — in a but not b
The trade-off: no order and no indexing. my_set[0] is an error.
Your turn: find the tags common to both lists, and everything either post used.
You start from this, and edit it in the browser:
post_one = ["python", "beginner", "tutorial"]
post_two = ["python", "advanced", "tutorial"]
# Set shared and everything (both as sets).
Step 2 of 2
De-duplicating while keeping order
set() loses the original order, which is often fine and sometimes not. When order
matters, walk the list and use a set only to remember what you have already seen:
seen = set()
result = []
for item in items:
if item not in seen:
seen.add(item)
result.append(item)
Your turn: write unique(items) returning the items with duplicates removed, in the
order they first appeared.
You start from this, and edit it in the browser:
def unique(items):
# Remove duplicates but keep the first-seen order.
pass