Data Structures
Lists, dictionaries, sets and tuples — what each is for, and how to work with it.
Choosing a container
| Need | Use |
|---|---|
| an ordered collection you will change | list |
| lookup by name or id | dict |
| uniqueness, or fast "is it in here?" | set |
| a fixed group of values that belong together | tuple |
Lists
items = ["a", "b", "c"]
items[0] # "a" — indexes start at zero
items[-1] # "c" — counts from the end
len(items) # 3
Slicing
items[1:3] # from 1, up to but not including 3
items[:2] # from the start
items[2:] # to the end
items[:] # a copy
items[::2] # every second item
items[::-1] # reversed copy
Methods
| Method | Effect | Returns |
|---|---|---|
.append(x) | add to the end | None |
.extend(other) | add every item of another list | None |
.insert(i, x) | add at position i | None |
.remove(x) | delete the first x | None |
.pop() / .pop(i) | remove and return | the item |
.sort() | sort in place | None |
.reverse() | reverse in place | None |
.count(x) | how many x | int |
.index(x) | position of first x | int |
sorted(items) # a NEW sorted list
sorted(items, reverse=True) # descending
sorted(items, key=len) # sort by a computed value
sum(nums), min(nums), max(nums)
The distinction that catches people: .sort() changes the list and returns
None; sorted() leaves it alone and returns a new one.
Iterating
for item in items:
for i, item in enumerate(items): # index and value
for i, item in enumerate(items, start=1): # numbering from one
for a, b in zip(list_a, list_b): # two lists together
Comprehensions
[n * 2 for n in nums] # transform
[n for n in nums if n > 0] # filter
[n * 2 for n in nums if n > 0] # both
Dictionaries
person = {"name": "Ada", "age": 36}
person["name"] # KeyError if missing
person.get("job") # None if missing
person.get("job", "unknown") # your own default
person["city"] = "London" # add or replace
del person["age"] # remove
"name" in person # checks KEYS, not values
Iterating
for key in person:
for value in person.values():
for key, value in person.items(): # the one you want most of the time
Idioms
counts[word] = counts.get(word, 0) + 1 # counting
groups.setdefault(key, []).append(value) # grouping
{k: v for k, v in d.items() if v > 0} # dict comprehension
Sets
tags = {"python", "beginner"}
empty = set() # {} makes an empty DICT
tags.add("web")
tags.discard("web") # no error if absent
| Operator | Meaning |
|---|---|
a | b | union — in either |
a & b | intersection — in both |
a - b | difference — in a only |
a ^ b | in exactly one |
No order and no indexing. To de-duplicate while keeping order:
seen = set()
result = [x for x in items if not (x in seen or seen.add(x))]
Tuples
point = (3, 4)
pair = 3, 4 # brackets optional
single = (3,) # the comma makes it a tuple
x, y = point # unpacking
a, b = b, a # swap
Immutable, so they can be dictionary keys and set members — lists cannot.
Nested data
people = [{"name": "Ada", "age": 36}]
people[0]["name"] # read left to right
config = {"server": {"port": 8080}}
config["server"]["port"]
Copying
shallow = items[:] # or list(items)
import copy
deep = copy.deepcopy(nested) # when the contents are themselves containers
b = a does not copy — both names point at the same list.