Data Analysis · NumPy Foundations · lesson 4 of 12
Aggregation and axes
about 15 minutes · free · runs in your browser
Summarising an array
The aggregations are methods on the array, and each has a NumPy function twin:
a.sum() a.mean() a.min() a.max() a.std()
a.argmin() a.argmax() # the POSITION of the smallest / largest
For a two-dimensional array, axis decides which direction to collapse. This is the
part worth slowing down for:
m = np.array([[1, 2, 3],
[4, 5, 6]])
m.sum() # 21 — everything
m.sum(axis=0) # array([5, 7, 9]) — down the columns
m.sum(axis=1) # array([6, 15]) — across the rows
Read axis=0 as "collapse the rows, leaving one value per column". People remember it
as "axis 0 is vertical", which works as long as you also remember the result has one
entry per column.
Your turn: sales has one row per shop and one column per quarter. Find the total
per shop, the total per quarter, and which shop sold most overall.
You start from this, and edit it in the browser:
import numpy as np
# rows = shops, columns = quarters
sales = np.array([
[120, 150, 130, 170],
[200, 180, 220, 210],
[90, 110, 105, 95],
])
# Set per_shop, per_quarter and best_shop_index.