Data Analysis · Analysis Workflow · lesson 11 of 12
Reading CSV and JSON
about 16 minutes · free · runs in your browser
Getting data in
pandas reads most formats in one call:
pd.read_csv("sales.csv")
pd.read_json("sales.json")
pd.read_csv(path, parse_dates=["date"]) # real dates, not strings
pd.read_csv(path, na_values=["", "n/a", "-"]) # what counts as missing
And writes them back the same way — df.to_csv(path, index=False), where
index=False stops pandas adding a column of row numbers nobody asked for.
When the data is already in memory, StringIO lets you read a string as if it were a
file, which is how you write a test for a parser:
from io import StringIO
pd.read_csv(StringIO(text))
Your turn: parse the CSV text into a DataFrame, treating n/a as missing, and
report the row count and the total revenue.
You start from this, and edit it in the browser:
import pandas as pd
from io import StringIO
text = """product,units,price
Widget,10,2.50
Gadget,n/a,10.00
Doohickey,4,7.25
"""
# Set df, row_count and total_revenue (units * price, ignoring the missing row).