Warm-up · Activity 1 of 7
Warm-up from lesson I3.2: a generator can be used once. What does this print?
g = (n for n in [1, 2])
print(list(g), list(g))// I3.6 · ~40 min · Intermediate
In this build you read an orders file lazily with csv and a generator, then count with Counter, group with defaultdict, name fields with namedtuple and keep the latest orders with deque.
Lesson 6 of 6 in I3 Iteration and functional tools
You will be able to
Warm-up · Activity 1 of 7
g = (n for n in [1, 2])
print(list(g), list(g))Predict · Activity 2 of 7
from collections import Counter
c = Counter("banana")
print(c["a"], c["z"])Practice · Activity 3 of 7
from collections import defaultdict
groups = defaultdict(____)
for word in ["apple", "bean", "avocado"]:
groups[word[0]].append(word)
print(dict(groups))Practice · Activity 4 of 7
from collections import deque
d = deque(maxlen=3)
for n in range(1, 6):
d.append(n)
print(list(d))Practice · Activity 5 of 7
Brain teaser · Activity 6 of 7
from collections import deque
def numbers():
for n in range(1, 4):
print("made", n, end="; ")
yield n
print(deque(numbers(), maxlen=1))Apply · Activity 7 of 7
Check your work against this list
Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.
Worked example
The visit log is a list of CSV lines, so it runs without a file; csv.DictReader accepts any iterable of strings, including an open file. visits() is a generator that turns each row into a Visit namedtuple. Every consumer calls visits(LINES) again for a fresh pass: Counter counts pages, a defaultdict groups pages by user, and a bounded deque keeps the two latest visits.
main.py
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Visit = namedtuple("Visit", ["page", "user", "seconds"])
LINES = [
"page,user,seconds",
"home,ada,5",
"docs,bo,30",
"home,bo,3",
"blog,ada,12",
"docs,ada,8",
"home,cy,4",
]
def visits(lines: Iterable[str]) -> Iterator[Visit]:
"""Yield one Visit per CSV row, converting seconds to int."""
for row in csv.DictReader(lines):
yield Visit(row["page"], row["user"], int(row["seconds"]))
first = next(visits(LINES))
print(first, first.page, first[2])
views = Counter(v.page for v in visits(LINES))
print(views.most_common(2), views["shop"])
pages_by_user: defaultdict[str, list[str]] = defaultdict(list)
for v in visits(LINES):
pages_by_user[v.user].append(v.page)
print(dict(pages_by_user))
recent = deque(visits(LINES), maxlen=2)
print([v.page for v in recent])
Run it with
python main.pyOutput
Visit(page='home', user='ada', seconds=5) home 5
[('home', 3), ('docs', 2)] 0
{'ada': ['home', 'blog', 'docs'], 'bo': ['docs', 'home'], 'cy': ['home']}
['docs', 'home']Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.
The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.
Exercise 1 of 4
orders.csv (second tab) has the header customer,product,qty. The starter splits each line on commas, which breaks on the quoted product "tea, green", keeps qty as a string and builds a whole list. Rewrite read_orders(path) as a generator: open the file with newline="", read it with csv.DictReader and yield one Order per row, with qty converted to int.
Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.
The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.
Inside the with block: for row in csv.DictReader(f):. Each row is a dict keyed by the header, so row["qty"] is the quantity.
Replace orders.append(...) with yield Order(...), drop the list, and change the return type to Iterator[Order].
yield Order(row["customer"], row["product"], int(row["qty"])). DictReader reads the header itself, so next(f) must go.
One way to solve it. Yours can look different and still pass the checks.
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
if __name__ == "__main__":
for order in read_orders("orders.csv"):
print(order)
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> list[Order]:
"""Read every row of the CSV file at path into a list of Orders."""
orders: list[Order] = []
with open(path, encoding="utf-8") as f:
next(f) # skip the header line
for line in f:
customer, product, qty = line.strip().split(",")
orders.append(Order(customer, product, qty))
return orders
if __name__ == "__main__":
for order in read_orders("orders.csv"):
print(order)
test_main.py
from main import Order, read_orders
def test_first_order():
"""The first order is ada, tea, 2"""
got = next(iter(read_orders("orders.csv")))
assert got == Order("ada", "tea", 2), f"the first order is {got!r}, expected ada, tea and the int 2"
def test_qty_is_int():
"""qty is converted to int"""
got = [order.qty for order in read_orders("orders.csv")]
assert got == [2, 1, 3, 1, 2, 1], f"the quantities are {got!r}, expected the ints [2, 1, 3, 1, 2, 1]"
def test_quoted_comma():
"""A quoted product may contain a comma"""
got = list(read_orders("orders.csv"))[4]
assert got == Order("bo", "tea, green", 2), f"the fifth order is {got!r}, expected bo, tea, green and 2"
def test_lazy():
"""read_orders is a generator, so it returns an iterator"""
orders = read_orders("orders.csv")
assert iter(orders) is orders, f"read_orders returned a {type(orders).__name__}: use yield instead of building a list"
def test_other_file():
"""Another file gives its own orders"""
with open("other.csv", "w", encoding="utf-8", newline="") as f:
f.write("customer,product,qty\nzoe,jam,5\n")
got = list(read_orders("other.csv"))
assert got == [Order("zoe", "jam", 5)], f"for other.csv read_orders gave {got!r}"
orders.csv
customer,product,qty
ada,tea,2
bo,coffee,1
ada,coffee,3
cy,tea,1
bo,"tea, green",2
ada,cake,1
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 2 of 4
units_by_product(orders) should add up the qty of all orders per product and return a Counter, so that most_common() works and a product nobody ordered counts 0. The starter uses a plain dict and even overwrites earlier orders of the same product. For orders.csv the result is coffee 4, tea 3, tea, green 2 and cake 1.
Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.
The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.
Start with units: Counter[str] = Counter() and change the return type to Counter[str].
Add instead of assigning: units[order.product] += order.qty. A Counter starts every new product at 0.
Counter is already imported from collections at the top of main.py.
One way to solve it. Yours can look different and still pass the checks.
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> Counter[str]:
"""Count the units ordered of each product."""
units: Counter[str] = Counter()
for order in orders:
units[order.product] += order.qty
return units
if __name__ == "__main__":
print(units_by_product(read_orders("orders.csv")))
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> dict[str, int]:
"""Count the units ordered of each product."""
units: dict[str, int] = {}
for order in orders:
units[order.product] = order.qty
return units
if __name__ == "__main__":
print(units_by_product(read_orders("orders.csv")))
test_main.py
from collections import Counter
from main import Order, read_orders, units_by_product
def test_given_file():
"""orders.csv gives coffee 4, tea 3, tea, green 2, cake 1"""
got = units_by_product(read_orders("orders.csv"))
assert got == {"coffee": 4, "tea": 3, "tea, green": 2, "cake": 1}, f"units_by_product gave {got!r}"
def test_adds_up():
"""Orders of the same product add up"""
got = units_by_product([Order("a", "jam", 2), Order("b", "jam", 5)])
jam = got["jam"]
assert jam == 7, f"jam has {jam!r} units, expected 2 + 5 = 7"
def test_is_counter():
"""The result is a Counter with most_common and zero for missing products"""
got = units_by_product([Order("a", "jam", 2), Order("b", "tea", 1), Order("c", "jam", 1)])
assert isinstance(got, Counter), f"units_by_product returned a {type(got).__name__}, expected a Counter"
top = got.most_common(1)
assert top == [("jam", 3)], f"most_common(1) is {top!r}, expected [('jam', 3)]"
assert got["cake"] == 0, "a product nobody ordered should count 0"
def test_generator_input():
"""A generator of orders works too"""
got = units_by_product(Order("a", "p", n) for n in range(4))
assert got == {"p": 6}, f"units_by_product gave {got!r}, expected p with 0 + 1 + 2 + 3 = 6"
orders.csv
customer,product,qty
ada,tea,2
bo,coffee,1
ada,coffee,3
cy,tea,1
bo,"tea, green",2
ada,cake,1
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 3 of 4
products_by_customer(orders) should map each customer to the set of different products they ordered. The starter replaces the set on every order, so only the last product survives. Use defaultdict(set), add each product, and return a plain dict with dict(...). For orders.csv, ada ordered tea, coffee and cake.
Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.
The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.
products: defaultdict[str, set[str]] = defaultdict(set) gives each new customer an empty set.
products[order.customer].add(order.product) adds to that set; a set ignores a product it already has.
End with return dict(products), so the caller gets a plain dict.
One way to solve it. Yours can look different and still pass the checks.
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> Counter[str]:
"""Count the units ordered of each product."""
units: Counter[str] = Counter()
for order in orders:
units[order.product] += order.qty
return units
def products_by_customer(orders: Iterable[Order]) -> dict[str, set[str]]:
"""Map each customer to the set of products they ordered."""
products: defaultdict[str, set[str]] = defaultdict(set)
for order in orders:
products[order.customer].add(order.product)
return dict(products)
if __name__ == "__main__":
print(products_by_customer(read_orders("orders.csv")))
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> Counter[str]:
"""Count the units ordered of each product."""
units: Counter[str] = Counter()
for order in orders:
units[order.product] += order.qty
return units
def products_by_customer(orders: Iterable[Order]) -> dict[str, set[str]]:
"""Map each customer to the set of products they ordered."""
products: dict[str, set[str]] = {}
for order in orders:
products[order.customer] = {order.product}
return products
if __name__ == "__main__":
print(products_by_customer(read_orders("orders.csv")))
test_main.py
from main import Order, products_by_customer, read_orders
def test_given_file():
"""orders.csv: ada ordered tea, coffee and cake"""
got = products_by_customer(read_orders("orders.csv"))
expected = {"ada": {"tea", "coffee", "cake"}, "bo": {"coffee", "tea, green"}, "cy": {"tea"}}
assert got == expected, f"products_by_customer gave {got!r}, expected {expected!r}"
def test_no_duplicates():
"""A product ordered twice appears once"""
got = products_by_customer([Order("a", "jam", 1), Order("a", "jam", 2)])
assert got == {"a": {"jam"}}, f"products_by_customer gave {got!r}, expected a with the set of jam"
def test_empty():
"""No orders give an empty dict"""
got = products_by_customer([])
assert got == {}, f"products_by_customer([]) gave {got!r}, expected {{}}"
def test_plain_dict():
"""The result is a plain dict"""
got = products_by_customer([Order("a", "jam", 1)])
assert type(got) is dict, f"products_by_customer returned a {type(got).__name__}: convert it with dict(...)"
orders.csv
customer,product,qty
ada,tea,2
bo,coffee,1
ada,coffee,3
cy,tea,1
bo,"tea, green",2
ada,cake,1
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 4 of 4
last_orders(orders, n) returns the last n orders. The starter turns all orders into a list, and for n = 0 it even returns every order, because [-0:] is the whole list. Use deque(orders, maxlen=n), which keeps only n orders in memory. The report at the bottom of main.py then prints the two best-selling products, the products ada ordered, and the last two orders.
Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.
The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.
deque(orders, maxlen=n) reads every order but keeps only the last n.
Return a list, as the type says: list(deque(orders, maxlen=n)).
deque with maxlen=0 keeps nothing, so n = 0 gives [] without a special case.
One way to solve it. Yours can look different and still pass the checks.
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> Counter[str]:
"""Count the units ordered of each product."""
units: Counter[str] = Counter()
for order in orders:
units[order.product] += order.qty
return units
def products_by_customer(orders: Iterable[Order]) -> dict[str, set[str]]:
"""Map each customer to the set of products they ordered."""
products: defaultdict[str, set[str]] = defaultdict(set)
for order in orders:
products[order.customer].add(order.product)
return dict(products)
def last_orders(orders: Iterable[Order], n: int) -> list[Order]:
"""Return the last n orders, reading the orders only once."""
return list(deque(orders, maxlen=n))
if __name__ == "__main__":
print(units_by_product(read_orders("orders.csv")).most_common(2))
print(sorted(products_by_customer(read_orders("orders.csv"))["ada"]))
for order in last_orders(read_orders("orders.csv"), 2):
print(order.customer, order.product, order.qty)
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import csv
from collections import Counter, defaultdict, deque, namedtuple
from collections.abc import Iterable, Iterator
Order = namedtuple("Order", ["customer", "product", "qty"])
def read_orders(path: str) -> Iterator[Order]:
"""Yield one Order per row of the CSV file at path, lazily."""
with open(path, encoding="utf-8", newline="") as f:
for row in csv.DictReader(f):
yield Order(row["customer"], row["product"], int(row["qty"]))
def units_by_product(orders: Iterable[Order]) -> Counter[str]:
"""Count the units ordered of each product."""
units: Counter[str] = Counter()
for order in orders:
units[order.product] += order.qty
return units
def products_by_customer(orders: Iterable[Order]) -> dict[str, set[str]]:
"""Map each customer to the set of products they ordered."""
products: defaultdict[str, set[str]] = defaultdict(set)
for order in orders:
products[order.customer].add(order.product)
return dict(products)
def last_orders(orders: Iterable[Order], n: int) -> list[Order]:
"""Return the last n orders, reading the orders only once."""
return list(orders)[-n:]
if __name__ == "__main__":
print(units_by_product(read_orders("orders.csv")).most_common(2))
print(sorted(products_by_customer(read_orders("orders.csv"))["ada"]))
for order in last_orders(read_orders("orders.csv"), 2):
print(order.customer, order.product, order.qty)
test_main.py
from learnrun import run_main
from main import Order, last_orders, read_orders
def test_last_two():
"""The last two orders of orders.csv"""
got = last_orders(read_orders("orders.csv"), 2)
assert got == [Order("bo", "tea, green", 2), Order("ada", "cake", 1)], f"last_orders(..., 2) gave {got!r}"
def test_zero():
"""n = 0 gives an empty list"""
got = last_orders(read_orders("orders.csv"), 0)
assert got == [], f"last_orders(..., 0) gave {got!r}, expected []"
def test_long_generator():
"""A long generator works"""
got = [order.qty for order in last_orders((Order("c", "p", n) for n in range(10000)), 3)]
assert got == [9997, 9998, 9999], f"the last three quantities are {got!r}, expected [9997, 9998, 9999]"
def test_report():
"""Running main.py prints the report"""
got = run_main().strip().splitlines()
expected = ["[('coffee', 4), ('tea', 3)]", "['cake', 'coffee', 'tea']", "bo tea, green 2", "ada cake 1"]
assert got == expected, f"main.py printed {got!r}, expected {expected!r}"
orders.csv
customer,product,qty
ada,tea,2
bo,coffee,1
ada,coffee,3
cy,tea,1
bo,"tea, green",2
ada,cake,1
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyimport csv
with open("data.csv", "w", encoding="utf-8", newline="") as f:
f.write("name,qty\ntea,2\n")
def read(path):
with open(path, encoding="utf-8", newline="") as f:
return csv.reader(f)
rows = read("data.csv")
print(next(rows))
What Python prints
ValueError: I/O operation on closed file.Why, and the fix
csv.reader is lazy: it reads the file only when asked for a row. return leaves the with block, which closes the file, so the first next() finds it closed. Make read a generator instead: for row in csv.reader(f): yield row keeps the file open until the rows are used up.
from collections import namedtuple
Order = namedtuple("Order", ["product", "qty"])
order = Order("tea", 2)
order.qty = 3
What Python prints
AttributeError: can't set attributeWhy, and the fix
A namedtuple is a tuple, and tuples cannot be changed. Make a changed copy instead: order = order._replace(qty=3). If records must change in place, use a dataclass from lesson I2.
counts = {}
for word in ["tea", "cake", "tea"]:
counts[word] += 1
print(counts)
What Python prints
KeyError: 'tea'Why, and the fix
counts[word] += 1 first reads counts[word], which does not exist for a new word. Use counts = Counter() from collections, which starts every key at 0, or count in one step: Counter(["tea", "cake", "tea"]).
Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source
5 questions, no hints. Score 80% or more to complete the lesson.
Finish every activity above to unlock the exit ticket.