Skip to content
aviral gupta

// I3.4 · ~35 min · Intermediate

itertools

After this lesson you can join, slice and chunk any iterable with chain, islice and batched, group sorted data with groupby, and build running totals and combinations with accumulate and product.

Lesson 4 of 6 in I3 Iteration and functional tools

You will be able to

  • Join, slice and chunk iterables with chain, islice and batched
  • Group sorted data with groupby, and explain why the input must be sorted by the same key
  • Build running results with accumulate and all combinations with product
  1. Warm-up · Activity 1 of 7

    Warm-up from lesson I3.3: what does this print?

    g = (n * 2 for n in [1, 2, 3])
    print(next(g), sum(g))
  2. Predict · Activity 2 of 7

    Predict before you read on: groupby on unsorted letters. What does this print?

    from itertools import groupby
    
    for key, group in groupby("AABAA"):
        print(key, len(list(group)), end="; ")
  3. Practice · Activity 3 of 7

    naturals() never ends. Fill in the itertools function that takes its first three values.

    from itertools import islice
    
    def naturals():
        n = 1
        while True:
            yield n
            n += 1
    
    print(list(____(naturals(), 3)))
    print(list((naturals(), 3)))
  4. Practice · Activity 4 of 7

    chain takes any iterables. What does this print?

    from itertools import chain
    
    print(list(chain([1, 2], "ab", range(2))))
  5. Practice · Activity 5 of 7

    Match each call to the list it gives.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. The groups are stored in a dict and read after the loop. What does this print?

    from itertools import groupby
    
    groups = {}
    for key, group in groupby("aabbb"):
        groups[key] = group
    print({k: list(g) for k, g in groups.items()})
  7. Apply · Activity 7 of 7

    Mini-task. sales = [("north", 120), ("south", 80), ("north", 50), ("east", 70), ("south", 30)]. Print the total per region, one line each in alphabetical order, with sorted and groupby. Then print the running total of all sales, in the order given, with accumulate.

    Check your work against this list

Build it yourself

Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.

Worked example

A café’s shift report

Orders from two shifts are joined with chain, and islice peeks at the first two. sorted and groupby list who ordered each drink, with one key function for both. accumulate gives the running number of cups, batched seats guests two to a table, and product builds the menu from sizes and drinks.

main.py

from itertools import accumulate, batched, chain, groupby, islice, product

morning = [("Ada", "tea"), ("Ben", "coffee"), ("Cy", "tea")]
evening = [("Dee", "coffee"), ("Eve", "tea")]

orders = list(chain(morning, evening))
print("first two:", list(islice(orders, 2)))


def drink(order: tuple[str, str]) -> str:
    return order[1]


for name, group in groupby(sorted(orders, key=drink), key=drink):
    print(name, [who for who, _ in group])

cups = [3, 5, 2, 4]
print("running total:", list(accumulate(cups)))
print("tables:", list(batched(["T1", "T2", "T3", "T4", "T5"], 2)))
sizes = ["small", "large"]
print("menu:", [f"{s} {d}" for s, d in product(sizes, ["tea", "coffee"])])

Run it with

python main.py

Output

first two: [('Ada', 'tea'), ('Ben', 'coffee')]
coffee ['Ben', 'Dee']
tea ['Ada', 'Cy', 'Eve']
running total: [3, 8, 10, 14]
tables: [('T1', 'T2'), ('T3', 'T4'), ('T5',)]
menu: ['small tea', 'small coffee', 'large tea', 'large coffee']
  • Sorting by drink put both coffees and all three teas next to each other, so each drink got one group.
  • The list comprehension read each group inside the loop, while it was still visible.
  • The last table has one guest: batched makes the final tuple shorter instead of padding it.
  • product changed the drink fastest, like an inner for loop: small tea, small coffee, then large.
Change it and run it

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Exercises

Exercise 1 of 3

Words by first letter

group_by_first_letter(words) returns a dict that maps each first letter to the words starting with it, each list in alphabetical order. The starter uses groupby without sorting, so it only works when the input is already sorted. Fix it so that ["bat", "apple", "bee", "axe"] gives {"a": ["apple", "axe"], "b": ["bat", "bee"]}.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    groupby only joins neighbours, so equal first letters must sit next to each other first.

  2. Hint 2

    Sorting the words alphabetically also sorts them by first letter, which is the groupby key.

  3. Hint 3

    Write groupby(sorted(words), key=lambda w: w[0]).

Show a solution

One way to solve it. Yours can look different and still pass the checks.

from itertools import groupby


def group_by_first_letter(words: list[str]) -> dict[str, list[str]]:
    """Map each first letter to the words that start with it."""
    result: dict[str, list[str]] = {}
    for letter, group in groupby(sorted(words), key=lambda w: w[0]):
        result[letter] = list(group)
    return result


if __name__ == "__main__":
    print(group_by_first_letter(["bat", "apple", "bee", "axe"]))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

from itertools import groupby


def group_by_first_letter(words: list[str]) -> dict[str, list[str]]:
    """Map each first letter to the words that start with it."""
    result: dict[str, list[str]] = {}
    for letter, group in groupby(words, key=lambda w: w[0]):
        result[letter] = list(group)
    return result


if __name__ == "__main__":
    print(group_by_first_letter(["bat", "apple", "bee", "axe"]))

test_main.py

from main import group_by_first_letter


def test_sorted_input():
    """Sorted input is grouped"""
    got = group_by_first_letter(["apple", "axe", "bat"])
    assert got == {"a": ["apple", "axe"], "b": ["bat"]}, f"got {got!r}"


def test_unsorted_input():
    """Unsorted input gives one entry per letter, with every word"""
    got = group_by_first_letter(["bat", "apple", "bee", "axe"])
    expected = {"a": ["apple", "axe"], "b": ["bat", "bee"]}
    assert got == expected, f"for ['bat', 'apple', 'bee', 'axe'] got {got!r}, expected {expected!r}"


def test_empty():
    """No words give an empty dict"""
    got = group_by_first_letter([])
    assert got == {}, f"for [] got {got!r}, expected {{}}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 2 of 3

Balances and peaks with accumulate

Write two functions with accumulate. balances(start, changes) returns the account balance before and after each change: balances(100, [20, -50, 10]) is [100, 120, 70, 80]. peaks(values) returns the highest value seen so far at each position: peaks([3, 1, 4, 1, 5]) is [3, 3, 4, 4, 5]. Both return lists.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    accumulate(changes) alone gives running totals of the changes, without the start balance.

  2. Hint 2

    The keyword argument initial=start puts the start balance in front and adds the changes to it.

  3. Hint 3

    For peaks, pass max as the second argument: accumulate(values, max) keeps the larger of the result so far and the next value.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

from itertools import accumulate


def balances(start: int, changes: list[int]) -> list[int]:
    """Return the balance before and after each change."""
    return list(accumulate(changes, initial=start))


def peaks(values: list[int]) -> list[int]:
    """Return the highest value seen so far, at each position."""
    return list(accumulate(values, max))


if __name__ == "__main__":
    print(balances(100, [20, -50, 10]))
    print(peaks([3, 1, 4, 1, 5]))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

from itertools import accumulate


def balances(start: int, changes: list[int]) -> list[int]:
    """Return the balance before and after each change."""
    return [start]


def peaks(values: list[int]) -> list[int]:
    """Return the highest value seen so far, at each position."""
    return list(values)


if __name__ == "__main__":
    print(balances(100, [20, -50, 10]))
    print(peaks([3, 1, 4, 1, 5]))

test_main.py

from main import balances, peaks


def test_balances():
    """100 with changes 20, -50 and 10 gives 100, 120, 70, 80"""
    got = balances(100, [20, -50, 10])
    assert got == [100, 120, 70, 80], f"balances(100, [20, -50, 10]) returned {got!r}"


def test_no_changes():
    """Without changes only the start balance is left"""
    got = balances(5, [])
    assert got == [5], f"balances(5, []) returned {got!r}, expected [5]"


def test_peaks():
    """peaks([3, 1, 4, 1, 5]) is [3, 3, 4, 4, 5]"""
    got = peaks([3, 1, 4, 1, 5])
    assert got == [3, 3, 4, 4, 5], f"peaks([3, 1, 4, 1, 5]) returned {got!r}"


def test_peaks_empty():
    """peaks([]) is []"""
    got = peaks([])
    assert got == [], f"peaks([]) returned {got!r}, expected []"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 3 of 3

One page of results

page(items, size, number) returns page number (counting from 1) of items as a list, with size items per page, or [] if there is no such page. page("abcdefg", 3, 3) is ["g"]. It must work for any iterable, including a generator. Use batched to make the pages and islice to skip to the right one.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    batched(items, size) gives the pages as tuples, one after the other.

  2. Hint 2

    islice(pages, number - 1, None) skips the pages before the one you want.

  3. Hint 3

    next(..., ()) takes that page, or an empty tuple when there is none. Return it as a list.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

from collections.abc import Iterable
from itertools import batched, islice


def page(items: Iterable[str], size: int, number: int) -> list[str]:
    """Return page number (counting from 1), size items per page."""
    batch = next(islice(batched(items, size), number - 1, None), ())
    return list(batch)


if __name__ == "__main__":
    print(page("abcdefg", 3, 2))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

from collections.abc import Iterable
from itertools import batched, islice


def page(items: Iterable[str], size: int, number: int) -> list[str]:
    """Return page number (counting from 1), size items per page."""
    return list(islice(items, size))


if __name__ == "__main__":
    print(page("abcdefg", 3, 2))

test_main.py

from main import page


def test_first_page():
    """Page 1 holds the first three letters"""
    got = page("abcdefg", 3, 1)
    assert got == ["a", "b", "c"], f"page('abcdefg', 3, 1) returned {got!r}"


def test_second_page():
    """Page 2 holds the next three"""
    got = page("abcdefg", 3, 2)
    assert got == ["d", "e", "f"], f"page('abcdefg', 3, 2) returned {got!r}, expected ['d', 'e', 'f']"


def test_short_last_page():
    """The last page may be shorter"""
    got = page("abcdefg", 3, 3)
    assert got == ["g"], f"page('abcdefg', 3, 3) returned {got!r}, expected ['g']"


def test_no_such_page():
    """A page past the end is empty"""
    got = page("abcdefg", 3, 4)
    assert got == [], f"page('abcdefg', 3, 4) returned {got!r}, expected []"


def test_generator():
    """A generator works too"""
    got = page((str(n) for n in range(100)), 10, 2)
    expected = [str(n) for n in range(10, 20)]
    assert got == expected, f"page 2 of a generator of 100 numbers was {got!r}, expected {expected!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Common mistakes

A negative index in islice

from itertools import islice

scores = [70, 85, 90, 65, 88]
print(list(islice(scores, -2, None)))

What Python prints

ValueError: Indices for islice() must be None or an integer: 0 <= x <= sys.maxsize.

Why, and the fix

islice works on iterators, which do not know their length, so counting from the end is impossible. On a list, use ordinary slicing: scores[-2:]. On a stream, deque(it, maxlen=2) keeps the last items; you will meet it in the collections lesson later in this module.

Indexing the result of chain

from itertools import chain

names = chain(["Ada", "Ben"], ["Cy"])
print(names[0])

What Python prints

TypeError: 'itertools.chain' object is not subscriptable

Why, and the fix

chain returns an iterator, not a list, so it has no positions. Take items with next(names), or build a list first when you need indexes: names = list(chain(...)). The same goes for every itertools function in this lesson.

An incomplete batch with strict=True

from itertools import batched

for pair in batched("abcde", 2, strict=True):
    print(pair)

What Python prints

ValueError: batched(): incomplete batch

Why, and the fix

strict=True asks batched to refuse a short last batch, and five items do not split into pairs. The first two pairs are printed before the error. Leave strict out if a shorter last batch is fine, or check that the length divides evenly.

Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

Joining, cutting and chunking

chain(a, b) runs through a, then b, as one iterator, without building a new list. islice(it, stop) or islice(it, start, stop, step) slices any iterable, even a generator or an endless one, which [ : ] cannot do; negative values are not allowed. batched(it, n) groups the items into tuples of n, and the last tuple may be shorter. Every itertools function in this lesson returns an iterator: lazy, good for one pass, and a list only when you call list() on it.

groupby needs sorted input

groupby(iterable, key) yields (key, group) pairs for runs of neighbouring items with the same key. It does not collect equal keys from the whole input: in "AABAA" the A’s form two groups. So sort by the same key first: groupby(sorted(data, key=f), key=f). Each group is an iterator that shares the input with groupby. Once groupby moves to the next group, the previous one is gone, so turn a group into a list, or sum it, straight away.

Running results and combinations

accumulate(it) yields running totals: 1, 3, 6 for [1, 2, 3]. With a function of two arguments as the second argument it keeps another running result: accumulate(prices, max) gives the highest price so far. initial=100 puts a start value in front. product(a, b) yields every pair (x, y), in the order two nested for loops would, with the right-hand item changing fastest. product(a, repeat=2) pairs a with itself.

Sources

Last reviewed September 29, 2026