Warm-up · Activity 1 of 7
// A3.6 · ~50 min · Advanced
Choosing a concurrency model
After this lesson you decide between threads, processes and asyncio from the kind of work and from measurements, and you have built one file hasher in all three models.
Lesson 6 of 6 in A3 Concurrency
You will be able to
- Tell I/O-bound from CPU-bound work and name the model that fits each
- Move the same function between thread pools, process pools and asyncio
- Measure a concurrent version against the sequential one with time.perf_counter
Predict · Activity 2 of 7
Predict before you read on. On a standard build with the GIL, you run 8 downloads that each wait 0.05 s, and 4 pure-Python counting loops, each in a ThreadPoolExecutor. Which gets clearly faster than running sequentially?
Practice · Activity 3 of 7
Fill in the method that gives the digest as a string of hexadecimal digits.
import hashlib import io f = io.BytesIO(b"hello") print(hashlib.file_digest(f, "sha256").____())hashlib.file_digest(f, "sha256").()Practice · Activity 4 of 7
What does this print?
import hashlib data = b"x" * 10_000_000 print(len(hashlib.sha256(data).hexdigest()), hashlib.sha256(data).digest_size)Practice · Activity 5 of 7
Match each job to the model that usually fits it best on a standard build.
Brain teaser · Activity 6 of 7
Teaser. 100 calls of abs, once in a list comprehension and once in a process pool. What does this print?
import time from concurrent.futures import ProcessPoolExecutor def main(): start = time.perf_counter() sequential = [abs(n) for n in range(-50, 50)] sequential_time = time.perf_counter() - start start = time.perf_counter() with ProcessPoolExecutor(max_workers=2) as pool: pooled = list(pool.map(abs, range(-50, 50))) pooled_time = time.perf_counter() - start print(sequential == pooled, pooled_time > sequential_time) if __name__ == "__main__": main()Apply · Activity 7 of 7
Mini task, on your own machine. In a temporary folder, write 8 files of 2 MB of random bytes (os.urandom). Hash them with hashlib.file_digest once in a list comprehension and once with ThreadPoolExecutor.map, time both with time.perf_counter, and print the two times and whether the digests are equal. Run it a few times: which is faster on your machine?
Check your work against this list
Build it yourself
Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.
Worked example
Measuring threads on waiting and on computing
work.py holds two module-level functions: wait sleeps like an I/O call, and count loops in pure Python. main.py times the waits sequentially and in a thread pool, measures how many cores the counting keeps busy in a thread pool, and runs it in a process pool too. It prints comparisons instead of raw times, because times change from run to run. Save both files and run python main.py (python3 main.py on macOS and Linux) on your own machine; add prints of the times to see your own numbers.
main.py
import sys
import time
from collections.abc import Callable
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
from typing import Any
from work import count, wait
def timed(run: Callable[[], list[Any]]) -> tuple[list[Any], float]:
"""Run once and return (results, seconds); perf_counter is made for short durations."""
start = time.perf_counter()
results = run()
return results, time.perf_counter() - start
def main() -> None:
waits = [0.05] * 8
_, sequential = timed(lambda: list(map(wait, waits)))
with ThreadPoolExecutor(max_workers=8) as pool:
_, threaded = timed(lambda: list(pool.map(wait, waits)))
print("I/O-bound, 8 waits of 0.05 s:")
print(" threads at least 3x faster than sequential:", sequential / threaded >= 3)
sizes = [1_000_000] * 4
seq_totals = list(map(count, sizes))
with ThreadPoolExecutor(max_workers=4) as pool:
cpu_start = time.process_time() # CPU seconds used by all threads of this process
thread_totals, wall = timed(lambda: list(pool.map(count, sizes)))
cores_busy = (time.process_time() - cpu_start) / wall
with ProcessPoolExecutor(max_workers=4) as pool:
process_totals = list(pool.map(count, sizes))
print(f"CPU-bound, 4 counts in pure Python (GIL enabled: {sys._is_gil_enabled()}):")
print(" threads kept more than 1.5 cores busy:", cores_busy > 1.5)
print(" all three gave the same totals:", seq_totals == thread_totals == process_totals)
if __name__ == "__main__":
main()
work.py
import time
def wait(seconds: float) -> float:
"""I/O-bound stand-in: the thread just waits, like a download or a disk read."""
time.sleep(seconds)
return seconds
def count(n: int) -> int:
"""CPU-bound: pure Python arithmetic, holding the GIL the whole time."""
total = 0
for i in range(n):
total += i % 7
return total
Run it with
python main.pyOutput
I/O-bound, 8 waits of 0.05 s:
threads at least 3x faster than sequential: True
CPU-bound, 4 counts in pure Python (GIL enabled: True):
threads kept more than 1.5 cores busy: False
all three gave the same totals: True- Eight waits of 0.05 s overlapped in eight threads: about 0.05 s instead of 0.4 s.
- The counting loops held the GIL, so four threads kept about one core busy: time.process_time() rose by about one CPU second per wall-clock second.
- The process pool ran the same count function, imported from work.py, and gave the same totals; on several cores it is the one that can be faster.
- On a free-threaded build, GIL enabled would print False, and the threads line could change: measure on the build you deploy.
Exercises
Exercise 1 of 4
Build step 1: hash one file, then all
Module build: a file hasher in three models. Start with the plain version. Write hash_file(path), which opens the file in binary mode and returns the SHA-256 hex digest from hashlib.file_digest. Then write hash_sequential(paths), which returns a dict from each path to its digest, in the order of paths. Run the tests on your own machine.
This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.
Hints
Hint 1
with open(path, "rb") as f: opens the file in binary mode, which file_digest requires.
Hint 2
hashlib.file_digest(f, "sha256") returns a hash object; .hexdigest() gives the string.
Hint 3
hash_sequential is one dict comprehension: {path: hash_file(path) for path in paths}.
Show a solution
One way to solve it. Yours can look different and still pass the checks.
import hashlib
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
Run it on your computer
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import hashlib
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
# Open the file in binary mode and pass it to hashlib.file_digest with "sha256".
return ""
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {}
test_main.py
import hashlib
import os
import tempfile
from main import hash_file, hash_sequential
def make_files(folder, contents):
paths = []
for number, data in enumerate(contents):
path = os.path.join(folder, f"file{number}.bin")
with open(path, "wb") as f:
f.write(data)
paths.append(path)
return paths
def test_hash_file():
"""hash_file gives the SHA-256 hex digest of the file's bytes"""
with tempfile.TemporaryDirectory() as folder:
[path] = make_files(folder, [b"hello"])
got = hash_file(path)
assert got == hashlib.sha256(b"hello").hexdigest(), f"hash_file gave {got!r}"
def test_large_file():
"""A file of 1 MB hashes to the digest of its bytes"""
data = bytes(range(256)) * 4096
with tempfile.TemporaryDirectory() as folder:
[path] = make_files(folder, [data])
got = hash_file(path)
assert got == hashlib.sha256(data).hexdigest(), f"hash_file gave {got!r}"
def test_hash_sequential():
"""hash_sequential maps every path to its digest, in the order of paths"""
with tempfile.TemporaryDirectory() as folder:
paths = make_files(folder, [b"b", b"a"])
got = hash_sequential(paths)
expected = [(paths[0], hashlib.sha256(b"b").hexdigest()), (paths[1], hashlib.sha256(b"a").hexdigest())]
assert list(got.items()) == expected, f"hash_sequential gave {got!r}"
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 2 of 4
Build step 2: threads
hash_threads(paths, workers) still hashes in the main thread. Hash the files in a ThreadPoolExecutor with workers threads, calling hash_file for each path, and return the same dict as hash_sequential. Reading waits on the disk and hashlib releases the GIL on large data, so threads can help here. Run the tests on your own machine.
This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.
Hints
Hint 1
Open the pool in a with block: with ThreadPoolExecutor(max_workers=workers) as pool:.
Hint 2
pool.map(hash_file, paths) returns the digests in the order of paths.
Hint 3
dict(zip(paths, digests)) pairs each path with its digest.
Show a solution
One way to solve it. Yours can look different and still pass the checks.
import hashlib
from concurrent.futures import ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
with ThreadPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
Run it on your computer
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import hashlib
from concurrent.futures import ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
# Still one after another in the main thread. Use a ThreadPoolExecutor.
return hash_sequential(paths)
test_main.py
import threading
import tempfile
import os
import main
def make_files(folder, contents):
paths = []
for number, data in enumerate(contents):
path = os.path.join(folder, f"file{number}.bin")
with open(path, "wb") as f:
f.write(data)
paths.append(path)
return paths
def test_same_digests():
"""hash_threads gives the same dict as hash_sequential"""
with tempfile.TemporaryDirectory() as folder:
paths = make_files(folder, [b"one", b"two", b"three"])
got = main.hash_threads(paths, workers=2)
expected = main.hash_sequential(paths)
assert list(got.items()) == list(expected.items()), f"hash_threads gave {got!r}"
def test_runs_in_pool_threads():
"""hash_threads calls hash_file in worker threads, not in the main thread"""
seen = []
def spy(path):
seen.append(threading.current_thread().name)
return path.upper()
real = main.hash_file
main.hash_file = spy
try:
got = main.hash_threads(["a", "b"], workers=2)
finally:
main.hash_file = real
assert got == {"a": "A", "b": "B"}, f"hash_threads gave {got!r}"
assert "MainThread" not in seen, f"hash_file ran in {seen!r}: use a ThreadPoolExecutor"
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 3 of 4
Build step 3: processes
hash_processes(paths, workers) still uses threads. Use the ProcessPoolExecutor imported at the top instead, with max_workers=workers, and return the same dict. hash_file is a module-level function, so the workers can unpickle it, and the __main__ guard at the bottom keeps them from rerunning the script. Run the tests on your own machine.
This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.
Hints
Hint 1
The body looks like hash_threads with one name changed.
Hint 2
with ProcessPoolExecutor(max_workers=workers) as pool: then dict(zip(paths, pool.map(hash_file, paths))).
Hint 3
Do not pass a lambda to pool.map: a process pool can only send module-level functions.
Show a solution
One way to solve it. Yours can look different and still pass the checks.
import hashlib
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
with ThreadPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
def hash_processes(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a process pool; hash_file is a module-level function, so it pickles."""
with ProcessPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
if __name__ == "__main__":
import sys
print(hash_processes(sys.argv[1:]))
Run it on your computer
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import hashlib
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
with ThreadPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
def hash_processes(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a process pool; hash_file is a module-level function, so it pickles."""
# Still threads. Use the ProcessPoolExecutor imported above.
return hash_threads(paths, workers)
if __name__ == "__main__":
import sys
print(hash_processes(sys.argv[1:]))
test_main.py
import tempfile
import os
import main
def make_files(folder, contents):
paths = []
for number, data in enumerate(contents):
path = os.path.join(folder, f"file{number}.bin")
with open(path, "wb") as f:
f.write(data)
paths.append(path)
return paths
def test_uses_process_pool():
"""hash_processes hashes in a ProcessPoolExecutor and gives the same dict as hash_sequential"""
used = []
real = main.ProcessPoolExecutor
class Spy(real):
def __init__(self, *args, **kwargs):
used.append(kwargs.get("max_workers", args[0] if args else None))
super().__init__(*args, **kwargs)
main.ProcessPoolExecutor = Spy
try:
with tempfile.TemporaryDirectory() as folder:
paths = make_files(folder, [b"one", b"two", b"three"])
got = main.hash_processes(paths, workers=2)
expected = main.hash_sequential(paths)
finally:
main.ProcessPoolExecutor = real
assert used == [2], f"ProcessPoolExecutor was created with {used!r}: create one with max_workers=workers"
assert list(got.items()) == list(expected.items()), f"hash_processes gave {got!r}"
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyExercise 4 of 4
Build step 4: asyncio
The last model: hash_async(paths) is a coroutine, but it calls hash_file directly, which blocks the event loop. Run each hash_file(path) with asyncio.to_thread, as a task in an asyncio.TaskGroup, and return the dict in the order of paths. Then time all four functions on your own files with time.perf_counter and compare. Run the tests on your own machine.
This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.
Hints
Hint 1
asyncio.to_thread(hash_file, path) is a coroutine that runs hash_file in a separate thread.
Hint 2
Inside async with asyncio.TaskGroup() as tg:, keep {path: tg.create_task(asyncio.to_thread(hash_file, path)) for path in paths}.
Hint 3
After the block, build {path: task.result() for path, task in tasks.items()}.
Show a solution
One way to solve it. Yours can look different and still pass the checks.
import asyncio
import hashlib
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
with ThreadPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
def hash_processes(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a process pool; hash_file is a module-level function, so it pickles."""
with ProcessPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
async def hash_async(paths: list[str]) -> dict[str, str]:
"""Hash the files from asyncio: each blocking hash_file call runs in a thread via to_thread."""
async with asyncio.TaskGroup() as tg:
tasks = {path: tg.create_task(asyncio.to_thread(hash_file, path)) for path in paths}
return {path: task.result() for path, task in tasks.items()}
if __name__ == "__main__":
import sys
print(hash_processes(sys.argv[1:]))
Run it on your computer
Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.
main.py
import asyncio
import hashlib
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def hash_file(path: str) -> str:
"""Return the SHA-256 hex digest of the file at path."""
with open(path, "rb") as f:
return hashlib.file_digest(f, "sha256").hexdigest()
def hash_sequential(paths: list[str]) -> dict[str, str]:
"""Hash the files one after another: {path: digest}, in the order of paths."""
return {path: hash_file(path) for path in paths}
def hash_threads(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a thread pool: file reads wait on the disk, and hashlib releases the GIL."""
with ThreadPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
def hash_processes(paths: list[str], workers: int = 4) -> dict[str, str]:
"""Hash the files in a process pool; hash_file is a module-level function, so it pickles."""
with ProcessPoolExecutor(max_workers=workers) as pool:
return dict(zip(paths, pool.map(hash_file, paths)))
async def hash_async(paths: list[str]) -> dict[str, str]:
"""Hash the files from asyncio: each blocking hash_file call runs in a thread via to_thread."""
# Calling hash_file directly blocks the event loop.
# Run each call with asyncio.to_thread, as a task in an asyncio.TaskGroup.
return hash_sequential(paths)
if __name__ == "__main__":
import sys
print(hash_processes(sys.argv[1:]))
test_main.py
import asyncio
import tempfile
import threading
import os
import main
def make_files(folder, contents):
paths = []
for number, data in enumerate(contents):
path = os.path.join(folder, f"file{number}.bin")
with open(path, "wb") as f:
f.write(data)
paths.append(path)
return paths
def test_same_digests():
"""hash_async gives the same dict as hash_sequential, in the order of paths"""
with tempfile.TemporaryDirectory() as folder:
paths = make_files(folder, [b"one", b"two", b"three"])
got = asyncio.run(main.hash_async(paths))
expected = main.hash_sequential(paths)
assert list(got.items()) == list(expected.items()), f"hash_async gave {got!r}"
def test_event_loop_not_blocked():
"""hash_async runs hash_file in worker threads, off the event loop"""
seen = []
def spy(path):
seen.append(threading.current_thread().name)
return path.upper()
real = main.hash_file
main.hash_file = spy
try:
got = asyncio.run(main.hash_async(["a", "b"]))
finally:
main.hash_file = real
assert got == {"a": "A", "b": "B"}, f"hash_async gave {got!r}"
assert "MainThread" not in seen, f"hash_file ran in {seen!r}: use asyncio.to_thread"
On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.
Run the program:
python main.pyRun the checks (needs learnrun.py in the same folder):
python learnrun.py testDownload learnrun.pyCommon mistakes
A lambda in a process pool
import hashlib
from concurrent.futures import ProcessPoolExecutor
texts = ["alpha", "beta"]
if __name__ == "__main__":
with ProcessPoolExecutor(max_workers=2) as pool:
digests = list(pool.map(lambda text: hashlib.sha256(text.encode()).hexdigest(), texts))
print(digests)
What Python prints
_pickle.PicklingError: Can't pickle <function <lambda>Why, and the fix
The thread version worked with a lambda, but a process pool pickles the function by its name, and a lambda has none a worker could look up. Define the work as a module-level function, such as def hash_text(text: str) -> str:, and pass that to pool.map.
Calling the function inside asyncio.to_thread
import asyncio
import hashlib
def hash_text(text: str) -> str:
return hashlib.sha256(text.encode()).hexdigest()
async def main() -> None:
digest = await asyncio.to_thread(hash_text("alpha"))
print(digest)
asyncio.run(main())
What Python prints
TypeError: 'str' object is not callableWhy, and the fix
hash_text("alpha") runs at once, in the event loop's thread, and to_thread receives its result, a string. Pass the function and its arguments separately: await asyncio.to_thread(hash_text, "alpha"). mypy reports this line before you run it.
Hashing a file opened in text mode
import hashlib
with open("notes.txt", "w", encoding="utf-8") as f:
f.write("hello")
with open("notes.txt") as f:
print(hashlib.file_digest(f, "sha256").hexdigest())
What Python prints
ValueError: '<_io.TextIOWrapper name='notes.txt' mode='r' encoding='utf-8'>' is not a file-like object in binary reading mode.Why, and the fix
file_digest hashes bytes, so the file must be opened for reading in binary mode: open("notes.txt", "rb"). Text mode would also decode and change line endings, which changes the digest.
Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source
Exit ticket
5 questions, no hints. Score 80% or more to complete the lesson.
Finish every activity above to unlock the exit ticket.