Skip to content
aviral gupta

// A3.2 · ~30 min · Advanced

Processes with multiprocessing

After this lesson you can spread CPU-bound work over several cores with Process and Pool.map, guard the entry point, and pass only data a child process can receive.

Lesson 2 of 6 in A3 Concurrency

You will be able to

  • Run functions in child processes with Process and Pool.map, and read their results and exit codes
  • Guard the entry point with if __name__ == "__main__" and explain why spawn and forkserver need it
  • Pass picklable arguments and module-level functions, knowing that a child works on copies
  1. Warm-up · Activity 1 of 7

    Warm-up from the last lesson: on the default CPython build, which statements about threads are true? Pick all that apply.

    Select all that apply.

  2. Predict · Activity 2 of 7

    Predict before you read on. The child process adds 100 to a global. What does the parent print?

    from multiprocessing import Process
    
    counter = 0
    
    
    def bump() -> None:
        global counter
        counter += 100
    
    
    if __name__ == "__main__":
        p = Process(target=bump)
        p.start()
        p.join()
        print(counter, p.exitcode)
  3. Practice · Activity 3 of 7

    Complete the guard so that the pool is created only when the file runs as a program, not when a child process imports it.

    from multiprocessing import Pool
    
    
    def square(x: int) -> int:
        return x * x
    
    
    if __name__ == ____:
        with Pool(2) as pool:
            print(pool.map(square, [1, 2, 3]))
    if __name__ == :
  4. Practice · Activity 4 of 7

    One call in the pool raises ZeroDivisionError. What does this print?

    from multiprocessing import Pool
    
    
    def inverse(x: int) -> float:
        return 1 / x
    
    
    if __name__ == "__main__":
        with Pool(2) as pool:
            try:
                pool.map(inverse, [1, 0, 2])
            except ZeroDivisionError as error:
                print("caught:", error)
  5. Practice · Activity 5 of 7

    Match each piece of multiprocessing to what it gives you.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. The child appends to the list it was given. What does the parent print?

    from multiprocessing import Process
    
    
    def add(items: list[int]) -> None:
        items.append(4)
    
    
    if __name__ == "__main__":
        items = [1, 2, 3]
        p = Process(target=add, args=(items,))
        p.start()
        p.join()
        print(items)
  7. Apply · Activity 7 of 7

    Mini-task, on your own computer. Write squares.py with a module-level function sum_of_squares(n) that returns the sum of i * i for i below n. Under the __main__ guard, use a Pool to compute it for 10, 100 and 1000 with one map call, and print the list. Then remove the guard and run it again to see the RuntimeError.

    Check your work against this list

Build it yourself

Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.

Worked example

Counting primes in four processes

Counting primes by trial division is CPU-bound, so a pool of four worker processes can use four cores. The worker functions live in primes.py and main.py imports them: module-level functions pickle by name. main.py creates the pool and a Process of its own only under the __main__ guard. Save both files in one folder and run python main.py (python3 main.py on macOS and Linux) on your own computer.

main.py

import multiprocessing as mp

from primes import check_limit, count_primes

LIMITS = [10_000, 20_000, 30_000, 40_000]


def main() -> None:
    # Four worker processes; each call of count_primes runs in one of them.
    with mp.Pool(processes=4) as pool:
        counts = pool.map(count_primes, LIMITS)
    for limit, count in zip(LIMITS, counts):
        print(f"primes below {limit}: {count}")

    # One process of our own: it exits with status 2 for a bad limit.
    checker = mp.Process(target=check_limit, args=(1,))
    checker.start()
    checker.join()
    print("checker exit code:", checker.exitcode)


if __name__ == "__main__":
    main()

primes.py

import sys


def count_primes(limit: int) -> int:
    """Count the primes below limit, the slow way: CPU-bound work."""
    count = 0
    for n in range(2, limit):
        if all(n % d for d in range(2, int(n**0.5) + 1)):
            count += 1
    return count


def check_limit(limit: int) -> None:
    """End the process with exit status 2 when there can be no prime below limit."""
    if limit < 3:
        sys.exit(2)

Run it with

python main.py

Output

primes below 10000: 1229
primes below 20000: 2262
primes below 30000: 3245
primes below 40000: 4203
checker exit code: 2
  • pool.map returned the counts in the order of LIMITS, whichever worker finished first.
  • The with block closes the pool: its worker processes end when the block does.
  • sys.exit(2) in the child ended only the child; the parent read the status from exitcode.
  • Everything that starts processes sits in main(), which runs only under the guard.

Exercises

Exercise 1 of 2

Squares in a pool

check(n) returns n * n and whether it ran in a child process. Complete squares_in_pool(numbers, workers) so that it runs check on every number in a Pool of workers processes and returns the list of results, in order. Use a with statement so the pool is closed. Run the tests on your own computer: the browser cannot start processes.

This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.

Hints
  1. Hint 1

    with mp.Pool(processes=workers) as pool: creates the workers and closes them at the end of the block.

  2. Hint 2

    pool.map(check, numbers) calls check once per number in the workers and returns the results in order.

  3. Hint 3

    Pass the function itself, check, not check(n): the pool calls it for you.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

import multiprocessing as mp


def check(n: int) -> tuple[int, bool]:
    """Square n, and report whether this call ran in a child process."""
    return n * n, mp.parent_process() is not None


def squares_in_pool(numbers: list[int], workers: int = 2) -> list[tuple[int, bool]]:
    """Run check on every number in a pool of worker processes, in order."""
    with mp.Pool(processes=workers) as pool:
        return pool.map(check, numbers)


if __name__ == "__main__":
    print(squares_in_pool([1, 2, 3]))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

import multiprocessing as mp


def check(n: int) -> tuple[int, bool]:
    """Square n, and report whether this call ran in a child process."""
    return n * n, mp.parent_process() is not None


def squares_in_pool(numbers: list[int], workers: int = 2) -> list[tuple[int, bool]]:
    """Run check on every number in a pool of worker processes, in order."""
    # This runs check in the main process. Use mp.Pool(processes=workers) and its map.
    return [check(n) for n in numbers]


if __name__ == "__main__":
    print(squares_in_pool([1, 2, 3]))

test_main.py

from main import squares_in_pool


def test_squares_in_order():
    """The squares come back in the order of the numbers"""
    got = [square for square, _ in squares_in_pool([3, 1, 2])]
    assert got == [9, 1, 4], f"the squares are {got!r}, expected [9, 1, 4]"


def test_in_child_processes():
    """Every call ran in a worker process, not in the main process"""
    got = [in_child for _, in_child in squares_in_pool([1, 2, 3, 4])]
    assert got == [True, True, True, True], f"ran in a child: {got!r}; hand check to pool.map"


def test_workers_argument():
    """workers=3 works too"""
    got = squares_in_pool([5, 6], workers=3)
    assert got == [(25, True), (36, True)], f"squares_in_pool([5, 6], workers=3) returned {got!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 2 of 2

A worker that can be pickled

normalise_all strips and lower-cases words in a pool, but it fails with PicklingError: its worker clean is defined inside the function, and pickle sends functions by name. Move clean to the top level of main.py, keeping what it does, so that pool.map can send it to the workers. Run the tests on your own computer.

This exercise needs Python on your computer (the browser version cannot run it). The files and commands are below.

Hints
  1. Hint 1

    pickle stores a function as its module and qualified name; normalise_all.<locals>.clean cannot be looked up from outside.

  2. Hint 2

    Cut the def clean block out of normalise_all and paste it above, with no indentation.

  3. Hint 3

    normalise_all then keeps only the with mp.Pool(...) block that calls pool.map(clean, words).

Show a solution

One way to solve it. Yours can look different and still pass the checks.

import multiprocessing as mp


def clean(word: str) -> str:
    """Strip and lower-case one word. At module level, so a worker can import it."""
    return word.strip().lower()


def normalise_all(words: list[str], workers: int = 2) -> list[str]:
    """Strip and lower-case every word, in a pool of worker processes."""
    with mp.Pool(processes=workers) as pool:
        return pool.map(clean, words)


if __name__ == "__main__":
    print(normalise_all(["  Apple", "BANANA "]))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

import multiprocessing as mp


def normalise_all(words: list[str], workers: int = 2) -> list[str]:
    """Strip and lower-case every word, in a pool of worker processes."""

    def clean(word: str) -> str:
        return word.strip().lower()

    with mp.Pool(processes=workers) as pool:
        return pool.map(clean, words)


if __name__ == "__main__":
    print(normalise_all(["  Apple", "BANANA "]))

test_main.py

import pickle

import main


def test_normalise():
    """Words are stripped and lower-cased, in order"""
    got = main.normalise_all(["  Apple", "BANANA ", "Cherry"])
    assert got == ["apple", "banana", "cherry"], f"normalise_all returned {got!r}"


def test_clean_can_be_pickled():
    """clean is a module-level function that pickle can send to a worker"""
    clean = getattr(main, "clean", None)
    assert clean is not None, "main.py has no module-level function clean: move it out of normalise_all"
    assert pickle.loads(pickle.dumps(clean)) is clean, "clean could not be pickled by reference"
    assert clean("  Kiwi ") == "kiwi", f"clean('  Kiwi ') returned {clean('  Kiwi ')!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Common mistakes

Handing a lambda to the pool

from multiprocessing import Pool

if __name__ == "__main__":
    with Pool(2) as pool:
        print(pool.map(lambda x: x * 2, [1, 2, 3]))

What Python prints

_pickle.PicklingError: Can't pickle <function <lambda>

Why, and the fix

The pool pickles the function to send it to the workers, and pickle stores a function by its module and name. Every lambda is called <lambda>, so the worker could not find it. Write a def at the top level of the module, such as def double(x): return x * 2, and pass double.

Sending generators to worker processes

from multiprocessing import Pool

if __name__ == "__main__":
    rows = [(x * x for x in range(3)), (x + 1 for x in range(3))]
    with Pool(2) as pool:
        print(pool.map(sum, rows))

What Python prints

TypeError: cannot pickle 'generator' object

Why, and the fix

A generator holds a paused frame, which cannot be pickled. Everything you pass to a process must be picklable: numbers, strings, and lists, tuples and dicts of them. Build lists instead, [x * x for x in range(3)], or send the parameters (such as 3) and let the worker create the generator itself.

Using map for a function with two parameters

from multiprocessing import Pool

if __name__ == "__main__":
    with Pool(2) as pool:
        print(pool.map(pow, [(2, 3), (3, 2)]))

What Python prints

TypeError: pow() missing required argument 'exp' (pos 2)

Why, and the fix

pool.map calls the function with exactly one argument per item, so pow got the tuple (2, 3) as its base. The worker's exception is raised again in the parent. Use pool.starmap(pow, [(2, 3), (3, 2)]), which unpacks each tuple into the arguments and returns [8, 9].

Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

A process per task, a GIL per process

multiprocessing side-steps the GIL by using processes instead of threads. Each child is a separate Python interpreter with its own memory and its own GIL, so CPU-bound code really runs on several cores. Process(target=f, args=(x,)) works like Thread: start(), join(), and afterwards exitcode (0 for success, 1 after an uncaught exception). Pool(processes=4) keeps worker processes ready; pool.map(f, items) is a parallel map() that blocks until every result is ready and returns them in the order of items. starmap unpacks tuples into several arguments.

The __main__ guard

The default start method is spawn on Windows and macOS and, since 3.14, forkserver on the other POSIX systems. Both start a fresh interpreter that imports your main module, under the name __mp_main__, to find the target function. Code at the top level runs again in every child. So create processes and pools only under if __name__ == "__main__":; without the guard, the child would try to start processes itself, and multiprocessing stops it with a RuntimeError. Top-level definitions, constants and imports are fine.

Everything travels by pickle

The target, its arguments and the results are pickled and sent between processes. Functions are pickled by their qualified name, not their code, so a worker must be a def at the top level of a module: a lambda or a function nested in another function cannot be pickled. Generators, locks and open files cannot be pickled either. Because the child gets a copy, a list or dict it changes stays unchanged in the parent: send results back as return values from pool.map, or through a multiprocessing.Queue.

Sources

Last reviewed September 29, 2026