Skip to content
aviral gupta

// I5.3 · ~30 min · Intermediate

doctest and what to test

After this lesson you can write docstring examples that doctest checks, choose edge cases such as empty input and boundaries, and split tests so that each one checks a single behaviour.

Lesson 3 of 5 in I5 Testing and project tooling

You will be able to

  • Write docstring examples, including an expected exception, and run them with doctest
  • Pick edge cases that find bugs: empty input, single items and both sides of a boundary
  • Write focused tests that each check one behaviour and say which one in their name
  1. Warm-up · Activity 1 of 7

    Warm-up from the previous lesson: which of these create a mock that raises ValueError when it is called? Pick all that apply.

    Select all that apply.

  2. Predict · Activity 2 of 7

    Predict before you read on: this example fails. What does doctest show under Got:?

    import doctest
    
    
    def greet(name):
        """
        >>> greet("Ada")
        Hello, Ada!
        """
        return f"Hello, {name}!"
    
    
    doctest.testmod()
  3. Practice · Activity 3 of 7

    Fill in the doctest function that checks every docstring of the module being run and returns the counts.

    print(doctest.____())
    print(doctest.())
  4. Practice · Activity 4 of 7

    The traceback example leaves out the stack. What is the last line printed?

    import doctest
    
    
    def parse_age(text):
        """
        >>> parse_age("42")
        42
        >>> parse_age("old")
        Traceback (most recent call last):
        ValueError: not a number: 'old'
        """
        if not text.isdigit():
            raise ValueError(f"not a number: {text!r}")
        return int(text)
    
    
    print(doctest.testmod())
  5. Practice · Activity 5 of 7

    You are testing word_count(text), which returns len(text.split()). Match each input to the edge case it covers.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. Two functions, three examples. What is the last line printed?

    import doctest
    
    
    def add(a, b):
        """
        >>> add(1, 2)
        3
        >>> add(0.1, 0.2)
        0.3
        """
        return a + b
    
    
    def half(n):
        """
        >>> half(4)
        2.0
        """
        return n / 2
    
    
    print(doctest.testmod())
  7. Apply · Activity 7 of 7

    Mini-task. clamp(value, low, high) limits value to the range low..high and raises ValueError when low > high. Write its docstring examples: a value inside, one below, one above, both boundaries exactly, and the error. On your machine, end the file with print(doctest.testmod()) and run it.

    Check your work against this list

Build it yourself

Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.

Worked example

Splitting a bill

split_bill shares a bill in cents as evenly as possible. Its docstring shows the typical use and the error. The unit tests go after the edges: a total that does not divide, zero cents, more people than cents and zero people, one behaviour per test. The program runs the function’s doctests with DocTestFinder and DocTestRunner, which testmod() uses too, then the unit tests, and prints both results.

main.py

import doctest
import unittest


def split_bill(total_cents: int, people: int) -> list[int]:
    """Split a bill in cents as evenly as possible.

    >>> split_bill(1000, 4)
    [250, 250, 250, 250]
    >>> split_bill(1000, 3)
    [334, 333, 333]
    >>> split_bill(5, 1)
    [5]
    >>> split_bill(1000, 0)
    Traceback (most recent call last):
    ...
    ValueError: people must be at least 1
    """
    if people < 1:
        raise ValueError("people must be at least 1")
    share, rest = divmod(total_cents, people)
    return [share + 1 if i < rest else share for i in range(people)]


class TestSplitBill(unittest.TestCase):
    def test_shares_add_up_to_total(self) -> None:
        self.assertEqual(sum(split_bill(1000, 7)), 1000)

    def test_shares_differ_by_at_most_one_cent(self) -> None:
        shares = split_bill(1000, 7)
        self.assertTrue(max(shares) - min(shares) <= 1)

    def test_zero_total(self) -> None:
        self.assertEqual(split_bill(0, 3), [0, 0, 0])

    def test_more_people_than_cents(self) -> None:
        self.assertEqual(split_bill(2, 3), [1, 1, 0])

    def test_zero_people_raises(self) -> None:
        with self.assertRaises(ValueError):
            split_bill(1000, 0)


# doctest.testmod() would check every docstring of a module run as a script;
# here one function's examples are found and run, and the counts printed.
finder = doctest.DocTestFinder()
runner = doctest.DocTestRunner()
for test in finder.find(split_bill, "split_bill", globs=globals()):
    runner.run(test)
print("doctests:", runner.summarize(verbose=False))

result = unittest.TestResult()
unittest.TestLoader().loadTestsFromTestCase(TestSplitBill).run(result)
print("unit tests run:", result.testsRun)
print("all passed:", result.wasSuccessful())

Run it with

python main.py

Output

doctests: TestResults(failed=0, attempted=4)
unit tests run: 5
all passed: True
  • The four doctest examples double as documentation: a reader sees at once what split_bill returns.
  • The unit tests check properties, such as "adds up to the total", which hold for any input.
  • Each test name says which behaviour broke if it fails, for example test_more_people_than_cents.
  • Change i < rest to i <= rest and run it again: the doctests and several focused tests fail.
Change it and run it

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Exercises

Exercise 1 of 2

Examples for median

median(values) works. Its docstring has one example; add two more: an even number of values, such as [4, 1, 3, 2], where the result is the mean of the two middle values, and an empty list, which raises ValueError: median of an empty list. The checks run your examples against the real median and against two broken versions, which they must catch.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    An example is a line >>> median([4, 1, 3, 2]) and, on the next line at the same indentation, the exact output: 2.5.

  2. Hint 2

    For the error, the expected output is three lines: Traceback (most recent call last):, then ..., then ValueError: median of an empty list.

  3. Hint 3

    On your own machine, doctest.testmod() at the end of the file runs the same examples.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

def median(values: list[float]) -> float:
    """Return the middle value; for an even count, the mean of the two middle values.

    >>> median([3, 1, 2])
    2
    >>> median([4, 1, 3, 2])
    2.5
    >>> median([])
    Traceback (most recent call last):
    ...
    ValueError: median of an empty list
    """
    if not values:
        raise ValueError("median of an empty list")
    ordered = sorted(values)
    middle = len(ordered) // 2
    if len(ordered) % 2 == 1:
        return ordered[middle]
    return (ordered[middle - 1] + ordered[middle]) / 2


if __name__ == "__main__":
    print(median([3, 1, 2]), median([4, 1, 3, 2]))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

def median(values: list[float]) -> float:
    """Return the middle value; for an even count, the mean of the two middle values.

    >>> median([3, 1, 2])
    2
    """
    # Add two more examples to the docstring above:
    # - an even number of values, such as [4, 1, 3, 2]
    # - an empty list, which raises ValueError: median of an empty list
    if not values:
        raise ValueError("median of an empty list")
    ordered = sorted(values)
    middle = len(ordered) // 2
    if len(ordered) % 2 == 1:
        return ordered[middle]
    return (ordered[middle - 1] + ordered[middle]) / 2


if __name__ == "__main__":
    print(median([3, 1, 2]), median([4, 1, 3, 2]))

test_main.py

import doctest

import main


def run_examples(implementation):
    parser = doctest.DocTestParser()
    test = parser.get_doctest(main.median.__doc__ or "", {"median": implementation}, "median", "main.py", 0)
    runner = doctest.DocTestRunner()
    return runner.run(test, out=lambda text: None)


def middle_only(values):
    ordered = sorted(values)
    return ordered[len(ordered) // 2]


def no_empty_check(values):
    ordered = sorted(values)
    middle = len(ordered) // 2
    if len(ordered) % 2 == 1:
        return ordered[middle]
    return (ordered[middle - 1] + ordered[middle]) / 2


def test_examples_pass():
    """Your docstring has at least 3 examples, and they pass"""
    failed, attempted = run_examples(main.median)
    assert attempted >= 3, f"the docstring has {attempted} examples; write at least 3"
    assert failed == 0, f"{failed} of your examples fail on the correct median"


def test_catches_even_bug():
    """Your examples catch a median that ignores even counts"""
    failed, _ = run_examples(middle_only)
    assert failed > 0, "every example passes when median([4, 1, 3, 2]) returns 3; add an even-length example"


def test_catches_empty_bug():
    """Your examples catch a median without the empty-list check"""
    failed, _ = run_examples(no_empty_check)
    assert failed > 0, "every example passes when median([]) raises IndexError; show the ValueError in an example"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 2 of 2

One behaviour per test

A valid username has 3 to 15 characters: letters, digits or _, starting with a letter. Replace test_everything with at least five focused tests, each named after the one behaviour it checks. Cover both length limits and the lengths just outside them, a name starting with a digit and a forbidden character. The checks run your tests against six broken versions, each wrong in one rule.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    One method per rule, for example def test_two_characters_is_too_short(self) -> None: with a single assertFalse inside.

  2. Hint 2

    Boundaries come in pairs: "ada" (3) must pass and "al" (2) must fail; "a" * 15 must pass and "a" * 16 must fail.

  3. Hint 3

    Two more tests: a name starting with a digit such as "1ada", and one with a hyphen such as "ada-lovelace".

Show a solution

One way to solve it. Yours can look different and still pass the checks.

import unittest


def is_valid_username(name: str) -> bool:
    """3 to 15 characters: letters, digits or _, starting with a letter."""
    if not 3 <= len(name) <= 15:
        return False
    if not name[0].isalpha():
        return False
    return all(ch.isalnum() or ch == "_" for ch in name)


class TestUsername(unittest.TestCase):
    def test_typical_name_is_valid(self) -> None:
        self.assertTrue(is_valid_username("ada_1815"))

    def test_three_characters_is_shortest_allowed(self) -> None:
        self.assertTrue(is_valid_username("ada"))

    def test_two_characters_is_too_short(self) -> None:
        self.assertFalse(is_valid_username("al"))

    def test_fifteen_characters_is_longest_allowed(self) -> None:
        self.assertTrue(is_valid_username("a" * 15))

    def test_sixteen_characters_is_too_long(self) -> None:
        self.assertFalse(is_valid_username("a" * 16))

    def test_must_start_with_a_letter(self) -> None:
        self.assertFalse(is_valid_username("1ada"))

    def test_hyphen_is_not_allowed(self) -> None:
        self.assertFalse(is_valid_username("ada-lovelace"))


if __name__ == "__main__":
    suite = unittest.TestLoader().loadTestsFromTestCase(TestUsername)
    unittest.TextTestRunner(verbosity=2).run(suite)
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

import unittest


def is_valid_username(name: str) -> bool:
    """3 to 15 characters: letters, digits or _, starting with a letter."""
    if not 3 <= len(name) <= 15:
        return False
    if not name[0].isalpha():
        return False
    return all(ch.isalnum() or ch == "_" for ch in name)


class TestUsername(unittest.TestCase):
    # Split this into focused tests, one behaviour each, and add the
    # missing edge cases: both length limits and their neighbours,
    # a name starting with a digit, and a character that is not allowed.
    def test_everything(self) -> None:
        self.assertTrue(is_valid_username("ada_1815"))
        self.assertFalse(is_valid_username("al"))


if __name__ == "__main__":
    suite = unittest.TestLoader().loadTestsFromTestCase(TestUsername)
    unittest.TextTestRunner(verbosity=2).run(suite)

test_main.py

import unittest

import main


def run_against(implementation):
    saved = main.is_valid_username
    main.is_valid_username = implementation
    try:
        result = unittest.TestResult()
        unittest.TestLoader().loadTestsFromTestCase(main.TestUsername).run(result)
    finally:
        main.is_valid_username = saved
    return result


def rules(name, low=3, high=15, digit_start=False, hyphen=False):
    if not low <= len(name) <= high:
        return False
    if not (name[0].isalpha() or (digit_start and name[0].isdigit())):
        return False
    return all(ch.isalnum() or ch == "_" or (hyphen and ch == "-") for ch in name)


BROKEN = {
    "3 characters are refused": lambda name: rules(name, low=4),
    "2 characters are accepted": lambda name: rules(name, low=2),
    "15 characters are refused": lambda name: rules(name, high=14),
    "16 characters are accepted": lambda name: rules(name, high=16),
    "a name may start with a digit": lambda name: rules(name, digit_start=True),
    "a hyphen is accepted": lambda name: rules(name, hyphen=True),
}


def test_correct_code_passes():
    """At least 5 tests, and all pass on the correct function"""
    result = run_against(main.is_valid_username)
    failing = [test.id() for test, _ in result.failures + result.errors]
    assert result.wasSuccessful(), f"these tests fail on correct code: {failing}"
    assert result.testsRun >= 5, f"{result.testsRun} tests ran; write at least 5 focused tests"


def test_every_bug_is_caught():
    """Each of six one-rule bugs makes a test fail"""
    missed = [bug for bug, broken in BROKEN.items() if run_against(broken).wasSuccessful()]
    if missed:
        assert False, f"no test fails when {missed[0]}; add that edge case"


def test_tests_are_focused():
    """A one-rule bug leaves the other tests passing"""
    for bug, broken in BROKEN.items():
        result = run_against(broken)
        bad = len(result.failures) + len(result.errors)
        assert bad < result.testsRun, f"when {bug}, every test fails; test one behaviour per test"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Common mistakes

No space after >>>

import doctest


def add(a, b):
    """Return a + b.

    >>>add(1, 2)
    3
    """
    return a + b


doctest.run_docstring_examples(add, globals(), name="add")

What Python prints

ValueError: line 3 of the docstring for add lacks blank after >>>: '>>>add(1, 2)'

Why, and the fix

doctest recognises an example only by >>> followed by a space, like the prompt in an interactive session. Without the space it refuses to parse the docstring at all. Write >>> add(1, 2).

Output indented differently from its example

import doctest


def add(a, b):
    """Return a + b.

      >>> add(1, 2)
    3
    """
    return a + b


doctest.run_docstring_examples(add, globals(), name="add")

What Python prints

ValueError: line 4 of the docstring for add has inconsistent leading whitespace: '3'

Why, and the fix

The expected output must start in the same column as the >>> of its example. Here the example is indented two spaces more than its output, so doctest cannot tell where the output begins. Line both up with the rest of the docstring.

Passing a module name instead of the module

import doctest


def add(a, b):
    """
    >>> add(1, 2)
    3
    """
    return a + b


doctest.testmod("main")

What Python prints

TypeError: testmod: module required; 'main'

Why, and the fix

testmod(m) needs a module object, not its name as a string. Call doctest.testmod() without arguments to test the module you run as a script, or import the module first and pass it: import shop, then doctest.testmod(shop).

Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

Examples that check themselves

doctest finds text in docstrings that looks like an interactive session: a line starting with >>> and, below it, the output you expect. It runs the code and compares the output exactly, character by character. A value is shown as its repr, so a string appears with quotes; printed text appears without. The expected output ends at the next >>> or blank line. For an exception, write Traceback (most recent call last): and then the exception line; the stack in between is ignored and may be ... .

Running doctests

The usual way is to end a module with if __name__ == "__main__": import doctest; doctest.testmod(). Run as a script, it checks every docstring in the module and prints nothing when all examples pass; a failure shows the example, what was expected and what it got. testmod() returns TestResults(failed, attempted). Doctests are best as documentation that stays true; detailed checks of many cases belong in unittest tests.

What to test, and how much per test

Bugs gather at the edges: an empty list, a single item, zero, negative numbers, and both sides of every boundary. If a rule says 18 or more, test 17 and 18. Let each test check one behaviour and name it after that behaviour, such as test_empty_list_raises. A test stops at its first failing assert, so a test with five rules hides the other four. With focused tests, a bug fails exactly the test that names it.

Sources

Last reviewed September 29, 2026