Skip to content
aviral gupta

// I4.1 · ~30 min · Intermediate

Files and folders with pathlib

After this lesson you can build paths with Path and /, create folders safely, read and write text files, and find files with glob patterns.

Lesson 1 of 6 in I4 The standard library for real programs

Start of the module

You will be able to

  • Build paths with / and read their parts: name, stem, suffix and parent
  • Create folders and files with mkdir, write_text and read_text, and check them with exists
  • Find files with glob patterns such as *.txt and **/*.log, and sort the results
  1. Warm-up · Activity 1 of 7

    Warm-up from module I3: a generator expression is lazy. What does this print?

    squares = (n * n for n in range(4))
    print(sum(squares), sum(squares))
  2. Predict · Activity 2 of 7

    Predict before you read on: what does this print?

    from pathlib import Path
    
    p = Path("reports") / "2026" / "summary.txt"
    print(p.name, p.stem, p.suffix)
  3. Practice · Activity 3 of 7

    The program may run many times. Fill in the argument so that mkdir does not fail when the folder already exists.

    folder.mkdir(parents=True, ____=True)
    folder.mkdir(parents=True, =True)
  4. Practice · Activity 4 of 7

    A file name with two dots. What does this print?

    from pathlib import Path
    
    p = Path("backup.tar.gz")
    print(p.stem, p.suffix)
  5. Practice · Activity 5 of 7

    Match each call to what it returns.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. The folder t holds exactly two .txt files. What does this print?

    from pathlib import Path
    
    folder = Path("t")
    folder.mkdir(exist_ok=True)
    (folder / "a.txt").write_text("a", encoding="utf-8")
    (folder / "b.txt").write_text("b", encoding="utf-8")
    
    texts = folder.glob("*.txt")
    print(len(list(texts)), len(list(texts)))
  7. Apply · Activity 7 of 7

    Mini-task. Make a folder notes with a few .txt files and one .csv file. Then write a program that prints every .txt file in notes with its number of lines, sorted by name, for example a.txt: 3 lines. The .csv file must not appear.

    Check your work against this list

Build it yourself

Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.

Worked example

Tidy up a downloads folder

The program first creates a messy downloads folder with five files. Then it moves each file into a subfolder named after its suffix: pdf, jpg, txt, or other for a file without one. Finally it checks the result with glob, read_text and exists. Run it twice: exist_ok=True lets the second run reuse the folders.

main.py

from pathlib import Path

# A messy downloads folder to tidy up.
downloads = Path("downloads")
downloads.mkdir(exist_ok=True)
for name in ["report.pdf", "photo.jpg", "notes.txt", "todo.txt", "README"]:
    (downloads / name).write_text(f"contents of {name}\n", encoding="utf-8")

# Move each file into a folder named after its suffix.
for path in sorted(downloads.glob("*")):
    if not path.is_file():
        continue
    kind = path.suffix.removeprefix(".") or "other"
    target = downloads / kind
    target.mkdir(exist_ok=True)
    path.replace(target / path.name)
    print(f"{path.name:<11} -> {kind}/")

print(sorted(p.stem for p in downloads.glob("*/*.txt")))
print((downloads / "txt" / "todo.txt").read_text(encoding="utf-8").strip())
print((downloads / "notes.txt").exists(), (downloads / "txt").is_dir())

Run it with

python main.py

Output

README      -> other/
notes.txt   -> txt/
photo.jpg   -> jpg/
report.pdf  -> pdf/
todo.txt    -> txt/
['notes', 'todo']
contents of todo.txt
False True
  • README has no suffix, so path.suffix is "" and or picks "other".
  • sorted() gives a fixed order; glob alone gives no particular order.
  • is_file() skips the subfolders, which "*" also matches on a second run.
  • replace moves the file and overwrites one of the same name, so a second run does not fail.
Change it and run it

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Exercises

Exercise 1 of 2

Count files by suffix

Write count_by_suffix(folder), which returns a dict from each suffix to the number of files directly in folder with that suffix. Files without a suffix count under "". Folders are not files: skip them, and do not look inside them. For a.txt, b.txt, c.csv, README and a subfolder, the result is {".txt": 2, ".csv": 1, "": 1}.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    folder.glob("*") yields everything directly in folder, files and subfolders.

  2. Hint 2

    path.is_file() is False for a folder, so an if skips subfolders.

  3. Hint 3

    counts[path.suffix] = counts.get(path.suffix, 0) + 1 counts each suffix; a file without one has suffix "".

Show a solution

One way to solve it. Yours can look different and still pass the checks.

from pathlib import Path


def count_by_suffix(folder: Path) -> dict[str, int]:
    counts: dict[str, int] = {}
    for path in folder.glob("*"):
        if path.is_file():
            counts[path.suffix] = counts.get(path.suffix, 0) + 1
    return counts


if __name__ == "__main__":
    print(count_by_suffix(Path(".")))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

from pathlib import Path


def count_by_suffix(folder: Path) -> dict[str, int]:
    counts: dict[str, int] = {}
    # Look at every file in folder and count its suffix.
    return counts


if __name__ == "__main__":
    print(count_by_suffix(Path(".")))

test_main.py

from pathlib import Path

from main import count_by_suffix


def make(folder, names):
    root = Path(folder)
    root.mkdir(exist_ok=True)
    for name in names:
        (root / name).write_text("x", encoding="utf-8")
    return root


def test_counts():
    """Two .txt files, one .csv and one file without a suffix"""
    root = make("t_counts", ["a.txt", "b.txt", "c.csv", "README"])
    got = count_by_suffix(root)
    assert got == {".txt": 2, ".csv": 1, "": 1}, f"count_by_suffix returned {got!r}"


def test_skips_folders():
    """Subfolders and the files inside them are not counted"""
    root = make("t_folders", ["a.txt"])
    (root / "sub").mkdir(exist_ok=True)
    (root / "sub" / "b.txt").write_text("x", encoding="utf-8")
    got = count_by_suffix(root)
    assert got == {".txt": 1}, f"with a.txt and sub/b.txt, count_by_suffix returned {got!r}, expected {{'.txt': 1}}"


def test_empty():
    """An empty folder gives an empty dict"""
    root = make("t_empty", [])
    got = count_by_suffix(root)
    assert got == {}, f"for an empty folder, count_by_suffix returned {got!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 2 of 2

A report archive

Write save_report(base, year, title, text), which saves text in base/<year>/<title>.md, creating the folders as needed, and returns the file's path. Saving again must work. Then write report_titles(base, year), which returns the sorted titles (names without .md) of the .md files in that year's folder, and [] when the folder does not exist.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    Build the folder path first: folder = base / str(year). Then call folder.mkdir(parents=True, exist_ok=True).

  2. Hint 2

    The file is folder / (title + ".md"); write_text creates it or replaces it.

  3. Hint 3

    In report_titles, check folder.exists() first, then sort path.stem for every path in folder.glob("*.md").

Show a solution

One way to solve it. Yours can look different and still pass the checks.

from pathlib import Path


def save_report(base: Path, year: int, title: str, text: str) -> Path:
    folder = base / str(year)
    folder.mkdir(parents=True, exist_ok=True)
    path = folder / (title + ".md")
    path.write_text(text, encoding="utf-8")
    return path


def report_titles(base: Path, year: int) -> list[str]:
    folder = base / str(year)
    if not folder.exists():
        return []
    return sorted(path.stem for path in folder.glob("*.md"))


if __name__ == "__main__":
    save_report(Path("reports"), 2026, "september", "All good.")
    print(report_titles(Path("reports"), 2026))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

from pathlib import Path


def save_report(base: Path, year: int, title: str, text: str) -> Path:
    path = base / str(year) / (title + ".md")
    path.write_text(text, encoding="utf-8")
    return path


def report_titles(base: Path, year: int) -> list[str]:
    return []


if __name__ == "__main__":
    save_report(Path("reports"), 2026, "september", "All good.")
    print(report_titles(Path("reports"), 2026))

test_main.py

from pathlib import Path

from main import report_titles, save_report


def test_saves_file():
    """save_report creates the folders and writes the text"""
    path = save_report(Path("t_save") / "reports", 2026, "may", "Sales up.")
    assert path == Path("t_save/reports/2026/may.md"), f"save_report returned {path!r}"
    text = path.read_text(encoding="utf-8")
    assert text == "Sales up.", f"the file holds {text!r}, expected 'Sales up.'"


def test_save_twice():
    """Saving a second report into the same year works"""
    base = Path("t_twice")
    save_report(base, 2025, "one", "1")
    save_report(base, 2025, "two", "2")
    got = (base / "2025" / "two.md").read_text(encoding="utf-8")
    assert got == "2", f"the second report holds {got!r}, expected '2'"


def test_titles_sorted():
    """report_titles lists only .md files, sorted, without .md"""
    base = Path("t_titles")
    for title in ["march", "april", "june"]:
        save_report(base, 2024, title, "x")
    (base / "2024" / "notes.txt").write_text("x", encoding="utf-8")
    got = report_titles(base, 2024)
    assert got == ["april", "june", "march"], f"report_titles returned {got!r}"


def test_missing_year():
    """A year without a folder gives []"""
    got = report_titles(Path("t_missing"), 1999)
    assert got == [], f"for a missing folder, report_titles returned {got!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Common mistakes

Joining a Path and a string with +

from pathlib import Path

folder = Path("data")
print(folder + "/scores.txt")

What Python prints

TypeError: unsupported operand type(s) for +: 'PosixPath' and 'str'

Why, and the fix

A Path is not a string, so + does not work. Join with /: folder / "scores.txt". On Windows the message names WindowsPath instead of PosixPath. If you need the path as text, for example for a message, use str(folder / "scores.txt").

Creating a folder that already exists

from pathlib import Path

Path("out").mkdir()
Path("out").mkdir()

What Python prints

File exists: 'out'

Why, and the fix

The full line is FileExistsError: [Errno 17] File exists: 'out' (the browser shows another number). mkdir fails when the folder is already there, so a program that works the first time fails the second time. Pass exist_ok=True when an existing folder is fine.

Creating nested folders in one step

from pathlib import Path

Path("backup/2026/09").mkdir()

What Python prints

No such file or directory: 'backup/2026/09'

Why, and the fix

The full line is FileNotFoundError: [Errno 2] No such file or directory: 'backup/2026/09'. mkdir creates only the last folder, and backup/2026 does not exist yet. Pass parents=True to create the missing folders above it, usually together with exist_ok=True.

Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

A path is an object, joined with /

Path("data") / "raw" / "sales.csv" builds a path from parts, the way a slash does on the command line, without string gluing. The object knows its parts: name is "sales.csv", stem is "sales", suffix is ".csv" and parent is the folder above. For "archive.tar.gz" the suffix is only the last part, ".gz". print shows data/raw/sales.csv here and in the browser; on Windows it shows backslashes, and the same code works on both. A str and a Path cannot be added with +.

Folders and files

mkdir creates one folder. It fails with FileExistsError if the folder is already there, unless you pass exist_ok=True, and with FileNotFoundError if a parent is missing, unless you pass parents=True. write_text(text, encoding="utf-8") replaces the whole file and returns the number of characters written; read_text(encoding="utf-8") returns the whole file as one string. Both open and close the file for you. exists, is_file and is_dir answer questions without changing anything.

Finding files with glob

folder.glob("*.txt") yields the paths in folder that match the pattern. * matches any part of a name, ? one character, and ** any number of folder levels, so "**/*.log" finds .log files at every depth, the folder itself included. The paths come in no particular order, so sort them when order matters. glob returns an iterator: it can be walked once. Store list(folder.glob(...)) if you need the paths twice. "*" also matches folders; is_file() filters them out.

Sources

Last reviewed September 29, 2026