Skip to content
aviral gupta

// Intermediate project ยท about 10 hours of work

Log analyser

You build loganalyser, a command-line tool that reads one or more log files, counts the entries per level, source or hour, and writes the counts as CSV or JSON. It reads lazily with generators, so a large file never has to fit in memory. Bad lines are skipped with a logged warning, or stop the run with --strict. Errors become exit statuses a script can test: 0 for success, 1 for bad input, 2 for bad arguments. A pyproject.toml makes it an installable package with its own command. It brings the Intermediate level together: custom exceptions, dataclasses, enums, generators, collections, pathlib, argparse, re, datetime, csv and logging.

What the finished program does

  • A log line reads 2026-09-28 14:03:12 LEVEL source: message. Levels are DEBUG, INFO, WARNING (or WARN), ERROR and CRITICAL (or FATAL), in any case. Blank lines are ignored.
  • Each parsed line becomes a frozen Entry dataclass; Level is an IntEnum, so levels compare by severity.
  • All deliberate errors derive from LogAnalyserError. ParseError carries the line number, and chains the ValueError that caused it; InputError reports a missing or unreadable file.
  • Files are read one line at a time by generators and can be given in any number; filters and counting never build a list of every entry.
  • Options: --by level|source|hour, --format csv|json, -o FILE, --since and --until (ISO dates or date-times; since inclusive, until exclusive), --level MIN, --strict, -v and --version.
  • Counts come in a stable order: levels by severity, sources most common first with ties by name, hours in time order. JSON also holds the number of entries and the first and last time.
  • Progress and warnings are logged to standard error, never mixed into the output. -v shows each file read and the number of entries counted.
  • main(argv) returns 0 on success and 1 for a missing file or, with --strict, a bad line; argparse exits with 2 on bad arguments.
  • pyproject.toml declares the package and a loganalyser console script, so python -m pip install -e . in a virtual environment installs the command.

Starter layout

pyproject.toml
Package metadata and the build backend. You add the [project.scripts] table.
loganalyser/__init__.py
Marks the package and holds __version__. Complete.
loganalyser/__main__.py
Lets python -m loganalyser run the command line. Complete.
loganalyser/errors.py
LogAnalyserError, InputError and ParseError. ParseError.__init__ is a stub.
loganalyser/model.py
The Level enum and the Entry dataclass. Level.parse is a stub.
loganalyser/parser.py
The line regex, parse_line and the generators parse_lines and read_entries. Stubs.
loganalyser/report.py
GroupBy, the within and at_least filters, the Report dataclass, summarise and the CSV and JSON writers. Stubs.
loganalyser/cli.py
The argparse parser, logging set-up, the output context manager and main(argv). Logging set-up is written; the rest is stubs and TODOs.
samples/app.log
A sample log with every level, a broken line and a blank line, to try the tool on.

Milestones

  1. Milestone 1

    The model and one line

    Write Level.parse and ParseError, then the LINE_RE regex with named groups and parse_line. Turn the ValueError from strptime or Level.parse into a ParseError with raise ... from err.

    Checks that pass once this milestone is done:

    • Level.parse reads names in any case, accepts WARN, and levels compare by severity
    • parse_line turns a log line into an Entry
    • parse_line raises ParseError, a LogAnalyserError, with the line number
  2. Milestone 2

    Lazy reading

    Write parse_lines and read_entries as generators: enumerate from 1, log and skip a bad line unless strict, and use path.open in a with statement and yield from for each file.

    Checks that pass once this milestone is done:

    • read_entries is a lazy generator that skips a bad line with a warning, or raises with strict
  3. Milestone 3

    Filters and counts

    Write within and at_least as generators, GroupBy.key with match, and summarise with a Counter in a single pass. Then Report.to_dict, write_csv and write_json.

    Checks that pass once this milestone is done:

    • within keeps since <= time < until, and at_least keeps a level and above
    • summarise counts per level in severity order, per source most common first, and per hour
    • write_csv writes a header and a row per group; write_json writes the whole report
  4. Milestone 4

    The command line

    Add the options to build_parser, write the when and level types, open_output with @contextmanager, and main(argv), which chains read_entries, the filters and summarise, then writes the report.

    Checks that pass once this milestone is done:

    • main([FILE]) prints counts by level as CSV and returns 0
    • --format json -o FILE writes the report to the file and prints nothing
    • --since, --until and --level narrow what is counted
    • Several files are read in turn and counted together
    • samples/app.log gives the expected counts
  5. Milestone 5

    Errors, exit codes and logging

    Catch LogAnalyserError in main, log it and return 1. Let argparse exit with 2 by raising ArgumentTypeError in your types. Log each file read and the count at INFO, shown with -v.

    Checks that pass once this milestone is done:

    • A missing file gives exit status 1 and an error on standard error
    • A bad line is a logged warning, or exit status 1 with --strict
    • Bad arguments make argparse exit with status 2
    • -v logs each file read and the count to standard error
  6. Milestone 6

    Install it

    Add [project.scripts] with loganalyser = "loganalyser.cli:main". In a new virtual environment run python -m pip install -e . and then loganalyser samples/app.log -v.

    Checks that pass once this milestone is done:

    • pyproject.toml declares the loganalyser console script

This project uses parts of Python that do not run in the browser, so you build it on your computer.

Build it on your computer

Make a folder with these starter files and Python 3.14, then work through the milestones. Run the acceptance tests at any point with:

Download the starter as one .zip (starter files, test_main.py and learnrun.py)
python learnrun.py test

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Download learnrun.py

pyproject.toml

[build-system]
requires = ["setuptools >= 77.0.3"]
build-backend = "setuptools.build_meta"

[project]
name = "loganalyser"
version = "1.0.0"
description = "Count log entries per level, source or hour, and write CSV or JSON."
requires-python = ">= 3.14"
dependencies = []

# TODO: a [project.scripts] table, so that pip install -e . creates a
# loganalyser command that calls main() in loganalyser/cli.py.

[tool.setuptools]
packages = ["loganalyser"]

loganalyser/__init__.py

"""Count log entries per level, source or hour, and write CSV or JSON."""

__version__ = "1.0.0"

loganalyser/__main__.py

"""python -m loganalyser runs the command line."""

import sys

from .cli import main

sys.exit(main())

loganalyser/errors.py

"""The exceptions of the log analyser; the command line turns each into exit status 1."""


class LogAnalyserError(Exception):
    """Base class: every error this package raises on purpose."""


class InputError(LogAnalyserError):
    """A log file is missing, is a folder, or cannot be read."""


class ParseError(LogAnalyserError):
    """A line does not have the expected log format."""

    def __init__(self, line_no: int, line: str, reason: str) -> None:
        # TODO: call super().__init__ with a message such as
        # "line 3: not a log line: 'the line'", and keep line_no, line and reason
        # as attributes.
        raise NotImplementedError

loganalyser/model.py

"""The domain: log levels and parsed log entries."""

from dataclasses import dataclass
from datetime import datetime
from enum import IntEnum


class Level(IntEnum):
    """Log levels, ordered by severity so they compare: Level.ERROR > Level.INFO."""

    DEBUG = 10
    INFO = 20
    WARNING = 30
    ERROR = 40
    CRITICAL = 50

    @classmethod
    def parse(cls, text: str) -> "Level":
        """The level named by text, in any case; WARN and FATAL are accepted too."""
        # TODO: look the name up with cls[name]; turn a KeyError into ValueError.
        raise NotImplementedError


@dataclass(frozen=True, slots=True)
class Entry:
    """One parsed log line."""

    timestamp: datetime
    level: Level
    source: str
    message: str
    line_no: int = 0

loganalyser/parser.py

"""Turning lines of text into Entry objects, lazily."""

import logging
import re
from collections.abc import Iterable, Iterator
from datetime import datetime
from pathlib import Path

from .errors import InputError, ParseError
from .model import Entry, Level

logger = logging.getLogger(__name__)

# 2026-09-28 14:03:12 INFO  auth: user alice logged in
# TODO: named groups timestamp, level, source and message.
LINE_RE = re.compile(r"TODO")
TIMESTAMP_FORMAT = "%Y-%m-%d %H:%M:%S"


def parse_line(line: str, line_no: int = 0) -> Entry:
    """Parse one log line. Raise ParseError if it is not in the expected format."""
    # TODO: match LINE_RE, then datetime.strptime and Level.parse; turn their
    # ValueError into a ParseError chained with `from`.
    raise NotImplementedError


def parse_lines(lines: Iterable[str], *, source: str = "<input>", strict: bool = False) -> Iterator[Entry]:
    """Yield an Entry per log line. Blank lines are ignored.

    A bad line raises ParseError when strict, and is logged as a warning and
    skipped otherwise.
    """
    # TODO: a generator; number the lines from 1 with enumerate.
    raise NotImplementedError


def read_entries(paths: Iterable[Path], *, strict: bool = False) -> Iterator[Entry]:
    """Yield the entries of each file in turn, reading one line at a time."""
    # TODO: raise InputError for a path that is not a file; open each with
    # `with path.open(encoding="utf-8")` and `yield from parse_lines(...)`.
    raise NotImplementedError

loganalyser/report.py

"""Filtering, counting and writing the results as CSV or JSON."""

import csv
import json
from collections import Counter
from collections.abc import Iterable, Iterator
from dataclasses import dataclass, field
from datetime import datetime
from enum import Enum
from typing import TextIO

from .model import Entry, Level


class GroupBy(Enum):
    """What to count entries by."""

    LEVEL = "level"
    SOURCE = "source"
    HOUR = "hour"

    def key(self, entry: Entry) -> str:
        """The group an entry falls into, as text: "ERROR", "db" or "2026-09-28 14:00"."""
        # TODO: match self.
        raise NotImplementedError


def within(entries: Iterable[Entry], since: datetime | None = None, until: datetime | None = None) -> Iterator[Entry]:
    """The entries at or after since and before until (either may be None)."""
    # TODO: a generator.
    raise NotImplementedError


def at_least(entries: Iterable[Entry], level: Level) -> Iterator[Entry]:
    """The entries whose level is level or more severe."""
    # TODO: a generator expression.
    raise NotImplementedError


@dataclass
class Report:
    """Counts per group, plus the size and time span of what was counted."""

    by: GroupBy
    counts: list[tuple[str, int]] = field(default_factory=list)
    entries: int = 0
    first: datetime | None = None
    last: datetime | None = None

    def to_dict(self) -> dict[str, object]:
        """A JSON-ready dict: by, entries, first, last and counts."""
        # TODO: first and last as "2026-09-28 09:00:01"; counts as
        # [{"level": "INFO", "count": 2}, ...], keyed by self.by.value.
        raise NotImplementedError


def summarise(entries: Iterable[Entry], by: GroupBy) -> Report:
    """Count the entries per group in one pass over them.

    Order the counts: levels by severity, sources most common first (ties by
    name), hours in time order.
    """
    # TODO: a collections.Counter, then sort.
    raise NotImplementedError


def write_csv(report: Report, out: TextIO) -> None:
    """One header row (the group name and count), then a row per group."""
    # TODO: csv.writer(out, lineterminator="\n").
    raise NotImplementedError


def write_json(report: Report, out: TextIO) -> None:
    """The report as an indented JSON object."""
    # TODO: json.dump(report.to_dict(), ...).
    raise NotImplementedError

loganalyser/cli.py

"""The command line: loganalyser [options] FILE [FILE ...]."""

import argparse
import logging
import sys
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import datetime
from pathlib import Path
from typing import TextIO

from . import __version__
from .errors import LogAnalyserError, ParseError
from .model import Level
from .parser import read_entries
from .report import GroupBy, at_least, summarise, within, write_csv, write_json

logger = logging.getLogger("loganalyser")

EXIT_OK = 0
EXIT_ERROR = 1  # bad input: a missing file, or a bad line with --strict
# argparse itself exits with status 2 on a usage error.


def when(text: str) -> datetime:
    """An argparse type: 2026-09-28, 2026-09-28T14:00 or "2026-09-28 14:00"."""
    # TODO: datetime.fromisoformat; raise argparse.ArgumentTypeError if it fails.
    raise NotImplementedError


def level(text: str) -> Level:
    """An argparse type: a level name such as warning or ERROR."""
    # TODO: Level.parse; raise argparse.ArgumentTypeError if it fails.
    raise NotImplementedError


def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(
        prog="loganalyser",
        description="Count log entries per level, source or hour, and write the counts as CSV or JSON.",
    )
    parser.add_argument("files", nargs="+", type=Path, metavar="FILE", help="log files to read, in order")
    # TODO: --by, --format, -o/--output, --since, --until, --level, --strict,
    # -v/--verbose (a count) and --version.
    return parser


def configure_logging(verbosity: int) -> None:
    """Progress and warnings go to standard error, so they never mix with the output."""
    handler = logging.StreamHandler(sys.stderr)
    handler.setFormatter(logging.Formatter("%(levelname)s: %(message)s"))
    logger.handlers[:] = [handler]
    logger.setLevel(logging.WARNING if verbosity == 0 else logging.INFO if verbosity == 1 else logging.DEBUG)
    logger.propagate = False


@contextmanager
def open_output(path: Path | None) -> Iterator[TextIO]:
    """Standard output, or the file at path, closed afterwards."""
    # TODO: yield sys.stdout for None; otherwise open the file with
    # newline="" (the csv module wants it) inside a with statement.
    raise NotImplementedError
    yield sys.stdout


def main(argv: list[str] | None = None) -> int:
    """Run the tool with argv (default: sys.argv[1:]); return the exit status."""
    # TODO: parse, configure logging, read -> filter -> summarise -> write,
    # and turn each LogAnalyserError into a logged error and EXIT_ERROR.
    raise NotImplementedError


if __name__ == "__main__":
    sys.exit(main())

samples/app.log

2026-09-28 08:59:58 INFO  web: server started on port 8080
2026-09-28 09:00:01 INFO  auth: user alice logged in
2026-09-28 09:00:07 DEBUG db: pool size 5
2026-09-28 09:02:13 WARN  web: slow response for /reports (2.4 s)
2026-09-28 09:15:42 ERROR db: connection lost, retrying
2026-09-28 09:15:43 INFO  db: reconnected
this line was cut off by a crash
2026-09-28 10:01:09 INFO  auth: user bob logged in
2026-09-28 10:05:30 ERROR web: 500 on /export (KeyError: 'month')
2026-09-28 10:05:31 WARNING auth: 3 failed logins for carol

2026-09-28 11:20:00 CRITICAL db: disk full
2026-09-28 11:20:05 INFO  web: server stopping

Acceptance tests

The project is done when every check in test_main.py passes. Read them before you start: they are the spec, written as code.

test_main.py

import contextlib
import csv
import inspect
import io
import json
import tempfile
import tomllib
from datetime import datetime
from pathlib import Path

from loganalyser.cli import main
from loganalyser.errors import InputError, LogAnalyserError, ParseError
from loganalyser.model import Entry, Level
from loganalyser.parser import parse_line, parse_lines, read_entries
from loganalyser.report import GroupBy, at_least, summarise, within, write_csv, write_json

LOG = """2026-09-28 09:00:01 INFO  auth: user alice logged in
2026-09-28 09:15:42 ERROR db: connection lost
not a log line
2026-09-28 10:05:30 WARN  web: slow response

2026-09-28 10:59:59 ERROR web: 500 on /export
2026-09-28 11:00:00 INFO  auth: user bob logged in
"""


def entries_of(text=LOG):
    """The entries of text, parsed leniently (the warning for line 3 is dropped)."""
    with contextlib.redirect_stderr(io.StringIO()):
        return list(parse_lines(io.StringIO(text)))


def run(argv):
    """Call main(argv) and return (status, stdout, stderr)."""
    out, err = io.StringIO(), io.StringIO()
    with contextlib.redirect_stdout(out), contextlib.redirect_stderr(err):
        status = main(argv)
    return status, out.getvalue(), err.getvalue()


def write_log(folder, text=LOG, name="app.log"):
    path = Path(folder) / name
    path.write_text(text, encoding="utf-8")
    return str(path)


def test_level_parse():
    """Level.parse reads names in any case, accepts WARN, and levels compare by severity"""
    got = Level.parse("error"), Level.parse("WARN"), Level.parse("Info")
    assert got == (Level.ERROR, Level.WARNING, Level.INFO), f"Level.parse gave {got!r}"
    assert Level.ERROR > Level.WARNING > Level.INFO > Level.DEBUG, "levels should compare by severity: ERROR > WARNING > INFO > DEBUG"
    try:
        Level.parse("loud")
    except ValueError:
        return
    raise AssertionError("Level.parse('loud') should raise ValueError")


def test_parse_line():
    """parse_line turns a log line into an Entry"""
    got = parse_line("2026-09-28 09:15:42 ERROR db: connection lost, retrying\n", 7)
    want = Entry(datetime(2026, 9, 28, 9, 15, 42), Level.ERROR, "db", "connection lost, retrying", 7)
    assert got == want, f"parse_line gave {got!r}, expected {want!r}"


def test_parse_line_rejects():
    """parse_line raises ParseError, a LogAnalyserError, with the line number"""
    for line in ["not a log line", "2026-02-30 09:00:00 INFO web: no such day", "2026-09-28 09:00:00 LOUD web: odd level"]:
        try:
            got = parse_line(line, 3)
        except ParseError as err:
            assert isinstance(err, LogAnalyserError), "ParseError should be a subclass of LogAnalyserError"
            assert err.line_no == 3, f"for {line!r} the error has line_no {err.line_no!r}, expected 3"
            continue
        raise AssertionError(f"parse_line({line!r}) returned {got!r}, expected a ParseError")


def test_read_entries_skips_bad_lines():
    """read_entries is a lazy generator that skips a bad line with a warning, or raises with strict"""
    with tempfile.TemporaryDirectory() as folder:
        path = Path(write_log(folder))
        gen = read_entries([path])
        assert inspect.isgenerator(gen), f"read_entries returned {type(gen).__name__}, expected a generator"
        with contextlib.redirect_stderr(io.StringIO()):
            got = [e.line_no for e in gen]
        assert got == [1, 2, 4, 6, 7], f"read_entries gave entries from lines {got!r}, expected [1, 2, 4, 6, 7]"
        try:
            list(read_entries([path], strict=True))
        except ParseError as err:
            assert err.line_no == 3, f"with strict the ParseError has line_no {err.line_no}, expected 3"
        else:
            raise AssertionError("with strict=True read_entries should raise ParseError at line 3")
        try:
            list(read_entries([Path(folder) / "missing.log"]))
        except InputError:
            pass
        else:
            raise AssertionError("for a missing file read_entries should raise InputError")


def test_filters():
    """within keeps since <= time < until, and at_least keeps a level and above"""
    entries = entries_of()
    got = [e.line_no for e in within(entries, datetime(2026, 9, 28, 9, 15, 42), datetime(2026, 9, 28, 11, 0))]
    assert got == [2, 4, 6], f"within 09:15:42 and 11:00 kept lines {got!r}, expected [2, 4, 6]"
    got = [e.line_no for e in at_least(entries, Level.WARNING)]
    assert got == [2, 4, 6], f"at_least WARNING kept lines {got!r}, expected [2, 4, 6]"


def test_summarise():
    """summarise counts per level in severity order, per source most common first, and per hour"""
    entries = entries_of()
    got = summarise(entries, GroupBy.LEVEL)
    assert got.counts == [("INFO", 2), ("WARNING", 1), ("ERROR", 2)], f"counts by level are {got.counts!r}"
    assert (got.entries, got.first, got.last) == (5, datetime(2026, 9, 28, 9, 0, 1), datetime(2026, 9, 28, 11, 0)), f"entries, first and last are {(got.entries, got.first, got.last)!r}"
    got = summarise(entries, GroupBy.SOURCE).counts
    assert got == [("auth", 2), ("web", 2), ("db", 1)], f"counts by source are {got!r}, expected [('auth', 2), ('web', 2), ('db', 1)]"
    got = summarise(entries, GroupBy.HOUR).counts
    want = [("2026-09-28 09:00", 2), ("2026-09-28 10:00", 2), ("2026-09-28 11:00", 1)]
    assert got == want, f"counts by hour are {got!r}, expected {want!r}"


def test_write_csv_and_json():
    """write_csv writes a header and a row per group; write_json writes the whole report"""
    report = summarise(entries_of(), GroupBy.SOURCE)
    out = io.StringIO()
    write_csv(report, out)
    rows = list(csv.reader(io.StringIO(out.getvalue())))
    assert rows == [["source", "count"], ["auth", "2"], ["web", "2"], ["db", "1"]], f"the CSV rows are {rows!r}"
    out = io.StringIO()
    write_json(report, out)
    got = json.loads(out.getvalue())
    want = {
        "by": "source",
        "entries": 5,
        "first": "2026-09-28 09:00:01",
        "last": "2026-09-28 11:00:00",
        "counts": [{"source": "auth", "count": 2}, {"source": "web", "count": 2}, {"source": "db", "count": 1}],
    }
    assert got == want, f"the JSON is {got!r}, expected {want!r}"


def test_cli_csv_to_stdout():
    """main([FILE]) prints counts by level as CSV and returns 0"""
    with tempfile.TemporaryDirectory() as folder:
        status, out, _ = run([write_log(folder)])
    assert status == 0, f"main returned {status}, expected 0"
    assert out.splitlines() == ["level,count", "INFO,2", "WARNING,1", "ERROR,2"], f"main printed {out!r}"


def test_cli_json_file():
    """--format json -o FILE writes the report to the file and prints nothing"""
    with tempfile.TemporaryDirectory() as folder:
        target = Path(folder) / "report.json"
        status, out, _ = run([write_log(folder), "--by", "hour", "--format", "json", "-o", str(target)])
        assert status == 0, f"main returned {status}, expected 0"
        assert out == "", f"with -o main still printed {out!r}"
        got = json.loads(target.read_text(encoding="utf-8"))
    assert got["by"] == "hour" and got["entries"] == 5, f"the JSON file holds {got!r}"


def test_cli_filters():
    """--since, --until and --level narrow what is counted"""
    with tempfile.TemporaryDirectory() as folder:
        path = write_log(folder)
        _, out, _ = run([path, "--since", "2026-09-28T09:10", "--until", "2026-09-28 11:00"])
        assert out.splitlines() == ["level,count", "WARNING,1", "ERROR,2"], f"with --since and --until main printed {out!r}"
        _, out, _ = run([path, "--level", "error", "--by", "source"])
        assert out.splitlines() == ["source,count", "db,1", "web,1"], f"with --level error --by source main printed {out!r}"


def test_cli_several_files():
    """Several files are read in turn and counted together"""
    with tempfile.TemporaryDirectory() as folder:
        first = write_log(folder, "2026-09-28 09:00:00 INFO a: one\n", "a.log")
        second = write_log(folder, "2026-09-29 09:00:00 INFO b: two\n2026-09-29 09:00:01 ERROR b: three\n", "b.log")
        status, out, _ = run([first, second, "--by", "source"])
    assert (status, out.splitlines()) == (0, ["source,count", "b,2", "a,1"]), f"main returned {status} and printed {out!r}"


def test_cli_missing_file():
    """A missing file gives exit status 1 and an error on standard error"""
    with tempfile.TemporaryDirectory() as folder:
        missing = str(Path(folder) / "nope.log")
        status, out, err = run([missing])
    assert status == 1, f"main returned {status} for a missing file, expected 1"
    assert "nope.log" in err and out == "", f"expected an error naming nope.log on stderr and nothing on stdout; stderr was {err!r}"


def test_cli_strict():
    """A bad line is a logged warning, or exit status 1 with --strict"""
    with tempfile.TemporaryDirectory() as folder:
        path = write_log(folder)
        status, _, err = run([path])
        assert status == 0 and "3" in err and "WARNING" in err, f"without --strict: status {status}, stderr {err!r}; expected 0 and a warning for line 3"
        status, out, err = run([path, "--strict"])
    assert status == 1 and out == "", f"with --strict main returned {status} and printed {out!r}, expected 1 and nothing"
    assert "line 3" in err, f"with --strict the error {err!r} should name line 3"


def test_cli_bad_arguments():
    """Bad arguments make argparse exit with status 2"""
    for argv in [[], ["app.log", "--since", "yesterday"], ["app.log", "--format", "xml"], ["app.log", "--level", "loud"]]:
        try:
            with contextlib.redirect_stderr(io.StringIO()):
                main(argv)
        except SystemExit as exc:
            assert exc.code == 2, f"for {argv!r} the exit status was {exc.code!r}, expected 2"
        else:
            raise AssertionError(f"main({argv!r}) returned instead of exiting with status 2")


def test_cli_logs_progress():
    """-v logs each file read and the count to standard error"""
    with tempfile.TemporaryDirectory() as folder:
        path = write_log(folder)
        status, out, err = run([path, "-v"])
    assert status == 0, f"main returned {status}, expected 0"
    assert "INFO" in err and "app.log" in err and "5" in err, f"with -v stderr was {err!r}, expected INFO lines naming app.log and the count 5"
    assert "INFO" not in out.replace("INFO,", ""), f"log lines leaked into standard output: {out!r}"


def test_sample_log():
    """samples/app.log gives the expected counts"""
    status, out, _ = run(["samples/app.log"])
    want = ["level,count", "DEBUG,1", "INFO,5", "WARNING,2", "ERROR,2", "CRITICAL,1"]
    assert (status, out.splitlines()) == (0, want), f"for samples/app.log main returned {status} and printed {out.splitlines()!r}, expected {want!r}"


def test_pyproject_script():
    """pyproject.toml declares the loganalyser console script"""
    with open("pyproject.toml", "rb") as f:
        project = tomllib.load(f).get("project", {})
    scripts = project.get("scripts", {})
    assert scripts.get("loganalyser") == "loganalyser.cli:main", f"[project.scripts] is {scripts!r}, expected loganalyser = \"loganalyser.cli:main\""
    assert project.get("requires-python"), "[project] should say which Python it needs (requires-python)"

Run the finished program

python -m venv .venv, activate it, then python -m pip install -e . and loganalyser samples/app.log --by source --format json -v (or python -m loganalyser samples/app.log without installing)
Reference solution

Try the milestones first. This solution passes every acceptance test and the type checker.

pyproject.toml

[build-system]
requires = ["setuptools >= 77.0.3"]
build-backend = "setuptools.build_meta"

[project]
name = "loganalyser"
version = "1.0.0"
description = "Count log entries per level, source or hour, and write CSV or JSON."
requires-python = ">= 3.14"
dependencies = []

[project.scripts]
loganalyser = "loganalyser.cli:main"

[tool.setuptools]
packages = ["loganalyser"]

loganalyser/__init__.py

"""Count log entries per level, source or hour, and write CSV or JSON."""

__version__ = "1.0.0"

loganalyser/__main__.py

"""python -m loganalyser runs the command line."""

import sys

from .cli import main

sys.exit(main())

loganalyser/errors.py

"""The exceptions of the log analyser; the command line turns each into exit status 1."""


class LogAnalyserError(Exception):
    """Base class: every error this package raises on purpose."""


class InputError(LogAnalyserError):
    """A log file is missing, is a folder, or cannot be read."""


class ParseError(LogAnalyserError):
    """A line does not have the expected log format."""

    def __init__(self, line_no: int, line: str, reason: str) -> None:
        super().__init__(f"line {line_no}: {reason}: {line!r}")
        self.line_no = line_no
        self.line = line
        self.reason = reason

loganalyser/model.py

"""The domain: log levels and parsed log entries."""

from dataclasses import dataclass
from datetime import datetime
from enum import IntEnum


class Level(IntEnum):
    """Log levels, ordered by severity so they compare: Level.ERROR > Level.INFO."""

    DEBUG = 10
    INFO = 20
    WARNING = 30
    ERROR = 40
    CRITICAL = 50

    @classmethod
    def parse(cls, text: str) -> "Level":
        """The level named by text, in any case; WARN and FATAL are accepted too."""
        name = text.strip().upper()
        name = {"WARN": "WARNING", "FATAL": "CRITICAL"}.get(name, name)
        try:
            return cls[name]
        except KeyError:
            raise ValueError(f"unknown log level {text!r}") from None


@dataclass(frozen=True, slots=True)
class Entry:
    """One parsed log line."""

    timestamp: datetime
    level: Level
    source: str
    message: str
    line_no: int = 0

loganalyser/parser.py

"""Turning lines of text into Entry objects, lazily."""

import logging
import re
from collections.abc import Iterable, Iterator
from datetime import datetime
from pathlib import Path

from .errors import InputError, ParseError
from .model import Entry, Level

logger = logging.getLogger(__name__)

# 2026-09-28 14:03:12 INFO  auth: user alice logged in
LINE_RE = re.compile(
    r"^(?P<timestamp>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2})\s+"
    r"(?P<level>[A-Za-z]+)\s+"
    r"(?P<source>[\w.-]+):\s+"
    r"(?P<message>.*?)\s*$"
)
TIMESTAMP_FORMAT = "%Y-%m-%d %H:%M:%S"


def parse_line(line: str, line_no: int = 0) -> Entry:
    """Parse one log line. Raise ParseError if it is not in the expected format."""
    match = LINE_RE.match(line)
    if match is None:
        raise ParseError(line_no, line.rstrip("\n"), "not a log line")
    try:
        timestamp = datetime.strptime(match["timestamp"], TIMESTAMP_FORMAT)
        level = Level.parse(match["level"])
    except ValueError as err:
        raise ParseError(line_no, line.rstrip("\n"), str(err)) from err
    return Entry(timestamp, level, match["source"], match["message"], line_no)


def parse_lines(lines: Iterable[str], *, source: str = "<input>", strict: bool = False) -> Iterator[Entry]:
    """Yield an Entry per log line. Blank lines are ignored.

    A bad line raises ParseError when strict, and is logged as a warning and
    skipped otherwise.
    """
    for line_no, line in enumerate(lines, start=1):
        if not line.strip():
            continue
        try:
            yield parse_line(line, line_no)
        except ParseError as err:
            if strict:
                err.add_note(f"in {source}")
                raise
            logger.warning("%s:%d: skipped (%s)", source, line_no, err.reason)


def read_entries(paths: Iterable[Path], *, strict: bool = False) -> Iterator[Entry]:
    """Yield the entries of each file in turn, reading one line at a time."""
    for path in paths:
        if not path.is_file():
            raise InputError(f"{path}: no such file")
        logger.info("reading %s", path)
        try:
            with path.open(encoding="utf-8") as lines:
                yield from parse_lines(lines, source=str(path), strict=strict)
        except (OSError, UnicodeDecodeError) as err:
            raise InputError(f"{path}: cannot read it ({err})") from err

loganalyser/report.py

"""Filtering, counting and writing the results as CSV or JSON."""

import csv
import json
from collections import Counter
from collections.abc import Iterable, Iterator
from dataclasses import dataclass, field
from datetime import datetime
from enum import Enum
from typing import TextIO

from .model import Entry, Level


class GroupBy(Enum):
    """What to count entries by."""

    LEVEL = "level"
    SOURCE = "source"
    HOUR = "hour"

    def key(self, entry: Entry) -> str:
        """The group an entry falls into, as text."""
        match self:
            case GroupBy.LEVEL:
                return entry.level.name
            case GroupBy.SOURCE:
                return entry.source
            case GroupBy.HOUR:
                return entry.timestamp.strftime("%Y-%m-%d %H:00")


def within(entries: Iterable[Entry], since: datetime | None = None, until: datetime | None = None) -> Iterator[Entry]:
    """The entries at or after since and before until (either may be None)."""
    for entry in entries:
        if since is not None and entry.timestamp < since:
            continue
        if until is not None and entry.timestamp >= until:
            continue
        yield entry


def at_least(entries: Iterable[Entry], level: Level) -> Iterator[Entry]:
    """The entries whose level is level or more severe."""
    return (entry for entry in entries if entry.level >= level)


@dataclass
class Report:
    """Counts per group, plus the size and time span of what was counted."""

    by: GroupBy
    counts: list[tuple[str, int]] = field(default_factory=list)
    entries: int = 0
    first: datetime | None = None
    last: datetime | None = None

    def to_dict(self) -> dict[str, object]:
        """A JSON-ready dict."""
        return {
            "by": self.by.value,
            "entries": self.entries,
            "first": self.first.isoformat(sep=" ") if self.first else None,
            "last": self.last.isoformat(sep=" ") if self.last else None,
            "counts": [{self.by.value: key, "count": count} for key, count in self.counts],
        }


def _order(by: GroupBy, counter: Counter[str]) -> list[tuple[str, int]]:
    match by:
        case GroupBy.LEVEL:
            return sorted(counter.items(), key=lambda item: Level[item[0]])
        case GroupBy.SOURCE:
            # Most common first; ties by name, so the output is stable.
            return sorted(counter.items(), key=lambda item: (-item[1], item[0]))
        case GroupBy.HOUR:
            return sorted(counter.items())


def summarise(entries: Iterable[Entry], by: GroupBy) -> Report:
    """Count the entries per group in one pass over them."""
    counter: Counter[str] = Counter()
    report = Report(by)
    for entry in entries:
        counter[by.key(entry)] += 1
        report.entries += 1
        if report.first is None or entry.timestamp < report.first:
            report.first = entry.timestamp
        if report.last is None or entry.timestamp > report.last:
            report.last = entry.timestamp
    report.counts = _order(by, counter)
    return report


def write_csv(report: Report, out: TextIO) -> None:
    """One header row (the group name and count), then a row per group."""
    writer = csv.writer(out, lineterminator="\n")
    writer.writerow([report.by.value, "count"])
    writer.writerows(report.counts)


def write_json(report: Report, out: TextIO) -> None:
    """The report as an indented JSON object."""
    json.dump(report.to_dict(), out, indent=2)
    out.write("\n")

loganalyser/cli.py

"""The command line: loganalyser [options] FILE [FILE ...]."""

import argparse
import logging
import sys
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import datetime
from pathlib import Path
from typing import TextIO

from . import __version__
from .errors import LogAnalyserError, ParseError
from .model import Level
from .parser import read_entries
from .report import GroupBy, at_least, summarise, within, write_csv, write_json

logger = logging.getLogger("loganalyser")

EXIT_OK = 0
EXIT_ERROR = 1  # bad input: a missing file, or a bad line with --strict
# argparse itself exits with status 2 on a usage error.


def when(text: str) -> datetime:
    """An argparse type: 2026-09-28, 2026-09-28T14:00 or "2026-09-28 14:00"."""
    try:
        return datetime.fromisoformat(text)
    except ValueError:
        raise argparse.ArgumentTypeError(f"not an ISO date or date and time: {text!r}") from None


def level(text: str) -> Level:
    """An argparse type: a level name such as warning or ERROR."""
    try:
        return Level.parse(text)
    except ValueError as err:
        raise argparse.ArgumentTypeError(str(err)) from None


def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(
        prog="loganalyser",
        description="Count log entries per level, source or hour, and write the counts as CSV or JSON.",
    )
    parser.add_argument("files", nargs="+", type=Path, metavar="FILE", help="log files to read, in order")
    parser.add_argument("--by", choices=[g.value for g in GroupBy], default=GroupBy.LEVEL.value, help="what to count by (default: level)")
    parser.add_argument("--format", choices=["csv", "json"], default="csv", help="output format (default: csv)")
    parser.add_argument("-o", "--output", type=Path, help="write to this file instead of standard output")
    parser.add_argument("--since", type=when, help="only entries at or after this time")
    parser.add_argument("--until", type=when, help="only entries before this time")
    parser.add_argument("--level", type=level, default=Level.DEBUG, help="only entries at this level or above")
    parser.add_argument("--strict", action="store_true", help="stop at the first line that cannot be parsed")
    parser.add_argument("-v", "--verbose", action="count", default=0, help="log progress (-vv for debug)")
    parser.add_argument("--version", action="version", version=f"%(prog)s {__version__}")
    return parser


def configure_logging(verbosity: int) -> None:
    """Progress and warnings go to standard error, so they never mix with the output."""
    handler = logging.StreamHandler(sys.stderr)
    handler.setFormatter(logging.Formatter("%(levelname)s: %(message)s"))
    logger.handlers[:] = [handler]
    logger.setLevel(logging.WARNING if verbosity == 0 else logging.INFO if verbosity == 1 else logging.DEBUG)
    logger.propagate = False


@contextmanager
def open_output(path: Path | None) -> Iterator[TextIO]:
    """Standard output, or the file at path, closed afterwards."""
    if path is None:
        yield sys.stdout
        return
    with path.open("w", encoding="utf-8", newline="") as out:
        yield out
    logger.info("wrote %s", path)


def main(argv: list[str] | None = None) -> int:
    """Run the tool with argv (default: sys.argv[1:]); return the exit status."""
    args = build_parser().parse_args(argv)
    configure_logging(args.verbose)
    if args.since and args.until and args.since >= args.until:
        logger.error("--since must be before --until")
        return EXIT_ERROR
    by = GroupBy(args.by)
    try:
        entries = read_entries(args.files, strict=args.strict)
        entries = at_least(within(entries, args.since, args.until), args.level)
        report = summarise(entries, by)
        logger.info("counted %d entries by %s", report.entries, by.value)
        with open_output(args.output) as out:
            (write_json if args.format == "json" else write_csv)(report, out)
    except ParseError as err:
        logger.error("%s (%s)", err, "; ".join(getattr(err, "__notes__", [])))
        return EXIT_ERROR
    except LogAnalyserError as err:
        logger.error("%s", err)
        return EXIT_ERROR
    except OSError as err:
        logger.error("cannot write the output: %s", err)
        return EXIT_ERROR
    return EXIT_OK


if __name__ == "__main__":
    sys.exit(main())

samples/app.log

2026-09-28 08:59:58 INFO  web: server started on port 8080
2026-09-28 09:00:01 INFO  auth: user alice logged in
2026-09-28 09:00:07 DEBUG db: pool size 5
2026-09-28 09:02:13 WARN  web: slow response for /reports (2.4 s)
2026-09-28 09:15:42 ERROR db: connection lost, retrying
2026-09-28 09:15:43 INFO  db: reconnected
this line was cut off by a crash
2026-09-28 10:01:09 INFO  auth: user bob logged in
2026-09-28 10:05:30 ERROR web: 500 on /export (KeyError: 'month')
2026-09-28 10:05:31 WARNING auth: 3 failed logins for carol

2026-09-28 11:20:00 CRITICAL db: disk full
2026-09-28 11:20:05 INFO  web: server stopping

Take it further

  • Read .gz files too, choosing gzip.open or Path.open from the suffix, so rotated logs can be analysed without unpacking them.
  • Add --top N to keep only the N largest groups, using itertools.islice over the ordered counts.
  • Add a --pattern REGEX option that keeps only entries whose message matches, and report how many lines it dropped.
  • Accept a second log format, such as the common web server access log, by trying several compiled patterns in turn.
  • Rewrite the acceptance tests as a unittest.TestCase, using assertLogs for the warnings and unittest.mock to fake a file that cannot be read.

Projects are practice: your checks run in your browser or on your computer and never count toward a certificate.