Skip to content
aviral gupta

// I4.3 · ~35 min · Intermediate

Regular expressions with re

After this lesson you can find, check and pull apart text with regular expressions, and rewrite it with re.sub.

Lesson 3 of 6 in I4 The standard library for real programs

You will be able to

  • Pick search, match or fullmatch, and write patterns as raw strings with \d, \w, [ ], + and {n}
  • Pull out parts of a match with groups and named groups, and collect every match with findall
  • Rewrite text with re.sub, using group references such as \1 and \g<name>
  1. Warm-up · Activity 1 of 7

    Warm-up from the string methods: what does this print?

    print("2026-09-29".split("-"))
  2. Predict · Activity 2 of 7

    Predict before you read on: what does this print?

    import re
    
    text = "Order 42 shipped"
    print(re.match(r"\d+", text), re.search(r"\d+", text).group())
  3. Practice · Activity 3 of 7

    A German postcode is exactly five digits. Fill in the function so that "123456" and "a12345" are rejected.

    ok = re.____(r"\d{5}", code) is not None
    ok = re.(r"\d{5}", code) is not None
  4. Practice · Activity 4 of 7

    What does this print?

    print(len("\n"), len(r"\n"))
  5. Practice · Activity 5 of 7

    Match each pattern to what it matches.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. The pattern has two groups. What does findall return?

    import re
    
    print(re.findall(r"(\w+)@(\w+)\.com", "ada@mail.com, bo@web.com"))
  7. Apply · Activity 7 of 7

    Mini-task. The text is "Coffee 3.20 EUR, cake 4.50 EUR, juice 2.75 USD". Use re.findall with one group to collect the amounts in EUR only, as strings like "3.20". Then add them up and print the total as 7.70 EUR.

    Check your work against this list

Build it yourself

Read the worked example, then write the exercises. Your code runs in your browser or on your computer and is never uploaded.

Worked example

Reading a log with named groups

Each line of the log has a date, a level in capitals and a message. The program checks each line with fullmatch and a compiled pattern with named groups, skips lines that do not fit, and hides e-mail addresses with sub. At the end, findall collects numbers and the domains of the addresses.

main.py

import re

# One log line: a date, a level in capitals, then the message.
LINE = re.compile(r"(?P<date>\d{4}-\d{2}-\d{2}) (?P<level>[A-Z]+) (?P<message>.*)")
EMAIL = re.compile(r"[\w.]+@[\w.]+\.\w+")

log = """2026-09-28 INFO user ada@example.com logged in
2026-09-28 ERROR disk full on /dev/sda1
not a log line
2026-09-29 WARNING retry 3 of 5 for bo@example.org"""

for line in log.splitlines():
    found = LINE.fullmatch(line)
    if found is None:
        print("skipped:", line)
        continue
    message = EMAIL.sub("<email>", found["message"])
    print(found["date"], found["level"].lower(), "|", message)

print(re.findall(r"\d+", "retry 3 of 5"))
print(sorted(re.findall(r"@([\w.]+)", log)))

Run it with

python main.py

Output

2026-09-28 info | user <email> logged in
2026-09-28 error | disk full on /dev/sda1
skipped: not a log line
2026-09-29 warning | retry 3 of 5 for <email>
['3', '5']
['example.com', 'example.org']
  • re.compile builds a pattern once; LINE.fullmatch(line) works like re.fullmatch with that pattern.
  • found["date"] is the same as found.group("date").
  • fullmatch returns None for "not a log line", so the if skips it before any group is read.
  • With one group, findall returned just the domains, without the @.
Change it and run it

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Exercises

Exercise 1 of 2

Usernames and hashtags

Write is_valid_username(name): True when the whole name is 3 to 16 characters, starts with an ASCII letter, and otherwise holds only ASCII letters, digits and _. Then write hashtags(text), which returns the words after each #, without the #: "Loving #python and #regex_101!" gives ["python", "regex_101"].

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    match only checks the start: "ada!" starts well. fullmatch checks the whole name.

  2. Hint 2

    One letter first, then 2 to 15 more characters: [A-Za-z][A-Za-z0-9_]{2,15}.

  3. Hint 3

    Put the word in a group, r"#(\w+)", and findall returns only the group.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

import re


def is_valid_username(name: str) -> bool:
    return re.fullmatch(r"[A-Za-z][A-Za-z0-9_]{2,15}", name) is not None


def hashtags(text: str) -> list[str]:
    return re.findall(r"#(\w+)", text)


if __name__ == "__main__":
    print(is_valid_username("ada_99"), is_valid_username("ada!"))
    print(hashtags("Loving #python and #regex_101!"))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

import re


def is_valid_username(name: str) -> bool:
    return re.match(r"[A-Za-z]", name) is not None


def hashtags(text: str) -> list[str]:
    return re.findall(r"#\w+", text)


if __name__ == "__main__":
    print(is_valid_username("ada_99"), is_valid_username("ada!"))
    print(hashtags("Loving #python and #regex_101!"))

test_main.py

from main import hashtags, is_valid_username


def test_valid_names():
    """ada_99 and Bob are valid usernames"""
    got = is_valid_username("ada_99"), is_valid_username("Bob")
    assert got == (True, True), f"for ada_99 and Bob you returned {got!r}"


def test_invalid_names():
    """Too short, a digit first, a ! inside, or 17 characters are all refused"""
    names = ["ab", "9lives", "ada!", "a" * 17]
    got = [is_valid_username(name) for name in names]
    assert got == [False, False, False, False], f"for {names!r} you returned {got!r}"


def test_hashtags():
    """Hashtags are returned without the #"""
    got = hashtags("Loving #python and #regex_101!")
    assert got == ["python", "regex_101"], f"hashtags returned {got!r}"


def test_no_hashtags():
    """A text without # gives []"""
    got = hashtags("no tags here")
    assert got == [], f"hashtags returned {got!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Exercise 2 of 2

Parse log lines and rewrite dates

Write parse_log_line(line). For a whole line like "2026-09-28 ERROR disk full" it returns {"date": "2026-09-28", "level": "ERROR", "message": "disk full"}; for any line that does not fit (a lower-case level, no message, text before the date) it returns None. Then write iso_to_german(text), which rewrites every date like 2026-09-30 as 30.09.2026 and leaves the rest alone.

Tab indents and Shift+Tab outdents. To leave the editor with the keyboard, press Esc, then Tab.

The first run downloads Python for your browser (up to 6.5 MB) and keeps it cached. Your code stays on your device.

Hints
  1. Hint 1

    Name the groups: (?P<date>\d{4}-\d{2}-\d{2}), then a space, (?P<level>[A-Z]+), a space and (?P<message>.+).

  2. Hint 2

    fullmatch rejects text before the date; if the result is None, return None, else found.groupdict().

  3. Hint 3

    For the dates, name year, month and day, and use r"\g<day>.\g<month>.\g<year>" as the replacement in re.sub.

Show a solution

One way to solve it. Yours can look different and still pass the checks.

import re

LINE = re.compile(r"(?P<date>\d{4}-\d{2}-\d{2}) (?P<level>[A-Z]+) (?P<message>.+)")


def parse_log_line(line: str) -> dict[str, str] | None:
    found = LINE.fullmatch(line)
    if found is None:
        return None
    return found.groupdict()


def iso_to_german(text: str) -> str:
    return re.sub(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", r"\g<day>.\g<month>.\g<year>", text)


if __name__ == "__main__":
    print(parse_log_line("2026-09-28 ERROR disk full"))
    print(iso_to_german("Due 2026-09-30."))
Run it on your computer

Install Python 3.14 or newer. Save these files in one folder, open a terminal in that folder, and run the commands below.

main.py

import re


def parse_log_line(line: str) -> dict[str, str] | None:
    # Match the whole line with named groups date, level and message.
    return None


def iso_to_german(text: str) -> str:
    # Rewrite every 2026-09-30 as 30.09.2026.
    return text


if __name__ == "__main__":
    print(parse_log_line("2026-09-28 ERROR disk full"))
    print(iso_to_german("Due 2026-09-30."))

test_main.py

from main import iso_to_german, parse_log_line


def test_parse():
    """A log line becomes a dict with date, level and message"""
    got = parse_log_line("2026-09-28 ERROR disk full on /dev/sda1")
    want = {"date": "2026-09-28", "level": "ERROR", "message": "disk full on /dev/sda1"}
    assert got == want, f"parse_log_line returned {got!r}"


def test_reject():
    """Lines that do not fit give None"""
    lines = ["2026-09-28 error disk full", "2026-09-28 INFO", "x 2026-09-28 INFO hi"]
    got = [parse_log_line(line) for line in lines]
    assert got == [None, None, None], f"for {lines!r} parse_log_line returned {got!r}"


def test_german_dates():
    """Every date is rewritten as day.month.year"""
    got = iso_to_german("from 2026-09-01 to 2026-09-30")
    assert got == "from 01.09.2026 to 30.09.2026", f"iso_to_german returned {got!r}"


def test_no_dates():
    """Text without a date stays the same"""
    got = iso_to_german("no dates here")
    assert got == "no dates here", f"iso_to_german returned {got!r}"

On macOS and Linux, type python3 wherever these commands say python, as in the first lesson.

Run the program:

python main.py

Run the checks (needs learnrun.py in the same folder):

python learnrun.py test
Download learnrun.py

Common mistakes

Using the result without checking for None

import re

found = re.match(r"\d+", "Order 42")
print(found.group())

What Python prints

AttributeError: 'NoneType' object has no attribute 'group'

Why, and the fix

match, search and fullmatch return None when nothing matches, and None has no group(). Here match fails because the string starts with a letter; search would find 42. Always check: if found is None: before you read groups.

An unclosed parenthesis

import re

price = re.compile(r"(\d+\.\d{2}")

What Python prints

re.PatternError: missing ), unterminated subpattern at position 0

Why, and the fix

Every ( that opens a group needs its ). The position tells you where the unclosed group starts. To match a real parenthesis, escape it: r"\(" and r"\)".

Asking for a group that does not exist

import re

found = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", "due 2026-09")
print(found.group("day"))

What Python prints

IndexError: no such group

Why, and the fix

The pattern only names year and month, so there is no group day. Check the names in (?P<...>) against the names you read. Numbered groups count the opening parentheses from the left, starting at 1; group 0 is the whole match.

Python in the browser: Pyodide 314.0.7, MPL-2.0. Licence and source

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

Patterns and raw strings

A pattern describes text. \d is a digit, \w a letter, digit or underscore, \s whitespace, and . any character except a newline. [A-Z] is one character from a set. + means one or more of what comes before, * zero or more, ? optional, and {3} exactly three. To match a special character itself, escape it: \. is a real dot. Write patterns as raw strings, r"\d+": in a normal string, Python itself would read escapes such as \b (a backspace) before re ever sees them.

search, match and fullmatch

re.search finds the first match anywhere in the string. re.match only accepts a match that starts at the beginning, and re.fullmatch only one that covers the whole string, which is what you want to validate input. All three return a Match object, or None when nothing matches, so test with if found is None before you use it. found.group() is the matched text, and found.span() where it sits.

Groups, findall and sub

Parentheses make a group: in r"(\d+)-(\d+)", group(1) and group(2) are the two numbers, and groups() returns both. (?P<year>\d{4}) names a group: read it with found["year"], or all of them with groupdict(). re.findall returns every match as a list: whole matches without groups, the group with one group, and tuples with several. re.sub(pattern, replacement, text) replaces every match; in the replacement, \1 or \g<year> inserts what a group matched.

Sources

Last reviewed September 29, 2026