Introduction to Regular Expressions (RegEx)
Learning outcomes
- Understand the basic syntax and functionality of regular expressions (regex) for pattern matching.
- Explore the use of special characters, ranges, and anchors in regex to match specific patterns within text.
- Apply regex to search, extract, and manipulate data in various formats using practical examples.
- Use regular expressions to navigate and organize files within the filesystem.
Imagine you’re working as a data scientist for an e-commerce company, and you have been given a dataset of user reviews for a product. The reviews contain noisy data: some reviews include email addresses, phone numbers, and URLs, which you want to remove before conducting sentiment analysis. Additionally, you need to extract specific pieces of information, such as any mentions of product model numbers that follow a specific pattern.
Would you use ⌘ + f (Mac) or Ctrl + f (Windows / Linux) to find and replace what you needed?
Introduction
Like with most things, the best way for you to learn Regex is to get practice using it. There are a few exercises included in the notebook, and at the end I have also included links interactive online exercises with are great to practice your regular expressions (RegEx)!
To see what a particular regex is matching and how, you can use one of these two webpages, which both do a great job visualizing and explaining the different parts of a regex match:
- https://regexr.com/
regexrinterprets text input as one big string by default, so you need to check “multiline” under “flags” (top right) for it to behave as expected with beginning and end of line matches (it hints at this in the output for both ^ and $).
- https://regex101.com/
- regex101 has the “multiline” flag set by default.
Basic matching
Basic matching: if you look for a regular string, like banana, regex will match the exact string (including its upper/lower case). Both JupyterLab and VS Code have built in regex functionality (bring up the search box and click the .* symbol to use regex rather than the default search). When learning regex it is helpful to use one of the two web tools mentioned in the previous cell in order to visualize how your regex is matching the text. We will use a list of fruits to learn about regex.
Fruit list
applesas
apple
apricot
banana
bilberry
blackberry
blackcurrant
blood orange
blueberry
canary melon
cantaloupe
cherry
clementine
cloudberry
coconut
cranberry
cucumber
currant
dragonfruit
durian
elderberry
gooseberry
grape
grapefruit
papaya
passionfruit
peach
orange
oranges unripe
persimmon
pineapple
pomegranate
pomelo
purple mangosteen
rock melon
salal berry
satsuma
star fruit
strawberry
watermelonThe square brackets: []
If you want to specify the set of possible characters you can use square brackets []; For example, [Aa]pple would match Apple and apple.
Find all the pairs of vowels in the fruit list.
Highlight the black box below to see than correct answer (the black box will not show up on GitHub, so download the notebook unless you want the answer displayed) Remember to use one of the websites linked above to help you understand what your regex is matching (https://regexr.com/ or https://regex101.com/).
Pattern
[aeiou][aeiou]Ranges within []
You can also define ranges when using brackets. For example:
- `[A-Z]`: will match any upper case letter
- `[a-z]`: will match any lower case letter
- `[0-9]`: will match any digit
- `[0-5]`: will match any digit between 0 and 5
The order cannot be reversed, [z-A] does not work. You can combine ranges: [A-Za-z].
You can use square brackets starting with a caret. For example:
- `[^A-Z]`: will match anything that is not an upper case letter
- `[^0-9]`: will match anything that is not a digit
The caret needs to be inside the bracket, if it is outside it will match the beginning of a line as described under the “Anchors” section below.
These ranges are ordered based on ASCI codes where every character is represented by a number. The first character in the list is (space) and the last is ~ (tilde). The full list is shown below:
Output
!"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~
Special matching characters
A common operation is to match any character (e.g. between two important characters). Instead of writing out the full range [ -~] (space to tilde), the special character . can be used to match any character in the list above. . does not match the newline character, So if you have an expression that continues on the next line it will not be matched.
To match a literal . (the period character), you can “escape” its special meaning by prefacing it with a backslash \. (most common) or surrounding it with square brackets [.].
Another useful special character is \w, which matches any character that normally occurs inside a word (so it does not match spaces, underlines, etc)
What is the difference between writing [A-Za-z] and [A-z]?
[A-z] will also match the characters [/]^_, as you can see in the list above.
Match any characters between two _.
Pattern
_.*_Anchors
The caret outside the brackets means beginning of line. For example, ^apple will match all lines that start with apple, including apple sauce and apples. The dollar sign $ means end of line, e.g., fruit$ will match lines that end with fruit. To remember this, you can use the mnemonic “Start with power (^) and end with money ($)” (originally from Jenny Bryan).
Another useful anchor is \b, which matches end of word.
Write a regex that will match a line that contains only pineapple. (Hint: you cannot just write pineapple - it will not work - why?)
Pattern
^pineapple$note that if you use just pineapple, lines that also contain other words would match too.
Repetitions
To match multiple of the same character, you can either repeat it or use the following syntax:
{n}: exactlynoccurrences{n,}: at leastnoccurrences{0,m}: at mostmoccurrences{n,m}: betweennandm(inclusive) occurrences
Special repetition characters
There are some shortcuts for the most common repetitions:
?: means 0 or 1 time ({0,1})*: means 0 or more time ({0,})+: means 1 or more time ({1,})
For example, apples? will match apple and apples. But apples+ will not match apple or appplesq, but it will match apples, appless, applesss, etc.
Find the fruits with names between 10 and 12 characters.
Pattern
.{10,12}Find the lines with no more than 4 letters.
Pattern
^.{0,4}$Find all the words that contain at least two consecutive vowels.
Pattern
[aeiou]{2,}Pattern
[aeiou][aeiou]+This is a bit harder and derives from all previous sections: Match entire words that end in _.
Pattern
\\w*_\\bGo through the interactive tutorials and practice sessions at https://regexone.com/ that correspond to the topics we have covered during class.
The Library Carpentry organization has many regex exercises in all sections of their regex course here: https://librarycarpentry.org/lc-data-intro/ (you can just to do the exercises).
Cheatsheet
Practice data: phone numbers
Copy this data if you want to follow along.
Phone number data generated by Claude Opus 5.5 on 2026-09-24
phone_numbers.csv
phone_number
"555-0123"
"5550123"
"(604) 555-0142"
"604-555-0142"
"604.555.0199"
"604 555 0187"
"6045550173"
"+1 604 555 0164"
"+1 (416) 555-0110"
"1-416-555-0125"
"1 (212) 555-0156"
"+12125550138"
"001 212 555 0147"
"212/555-0191"
"(212)555-0108"
" (778) 555 - 0119 "
"+1-778-555-0177"Building up a phone number regex
Each row adds one new piece to the regex in the row above it, so that it matches one more way of writing a phone number. The examples show the text each row picks out of the practice data. Earlier rows often match only part of a longer number (for example, row 5 finds 555.0199 inside 604.555.0199). Those partial matches are why the anchors in row 11 matter.
Row 11 onward anchors the pattern to the start and end of the line. To try those rows, paste the numbers without the surrounding " quotes.
| What we’re matching | New piece | Example matches | Regex | |
|---|---|---|---|---|
| 1 | One exact number | literal characters | 555-0123 |
555-0123 |
| 2 | Any 7-digit number with a dash | ranges [0-9] |
555-0123, 555-0142 |
[0-9][0-9][0-9]-[0-9][0-9][0-9][0-9] |
| 3 | The same thing, written shorter | repetition {n} |
555-0123, 555-0142 |
[0-9]{3}-[0-9]{4} |
| 4 | The dash is optional | optional ? |
5550123 |
[0-9]{3}-?[0-9]{4} |
| 5 | Dash, dot, or space as the separator | a set of characters [-. ] |
555.0199, 555 0187 |
[0-9]{3}[-. ]?[0-9]{4} |
| 6 | Add a 3-digit area code | chaining pieces together | 604-555-0142, 604.555.0199, 6045550173, 212/555-0191 |
[0-9]{3}[-. /]?[0-9]{3}[-. ]?[0-9]{4} |
| 7 | Area code in brackets | escaping \( \) |
(604) 555-0142, (212)555-0108 |
\(?[0-9]{3}\)?[-. /]?[0-9]{3}[-. ]?[0-9]{4} |
| 8 | A leading 1 or +1 country code |
grouping ( )?, escaping \+ |
+1 604 555 0164, 1-416-555-0125, +12125550138 |
(\+?1[-. ]?)?\(?[0-9]{3}\)?[-. /]?[0-9]{3}[-. ]?[0-9]{4} |
| 9 | 00 as the international prefix |
alternation | |
001 212 555 0147 |
(\+|00)?(1[-. ]?)?\(?[0-9]{3}\)?[-. /]?[0-9]{3}[-. ]?[0-9]{4} |
| 10 | Spaces around the separators | zero or more * |
(778) 555 - 0119 |
(\+|00)?(1[-. ]?)?\(?[0-9]{3}\)? *[-./]? *[0-9]{3} *[-.]? *[0-9]{4} |
| 11 | Only the phone number on the line | anchors ^ $ |
" (778) 555 - 0119 ", +1-778-555-0177 |
^ *(\+|00)?(1[-. ]?)?\(?[0-9]{3}\)? *[-./]? *[0-9]{3} *[-.]? *[0-9]{4} *$ |
| 12 | Tidy up [0-9] |
shorthand \d for a digit |
the same as row 11 | ^ *(\+|00)?(1[-. ]?)?\(?\d{3}\)? *[-./]? *\d{3} *[-.]? *\d{4} *$ |
| 13 | Bring back the 7-digit numbers | making the whole area-code part optional | every number in the practice data | ^ *((\+|00)?(1[-. ]?)?\(?\d{3}\)? *[-./]? *)?\d{3} *[-.]? *\d{4} *$ |
| 14 | Any whitespace, not just spaces | whitespace \s |
every number in the practice data, plus numbers with tabs | ^\s*((\+|00)?(1[-.\s]?)?\(?\d{3}\)?\s*[-./]?\s*)?\d{3}\s*[-.]?\s*\d{4}\s*$ |
The kitchen sink
The table stops at a regex that handles our practice data. For comparison, we asked Claude (Opus 5.5, on 2026-09-24) to write a regex that matches any valid North American telephone number, and to make the answer as complete as possible.
This is what it came up with:
Pattern
(?i)^\s*(?:(?:(?:\+|00)\s*)?1[-.\s]*)?(?:(?P<open>\()?(?P<area>(?![2-9]11)[2-9][0-8]\d)(?(open)\))\s*[-./]?\s*)?(?P<exchange>(?![2-9]11)[2-9A-Z][0-9A-Z]{2})\s*[-.]?\s*(?P<line>[0-9A-Z]{4})(?:\s*,?\s*(?:ext(?:ension)?\.?|x|\#)\s*(?P<extension>\d{1,6}))?\s*$On top of everything in the table, this regex also checks the rules for real phone numbers: area codes and exchanges start with 2–9, area codes never have 9 as their middle digit, and neither can be a service code like 911. It also allows letters (1-800-555-FILM), extensions (ext. 204, x12), and brackets only when both ( and ) are there.
It uses several pieces we have not covered: non-capturing groups (?:...), named groups (?P<area>...), lookaheads (?!...), a conditional (?(open)\)), and the ignore-case flag (?i). Some of these are specific to Python and PCRE, so choose the “Python” or “PCRE” flavour on https://regex101.com/ to try it.
You can use LLMs and the online tools to help explain what each part of the regular expression is doing.
References
[1] Kery, M. B., Radensky, M., Arya, M., John, B. E., & Myers, B. A. (2018, April). The story in the notebook: Exploratory data science using a literate programming tool. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (pp. 1-11).