Lesson 3 of 5 · 18 min
Reading a codebase
Open the repository of a mature robotics project for the first time and the feeling is usually the same: thousands of files, unfamiliar folder names, and no obvious place to begin. Professional engineers feel it too. What separates them from beginners is not that they understand the whole thing; nobody does. They have a method for finding the small part that matters for the task at hand and ignoring the rest. This lesson gives you that method.
Why reading comes before writing
Most accepted contributions are small, and most of the effort in a small contribution is finding the right place. The edit itself is often three lines. If you can navigate quickly, you can fix a typo in a driver, correct a wrong pin in an example, or add a missing check, which are exactly the changes maintainers are glad to receive. If you cannot, you will either give up or send a change in the wrong place.
Reading code is also a skill with a different flavour from writing it. You are not trying to understand every line. You are trying to build a map and then answer one specific question with it.
Step 1: start with the README
The README tells you what the project is for and often links to the rest: documentation, a community forum, a list of supported boards. Read it for three answers:
- What problem does this project solve, and what does it explicitly not do?
- Which hardware and operating systems are supported?
- Where does the longer documentation live?
If the project has a docs folder or a documentation website, skim its table of contents. Architecture overviews, when they exist, are worth more than any amount of grepping.
Step 2: get it to build
Code you cannot build is code you cannot trust. Find the build instructions, usually in the README or a BUILDING or docs page, and follow them exactly, on a clean checkout. Typical shapes in this field are a make or cmake flow for firmware, a colcon workspace for ROS 2 style projects, west for Zephyr style projects, or a pip install for Python tooling. Do not memorise these; the project tells you which one it uses.
Write down every step that failed and how you fixed it. Build instructions rot as dependencies change, and an honest note of "step 3 fails on a fresh install unless you also do this" is itself a valuable contribution, which you will meet again in lesson 4.
Step 3: run the tests
Tests are the most honest documentation in a repository, because unlike prose they fail when they go out of date. Find the test folder and look at how a test is written. A good test shows you, in a few lines:
- the inputs a function expects,
- the outputs it promises,
- and the edge cases the authors worried about.
Run the full suite once, so you know the baseline: if tests fail before you change anything, you must not later blame yourself for them. Then read two or three tests near the area you care about. You learn the intended behaviour faster than from the implementation.
Step 4: map the tree
Look at the top-level folders and ask what each is for. Without reading any file, you can usually guess a structure like this:
| Folder name pattern | Usually contains |
|---|---|
src, lib | The main implementation |
include | Public headers for C and C++ |
drivers, boards, hal | Hardware-specific code, one folder per chip or board |
examples, samples | Small programs showing how to use things |
tests, test | Automated tests |
docs | Documentation sources |
tools, scripts | Build and helper scripts |
The drivers and boards pattern matters for robotics. Good projects separate code that is the same on every robot from code that touches a specific chip, so that one control algorithm runs on many boards. This layer is often called a hardware abstraction layer. Knowing it exists tells you where a bug lives: if it appears only on one board, look in that board's folder.
Step 5: search, do not browse
You cannot read thousands of files, so search. Your editor's project-wide search works, and so do these command-line tools, which are fast on large trees.
# Where is this symbol defined or used?
rg -n "set_motor_speed" src/
# Search only C and C++ files, ignore case
rg -n -i "watchdog" --glob "*.c" --glob "*.h"
# The same idea with plain git, searching tracked files
git grep -n "PWM_FREQUENCY"
# Which commit introduced a string? Often explains why it exists
git log -S"PWM_FREQUENCY" --oneline
# Who last changed each line of a file, and in which commit
git blame -L 40,80 src/motor.c
rg is ripgrep, a fast search tool you install separately; git grep ships with Git. Searching for error messages, configuration option names and function names you saw in the docs are the three most productive starting points, because they are unique strings.
Step 6: follow one feature from entry point to hardware
Now put the map to work with a method that works on any codebase. Pick one behaviour and trace it all the way down, ignoring everything else. Say you want to understand how a motor command reaches the motor.
- Find the entry point. For a firmware project that is a
mainfunction or a startup file; for a ROS 2 style project it is a node's main or launch file. Searching formain(is a fine start. - Find where your behaviour is triggered. Search for a name you know from the docs: a command name, a message topic, a configuration option, a menu string.
- Follow the calls downward. At each function, ask only: "what does it call next that relates to my behaviour?" Skip logging, error handling and unrelated branches on the first pass.
- Watch the layers change. You will usually move from application logic, to a more general control layer, to a driver interface, to the code that finally writes a register or sets a pin or sends bytes over a bus.
- Stop at the hardware call. When you reach the line that touches the chip, you have the whole path. Write it down as a short chain of function names with file paths.
- Confirm with a test or a print. Add a temporary log line or run a test that exercises the path, to check your map matches reality.
You now understand one vertical slice of the project completely, and that is usually all a first contribution needs. Each trace after this one is faster, because the layers repeat.
Check yourself
You want to know why a particular line of firmware looks odd. Which approach gives you the author's own reasoning most directly?
Check yourself
Why are tests considered the most reliable documentation in a repository?