Lesson 5 of 5 · 26 min
UART, DMA and debugging
A serial port is the oldest and still most useful window into a running program, and DMA is how you stop paying the CPU for moving bytes. This lesson ties both to the third tool a Cortex-M developer cannot live without: the debugger. Everything uses the NUCLEO-L476RG, where USART2 (PA2 as TX, PA3 as RX, alternate function 7) is wired through the ST-Link to a virtual COM port on your PC.
UART baud-rate maths
UART has no clock wire. Both sides agree on a bit time in advance and the receiver re-synchronises on each start bit. The transmitter on the STM32 divides its kernel clock by a value in the BRR register. With the default 16x oversampling:
baud = f_clk / USARTDIV, and BRR holds USARTDIV as an integer
With USART2 clocked from PCLK1 at 80 MHz and a target of 115200:
USARTDIV = 80,000,000 / 115,200 = 694.44, so BRR = 694
The actual baud rate is 80,000,000 / 694 = 115,274, an error of (115,274 - 115,200) / 115,200 = 0.064 percent. The receiver samples near the middle of each bit and the error accumulates over a 10-bit frame (start, 8 data, stop), so anything below about 2 percent is comfortable. This is why some clock and baud combinations are better than others: at 72 MHz, 115200 divides to exactly 625 with no error, while at 80 MHz it leaves a remainder.
The 8N1 frame has 10 bits for 8 bits of data, so throughput at 115200 baud is 11,520 bytes per second, 86.8 microseconds per byte.
RCC->AHB2ENR |= RCC_AHB2ENR_GPIOAEN;
RCC->APB1ENR1 |= RCC_APB1ENR1_USART2EN;
// PA2 and PA3: alternate function mode (10), AF7
GPIOA->MODER = (GPIOA->MODER & ~((3u << 4) | (3u << 6))) | (2u << 4) | (2u << 6);
GPIOA->AFR[0] = (GPIOA->AFR[0] & ~((0xFu << 8) | (0xFu << 12))) | (7u << 8) | (7u << 12);
USART2->BRR = 694; // 80 MHz / 115200
USART2->CR1 = USART_CR1_TE | USART_CR1_RE | USART_CR1_UE;
void uart_putc(char c) {
while (!(USART2->ISR & USART_ISR_TXE)) {} // wait for the transmit register to empty
USART2->TDR = (uint8_t)c;
}
Why DMA
Sending 64 bytes with uart_putc keeps the CPU spinning for 64 x 86.8 microseconds = 5.56 ms. At 80 MHz that is about 445,000 clock cycles spent waiting. The DMA (Direct Memory Access) controller is a small second engine on the bus that moves data between memory and a peripheral without involving the core. You describe the job once, then the core is free; DMA raises an interrupt when it is finished.
The job has four parts: a peripheral address, a memory address, a count, and a direction. On the L476, the request routing is selectable per channel; the reference manual's DMA request table says USART2_TX is on DMA1 channel 7, request 2.
static const char msg[] = "Hello from DMA\r\n";
RCC->AHB1ENR |= RCC_AHB1ENR_DMA1EN;
DMA1_CSELR->CSELR = (DMA1_CSELR->CSELR & ~(0xFu << 24)) | (2u << 24); // channel 7 = request 2
DMA1_Channel7->CPAR = (uint32_t)&USART2->TDR; // peripheral address, fixed
DMA1_Channel7->CMAR = (uint32_t)msg; // memory address, increments
DMA1_Channel7->CNDTR = sizeof(msg) - 1; // number of bytes
DMA1_Channel7->CCR = DMA_CCR_MINC // step through memory
| DMA_CCR_DIR // memory to peripheral
| DMA_CCR_TCIE; // interrupt when done
USART2->CR3 |= USART_CR3_DMAT; // UART requests a byte when TDR is empty
DMA1_Channel7->CCR |= DMA_CCR_EN; // go
Configuring this costs a handful of stores, tens of cycles, instead of 445,000. With HAL the same thing is HAL_UART_Transmit_DMA(&huart2, buf, len). The buffer must stay valid and unchanged until the transfer completes: DMA reads it long after your function returned, so never pass a local array that goes out of scope.
Debugging with SWD
Printing is a poor first tool. A real debugger over SWD (STM32CubeIDE, or OpenOCD with GDB in VS Code) lets you:
- Halt the core and inspect all registers, variables and memory, including peripheral registers. Compare the GPIO or TIM2 registers with what you intended.
- Set hardware breakpoints. The Cortex-M4 has a handful (6 instruction comparators in the FPB), so breakpoints work in flash without rewriting code.
- Set watchpoints on a variable (the DWT unit), halting when anything writes it. This finds memory corruption in minutes.
- Single-step, and see the call stack.
Build with -Og or -O0 plus -g while debugging; at -O2 variables are optimised into registers and lines are reordered, which is confusing.
Reading a fault
When code does something illegal, such as dereferencing a bad pointer, executing at an invalid address, or dividing by zero when trapping is enabled, the core takes a fault. Unhandled it escalates to HardFault, and the default handler is an endless loop that looks like a freeze. The fault status registers tell you why:
| Register | Address | Contents |
|---|---|---|
| CFSR | 0xE000 ED28 | Memory, bus and usage fault status bits |
| HFSR | 0xE000 ED2C | Bit 30 FORCED means a lesser fault escalated |
| BFAR | 0xE000 ED38 | Offending address, valid if CFSR bit 15 is set |
Suppose you write through a bad pointer to a reserved address such as 0x3000 0000. In CFSR, bit 9 (PRECISERR) and bit 15 (BFARVALID) are set, which reads 0x0000 8200, and BFAR holds 0x3000 0000. PRECISERR means the faulting instruction is exactly the stacked PC. The stacked frame (R0 to R3, R12, LR, PC, xPSR) is at the stack pointer, so PC is word 6.
__attribute__((naked)) void HardFault_Handler(void) {
__asm volatile (
"tst lr, #4 \n" // which stack was in use?
"ite eq \n"
"mrseq r0, msp \n"
"mrsne r0, psp \n"
"b hard_fault_c \n"
);
}
void hard_fault_c(uint32_t *frame) {
volatile uint32_t pc = frame[6]; // address of the faulting instruction
volatile uint32_t cfsr = SCB->CFSR;
volatile uint32_t bfar = SCB->BFAR;
(void)pc; (void)cfsr; (void)bfar;
__BKPT(0); // debugger stops here; read the locals
for (;;) {}
}
Look up pc in the map file or disassembly and you have the exact line.
printf over ITM or UART
Redirect the C library's output by defining _write. Over UART, reuse the code above:
int _write(int fd, char *ptr, int len) {
for (int i = 0; i < len; i++) uart_putc(ptr[i]);
return len;
}
The alternative is ITM, the trace unit in the core, which sends bytes out of the SWO pin to the ST-Link at several megabits per second without a UART or CPU waiting on a slow line. Replace the loop body with ITM_SendChar(ptr[i]) and enable SWV in your debugger's configuration with the correct core clock. Check the board's user manual for the solder bridge that connects SWO on your revision. ITM_SendChar does nothing when no debugger is attached, so leaving it in costs almost nothing. Floating-point printf needs the linker option -u _printf_float, and a newline \n flushes line-buffered output.
Check yourself
With f_clk = 80 MHz, what BRR value gives roughly 115200 baud?
Check yourself
A HardFault occurs and CFSR reads 0x0000 8200. What is the most likely cause?