Lesson 5 of 5 · 24 min
Running models on robots
Everything so far ran on a laptop with gigabytes of memory and a floating-point unit that never complains. The robot is different. Its controller may have 2 KB of RAM, no floating-point hardware and a motor that must be updated 20 times every second whether or not your model is finished thinking. Running a trained model on that device is called edge inference, and it is mostly a game of budgets: memory, time and accuracy. This lesson gives you the arithmetic to play it.
What the hardware gives you
Different boards live in different worlds, and the numbers decide which models are possible at all.
| Board | RAM | Flash (program + weights) | Clock |
|---|---|---|---|
| Arduino Uno (ATmega328P) | 2 KB | 32 KB | 16 MHz |
| Typical Cortex-M4 board | 64 KB | 256 KB | 80 MHz |
| ESP32 | 520 KB | 4 MB (typical module) | up to 240 MHz |
| Raspberry Pi 4 | 1 GB or more | SD card | 1.5 GHz |
Trained weights are constants, so they live in flash and are read in place. What needs RAM is the activations, the intermediate numbers each layer produces while one prediction is computed. Frameworks for small devices, such as TensorFlow Lite for Microcontrollers, reserve one fixed buffer for them, often called the tensor arena. There is no dynamic allocation at run time, which is exactly what you want on a robot.
Quantisation: float32 to int8
A float32 weight takes 4 bytes. An int8 weight takes 1, and integer multiplication is faster and cheaper than floating point on most small chips. Quantisation maps the real weights onto the 256 integers from -128 to 127. The simplest scheme uses a single scale for the whole tensor:
scale = max(abs(w)) / 127
q = round(w / scale), clamped to the range -128 to 127
restored = q * scale
The largest weight maps to exactly 127, and every other weight snaps to the nearest multiple of the scale. Rounding can be off by at most half a step, so the worst error is scale / 2, which is max(abs(w)) / 254, about 0.39 percent of the largest weight. Try it on eight weights.
The scale comes out as 0.010236, because the largest weight magnitude is 1.30 and 1.30 / 127 is 0.010236. That weight maps to -127 and is restored exactly. The worst error is 0.00488 (on 0.66, which becomes 64 and is restored as 0.6551), just under the theoretical bound of 0.00512. Memory drops from 32 bytes to 8, a factor of four.
Look at the 0.003 row, though. It rounds to 0, so the weight vanishes: a 100 percent relative error. Small weights are lost when one big weight sets the scale for everything. That is why practical toolchains often use a separate scale per output channel, and why you must always re-measure accuracy on your test set after quantising. A model that lost 0.1 percent accuracy is a bargain. One that lost 15 percent is not.
At run time, the multiply-accumulate also stays in integers. Multiply int8 weights by int8 inputs, add the products into a 32-bit accumulator, then convert once at the end by multiplying with the product of the two scales. Each product is at most 127 x 127 = 16,129, and a layer summing 256 of them reaches 4,129,024, far beyond the 16-bit range, which is why the accumulator must be 32-bit.
A worked memory budget
Suppose your model has 50,000 parameters, roughly a small network with a few hundred neurons per layer.
- As float32: 50,000 x 4 = 200,000 bytes, about 195 KiB.
- As int8: 50,000 x 1 = 50,000 bytes, about 49 KiB.
On the Arduino Uno with 32 KB of flash, neither version fits, even before counting your own program. On the Cortex-M4 board with 256 KB of flash, float32 technically fits but leaves only about 60 KiB for code, while int8 leaves about 207 KiB. On an ESP32 both fit with room to spare. You also need an activation buffer in RAM, and the model must fit that too. The conclusion for the Uno is not "give up" but "shrink the model": a network with 6,000 parameters in int8 needs about 6 KB of flash and runs comfortably there.
The latency budget at 20 Hz
A control loop at 20 Hz has a period of 1000 / 20 = 50 ms. Everything must happen inside that window, every time: read sensors, run the model, update motors, and any communication. Inference is one line item, not the whole budget. Estimating its cost starts from counting multiply-accumulates (MACs). A dense layer with n_in inputs and n_out outputs needs n_in * n_out MACs, so a model with 50,000 weights needs about 50,000 MACs per prediction. Assume a plain C implementation on an 80 MHz core takes about 4 cycles per MAC.
The 50,000-MAC model costs 2.50 ms: 4.00 + 2.50 + 2.00 + 3.00 = 11.50 ms used out of 50 ms, which leaves 38.50 ms or 77 percent as headroom. A model ten times bigger, 500,000 MACs, costs 25 ms and still fits, with 9 ms of other work added. At 5 million MACs inference alone takes 250 ms, which is five whole control periods, so the loop would run at 4 Hz at best. These estimates are rough, since memory access, interrupts and library overhead all add time, so always measure on the real board. Budget for the worst case, not the average, because a control loop that occasionally misses its deadline behaves worse than one that is steadily slow.
When not to use ML
A trained model is a black box with a failure mode that is hard to predict, so it has to earn its place. Prefer a plain rule or a PID controller when:
- you can write the rule on paper in five minutes, such as "stop if the distance sensor reads under 15 cm";
- you have almost no data, or cannot collect examples of the failure cases;
- the task is safety critical and needs a guarantee you can prove, not a score you can measure;
- the environment will change in ways your data did not cover, so the model will fail silently.
ML is worth its cost for perception, where hand-written rules are hopeless: recognising a spoken command, classifying a gesture from IMU data, spotting an object in an image. Even then, keep the model advisory. Put hard limits (current limits, an emergency stop, a speed cap) in ordinary code that the model cannot override.
Check yourself
A model has 80,000 parameters. How much flash do its weights need as int8, and as float32?
Check yourself
A robot control loop runs at 20 Hz. What is the maximum time one full cycle (sensing, model, actuation) can take?