How it works
The pipeline from camera frame to notification, why alerting is deliberately delayed, and how the model performs on a held-out test set.
Pipeline
- OctoPrint captures a timelapse frame
- The frame is cropped to a centred square
- The crop is resized to 224 × 224 pixels
- MobileNet runs through the TFLite interpreter
- A single “failed” probability comes out, between 0 and 1
- The probability is compared against the threshold, 0.8 by default
- An alert fires only once 3 of the last 5 frames are over the threshold
- An optional ntfy push goes out with the offending frame attached
The same preprocessing code path runs during training, in the plugin and in the hosted API, so a frame scored locally and the identical frame sent to the API produce the same number.
Why a sliding window
Single frames are noisy. A print can look fine one frame and ambiguous the next, and a nozzle passing through the shot is enough to flip a borderline score. Requiring 3 of the last 5 frames to cross the threshold removes almost all of that noise, at the cost of a few seconds of extra latency. A single good frame in the middle of a failure sequence also will not reset the alert.
The Attention page shows a real pair of consecutive frames where the prediction flips, which is what motivated this design.
Model performance
The model returns a probability rather than a hard label, so the operating point is a deployment decision rather than a property of the model. The ROC curve below shows the trade-off available on the held-out test set.
The plugin default of 0.8 deliberately favours specificity: a monitoring tool
that cries wolf gets switched off, so a missed failure costs less than a false alarm.
Training data
The classifier is trained on tens of thousands of crowdsourced print images rather than a single controlled camera rig. A model fitted to one printer in one room can reach excellent validation accuracy and still collapse in someone else's workshop, because it has learned the background rather than the print.
- Many printer models, bed surfaces and enclosure types
- Webcams mounted at different distances and angles
- Daylight, LED strips and near-dark rooms
- Filament colours that blend into the bed as well as ones that contrast with it
Test images are held out by source, not sampled at random, so frames from a printer seen during training never appear in the evaluation set.
Inference cost
| Hardware | Inference time | Details |
|---|---|---|
| Raspberry Pi 5 | ~40 ms | MobileNetV2 float32, multi-thread TFLite runtime |
| Raspberry Pi Zero 2 W | ~480 ms | Single thread; well within normal timelapse capture intervals |
Because scoring happens once per captured frame rather than per video frame, even the slowest supported board spends a small fraction of its CPU on detection.