Technical details

How it works

The pipeline from camera frame to notification, why alerting is deliberately delayed, and how the model performs on a held-out test set.

Pipeline

  1. OctoPrint captures a timelapse frame
  2. The frame is cropped to a centred square
  3. The crop is resized to 224 × 224 pixels
  4. MobileNet runs through the TFLite interpreter
  5. A single “failed” probability comes out, between 0 and 1
  6. The probability is compared against the threshold, 0.8 by default
  7. An alert fires only once 3 of the last 5 frames are over the threshold
  8. An optional ntfy push goes out with the offending frame attached

The same preprocessing code path runs during training, in the plugin and in the hosted API, so a frame scored locally and the identical frame sent to the API produce the same number.

Why a sliding window

Single frames are noisy. A print can look fine one frame and ambiguous the next, and a nozzle passing through the shot is enough to flip a borderline score. Requiring 3 of the last 5 frames to cross the threshold removes almost all of that noise, at the cost of a few seconds of extra latency. A single good frame in the middle of a failure sequence also will not reset the alert.

The Attention page shows a real pair of consecutive frames where the prediction flips, which is what motivated this design.

Model performance

The model returns a probability rather than a hard label, so the operating point is a deployment decision rather than a property of the model. The ROC curve below shows the trade-off available on the held-out test set.

Receiver operating characteristic curve for the PrintSheriff failure classifier, evaluated on the held-out test set
ROC curve on the held-out test set. Raise the threshold to suppress false alarms on long unattended prints; lower it to catch failures earlier when someone is nearby to intervene.

The plugin default of 0.8 deliberately favours specificity: a monitoring tool that cries wolf gets switched off, so a missed failure costs less than a false alarm.

Training data

The classifier is trained on tens of thousands of crowdsourced print images rather than a single controlled camera rig. A model fitted to one printer in one room can reach excellent validation accuracy and still collapse in someone else's workshop, because it has learned the background rather than the print.

  • Many printer models, bed surfaces and enclosure types
  • Webcams mounted at different distances and angles
  • Daylight, LED strips and near-dark rooms
  • Filament colours that blend into the bed as well as ones that contrast with it

Test images are held out by source, not sampled at random, so frames from a printer seen during training never appear in the evaluation set.

Inference cost

Hardware Inference time Details
Raspberry Pi 5 ~40 ms MobileNetV2 float32, multi-thread TFLite runtime
Raspberry Pi Zero 2 W ~480 ms Single thread; well within normal timelapse capture intervals

Because scoring happens once per captured frame rather than per video frame, even the slowest supported board spends a small fraction of its CPU on detection.