The "auto-add from photos" feature is the capstone because it exercises every layer of this book at once β and because it's a queueing system, so you get to grade it against the real ones you've run. The user selects a batch of food photos; each becomes its own meal, at the time the photo was taken, with an AI description and estimate β and the user can close the tab the moment the upload finishes.
An AI vision call takes up to ~25 seconds. The
JavaScript chapter's event loop says the
browser can't block, and the backend
chapter's gunicorn config says a
request thread shouldn't sit occupied for a 20-photo batch either
(8 threads total, 60s timeout). Classic answer: make the upload cheap
and synchronous, move the slow work to a background consumer, let the
client poll for progress. That's precisely the shape of
autolog.py.
browser POST /api/auto-log worker thread (one)
β user picks N photos β β
β per photo: read EXIF time, β β
β upload ββββββββββββββββββββββββββββΆβ normalize to JPEG β
β β dedup check (below) β
β β write photo + sidecar JSON β
β β to the LOCAL spool dir βββββββΆ β loop:
β βββ 202 {queued: β¦} ββββββββββββββββ β re-read user row (still approved?)
β β β consume 1 AI use (cap)
β polls /api/auto-log/pending β β OpenAI vision estimate
β to render per-photo status β β put_photo + put_meal
β β β delete spool entry
β β β (on failure: retry; then degrade
β β β or dead-letter as *.json.dead)
The spool is just files on local disk: the JPEG plus a tiny sidecar
JSON holding user_id/date/time/attempts β deliberately no
meal content, with one earned exception: after a failed commit the
paid-for AI estimate (description included) is cached onto the sidecar
so the retry never re-bills. One daemon thread consumes it. It's started lazily
(ensure_worker) because gunicorn's
--preload imports the app before forking, and a thread
started at import time would be lost in the fork β a genuinely sneaky
deployment interaction. The kick happens from the auto-log routes
and from /log, which is also the crash-recovery
story: after a restart, the first page view revives the worker and it
finds whatever the spool still holds.
| Property | How Ndiro gets it |
|---|---|
| Durability | Spooled to disk before the 202 Accepted (the status that says "taken, not yet done") β but local disk, ephemeral on most hosts. A stated PoC trade: a redeploy may drop queued work, and the design accepts it rather than paying for SQS. |
| Idempotent producer | Re-uploading the same photos is safe: the EXIF date+minute is the dedup fingerprint, checked against both existing PHOTO meals (a range query over the batch's whole date window β clamped camera clocks drift days) and the queued spool. |
| At-most-once billing per handled retry | The AI cap is consumed per photo before the call, refunded on upstream failure, and the estimate is cached on the sidecar so handled commit retries never re-bill. (A process death in the window after the AI reply but before that cache write can still bill twice on restart β at-least-once delivery leaves exactly this residue.) |
| Poison messages | MAX_ATTEMPTS, then the sidecar is renamed *.json.dead β a dead-letter queue, kept for the operator. |
| Graceful degradation | Cap exhausted or retries spent? The meal is committed anyway with a placeholder description β the photo and its timestamp are never lost; only the AI garnish is. |
| Revocation race | The worker re-reads the user row before each commit; a rejected or deleted account's queued photos are discarded, never logged. Account deletion also purges the user's spool (drop_user) before the S3 wipe. |
A single-broker, single-partition, at-least-once queue with a
DLQ, producer-side dedup by natural key, and consumer-side quota
enforcement β the entire Kafka-consumer checklist, in ~320 lines of
standard library. The interesting engineering isn't any one
mechanism; it's that each one exists because a specific failure was
imagined first. Read autolog.py top to bottom and
you'll find the failure story attached to every design choice, and
the trade-offs it declines to pay for stated out loud.
Back in log.html, the batch UI shows the other end:
reading EXIF timestamps client-side, uploading sequentially with
per-photo status rows, then polling /api/auto-log/pending
so a user who does keep the tab open watches photos flip from
"queued" to done. Everything you learned lands here at once: file
inputs (HTML), status colors from theme
tokens (CSS), createElement
rendering (JavaScript), fetch with the
standard error ladder (the API chapter).
How do you test a background thread deterministically? You don't β
you make the thread optional. Tests set
autolog.WORKER_ENABLED = False and call
process_once() directly, single-stepping the consumer:
every retry, refund, dead-letter and race in
tests/test_m11_autolog.py is asserted without a sleep or a
flake. Separating "the loop" from "one step of the loop" for
testability is a pattern worth stealing in any language.
Read autolog.py with three questions: (1) where
exactly is the boundary between "the upload request" and "the
worker" β which failure before it loses work, and which after it
merely delays work? (2) find the consume/refund pair and convince
yourself no handled path bills twice β then find the
unhandled window where a crash still can; (3) work out what happens
if the process dies between put_meal and deleting the
spool entry: the redo mints a fresh meal_id, so a
duplicate meal can appear β why does the design accept that residue
instead of engineering it away, and what would making the meal ID
stable per spool entry cost?