πŸ₯¦Ndiro
Chapter 11 of 13

Capstone: the auto-log pipeline

The "auto-add from photos" feature is the capstone because it exercises every layer of this book at once β€” and because it's a queueing system, so you get to grade it against the real ones you've run. The user selects a batch of food photos; each becomes its own meal, at the time the photo was taken, with an AI description and estimate β€” and the user can close the tab the moment the upload finishes.

Why async at all

An AI vision call takes up to ~25 seconds. The JavaScript chapter's event loop says the browser can't block, and the backend chapter's gunicorn config says a request thread shouldn't sit occupied for a 20-photo batch either (8 threads total, 60s timeout). Classic answer: make the upload cheap and synchronous, move the slow work to a background consumer, let the client poll for progress. That's precisely the shape of autolog.py.

The pipeline, end to end

browser                          POST /api/auto-log                worker thread (one)
   β”‚ user picks N photos                β”‚                                β”‚
   β”‚ per photo: read EXIF time,         β”‚                                β”‚
   β”‚ upload ───────────────────────────▢│ normalize to JPEG              β”‚
   β”‚                                    β”‚ dedup check (below)            β”‚
   β”‚                                    β”‚ write photo + sidecar JSON     β”‚
   β”‚                                    β”‚ to the LOCAL spool dir ──────▢ β”‚ loop:
   β”‚ ◀── 202 {queued: …} ───────────────│                                β”‚  re-read user row (still approved?)
   β”‚                                    β”‚                                β”‚  consume 1 AI use (cap)
   β”‚ polls /api/auto-log/pending        β”‚                                β”‚  OpenAI vision estimate
   β”‚ to render per-photo status         β”‚                                β”‚  put_photo + put_meal
   β”‚                                    β”‚                                β”‚  delete spool entry
   β”‚                                    β”‚                                β”‚  (on failure: retry; then degrade
   β”‚                                    β”‚                                β”‚   or dead-letter as *.json.dead)

The spool is just files on local disk: the JPEG plus a tiny sidecar JSON holding user_id/date/time/attempts β€” deliberately no meal content, with one earned exception: after a failed commit the paid-for AI estimate (description included) is cached onto the sidecar so the retry never re-bills. One daemon thread consumes it. It's started lazily (ensure_worker) because gunicorn's --preload imports the app before forking, and a thread started at import time would be lost in the fork β€” a genuinely sneaky deployment interaction. The kick happens from the auto-log routes and from /log, which is also the crash-recovery story: after a restart, the first page view revives the worker and it finds whatever the spool still holds.

Grade it like a queue system

PropertyHow Ndiro gets it
DurabilitySpooled to disk before the 202 Accepted (the status that says "taken, not yet done") β€” but local disk, ephemeral on most hosts. A stated PoC trade: a redeploy may drop queued work, and the design accepts it rather than paying for SQS.
Idempotent producerRe-uploading the same photos is safe: the EXIF date+minute is the dedup fingerprint, checked against both existing PHOTO meals (a range query over the batch's whole date window β€” clamped camera clocks drift days) and the queued spool.
At-most-once billing per handled retryThe AI cap is consumed per photo before the call, refunded on upstream failure, and the estimate is cached on the sidecar so handled commit retries never re-bill. (A process death in the window after the AI reply but before that cache write can still bill twice on restart β€” at-least-once delivery leaves exactly this residue.)
Poison messagesMAX_ATTEMPTS, then the sidecar is renamed *.json.dead β€” a dead-letter queue, kept for the operator.
Graceful degradationCap exhausted or retries spent? The meal is committed anyway with a placeholder description β€” the photo and its timestamp are never lost; only the AI garnish is.
Revocation raceThe worker re-reads the user row before each commit; a rejected or deleted account's queued photos are discarded, never logged. Account deletion also purges the user's spool (drop_user) before the S3 wipe.
🐘 From your world

A single-broker, single-partition, at-least-once queue with a DLQ, producer-side dedup by natural key, and consumer-side quota enforcement β€” the entire Kafka-consumer checklist, in ~320 lines of standard library. The interesting engineering isn't any one mechanism; it's that each one exists because a specific failure was imagined first. Read autolog.py top to bottom and you'll find the failure story attached to every design choice, and the trade-offs it declines to pay for stated out loud.

The frontend half

Back in log.html, the batch UI shows the other end: reading EXIF timestamps client-side, uploading sequentially with per-photo status rows, then polling /api/auto-log/pending so a user who does keep the tab open watches photos flip from "queued" to done. Everything you learned lands here at once: file inputs (HTML), status colors from theme tokens (CSS), createElement rendering (JavaScript), fetch with the standard error ladder (the API chapter).

And the tests

How do you test a background thread deterministically? You don't β€” you make the thread optional. Tests set autolog.WORKER_ENABLED = False and call process_once() directly, single-stepping the consumer: every retry, refund, dead-letter and race in tests/test_m11_autolog.py is asserted without a sleep or a flake. Separating "the loop" from "one step of the loop" for testability is a pattern worth stealing in any language.

πŸ”¬ Try it

Read autolog.py with three questions: (1) where exactly is the boundary between "the upload request" and "the worker" β€” which failure before it loses work, and which after it merely delays work? (2) find the consume/refund pair and convince yourself no handled path bills twice β€” then find the unhandled window where a crash still can; (3) work out what happens if the process dies between put_meal and deleting the spool entry: the redo mints a fresh meal_id, so a duplicate meal can appear β€” why does the design accept that residue instead of engineering it away, and what would making the meal ID stable per spool entry cost?

πŸ