๐ŸฅฆNdiro
Chapter 3 of 13

The backend: Flask on one worker

The backend is one Python process. Flask's job is small and precise: turn an incoming HTTP request into a call to one of your functions, and turn whatever that function returns into an HTTP response. Everything else โ€” storage, auth, templates โ€” is libraries and your own code.

Routes: URL โ†’ function

app.py โ†—
@app.route('/privacy')
def privacy():
    return render_template('privacy.html', user=auth.current_user())

@app.route('/api/meals/<date_str>/<meal_id>', methods=['PUT'])
@auth.approved_required
def update_meal(date_str, meal_id):
    ...

The decorator registers the function in Flask's routing table. <date_str> segments are extracted from the path and passed as arguments โ€” but note they arrive as untrusted strings; the first thing the handler does is validate them (_valid_date rejects anything that isn't a canonical YYYY-MM-DD). methods=['PUT'] means the same path can route to different functions by verb.

Decorators stack, and each has its own enforcement point:

app.py โ†—
@app.route('/api/estimate-fiber', methods=['POST'])
@limiter.limit('6 per minute')     # enforced in a before-request hook โ€”
                                   # runs FIRST, before any auth check
@auth.approved_required            # wraps the view โ€” runs when it's invoked
def estimate_fiber():
    ...

The order of enforcement is a subtlety worth knowing: decorators wrap bottom-up, but Flask-Limiter doesn't do its counting in the wrapper โ€” it registers the limit and checks it in a hook that runs before the view is called at all. So an anonymous client hammering this endpoint sees 429 on the seventh call in a minute, not 401: the rate limiter shields the auth guard (and the database read inside it), not the other way around.

approved_required (in auth.py) is ~20 lines and worth reading now: it resolves the session to a fresh users-table row, stashes it in g.user (Flask's per-request scratch space), and short-circuits with a redirect or JSON 401/403 when the check fails. That one decorator is the app's entire authorization story.

Request in, response out

Inside a handler, the global-looking request object is actually request-scoped (Flask proxies it per thread). Everything the client sent hangs off it:

AccessorWhatExample in Ndiro
request.argsQuery stringrequest.args.get('month') in the meals API
request.formPOSTed form fieldsform.get('description') in _read_meal_form
request.filesUploaded filesrequest.files.get('photo')
request.get_json()JSON bodyThe text estimator's {'description': โ€ฆ}

Return values are flexible: a rendered template string, a jsonify(โ€ฆ) response, or a (body, status) tuple like jsonify({'error': 'Description is required'}), 400. App-wide error shapes live in @app.errorhandler functions โ€” see the 429 and 413 handlers at the top of app.py.

Startup: fail fast, then be lazy

Importing app.py runs the module top to bottom once: config.py loads the environment and raises if SECRET_KEY is unset (a misconfigured deploy should die at boot, not limp), db.ensure_tables() creates any missing DynamoDB tables, and the route table is built. Per-request work stays out of import time; the auto-log worker thread is even started lazily on first use because gunicorn forks after preload (the capstone chapter).

The single-worker design

This is the part of the backend written specifically for someone with your instincts, because it looks wrong until you read why:

Dockerfile โ†—
# SINGLE worker is load-bearing: the rate limiter is in-memory (memory://),
# so a second worker would hold a divergent limit state. Threads provide the
# concurrency; the 60s timeout leaves room for the AI calls' 20-25s reads.
CMD gunicorn --bind 0.0.0.0:$PORT --workers 1 --threads 8 --timeout 60 \
    --no-control-socket --preload app:app

gunicorn is the production process manager: it speaks WSGI (the Python standard interface between HTTP servers and apps โ€” python app.py's dev server is for development only). Workers are processes; threads share one process. Ndiro pins one worker, eight threads, and three pieces of state silently depend on that:

Python's GIL doesn't hurt here because the workload is I/O-bound โ€” threads spend their time waiting on DynamoDB, S3, and OpenAI, and the GIL is released during I/O.

๐Ÿ˜ From your world

This is a deliberately single-node system, the way a NameNode without HA is: correctness properties are pinned to "there is exactly one of me". Scaling to two containers would be a real project โ€” the limiter and photo cache move to Redis, the spool moves to a real queue โ€” and at ~100 users that machinery would be pure liability. The discipline is documenting the assumption where it's made, so nobody adds --workers 4 in a well-meaning "performance fix". That comment in the Dockerfile is doing the same job as an invariant comment in a consensus implementation.

โš ๏ธ Review trap

Any diff that adds cross-request, in-process state (a cache dict, a counter, a background thread) is implicitly betting on the single worker. That can be fine โ€” but the bet must be stated in a comment, and it must survive the question "what happens under --workers 2?"

๐Ÿ”ฌ Try it

Clone the repo, pip install -r requirements.txt, set a throwaway SECRET_KEY in .env, and run python app.py. Without AWS credentials most pages fail โ€” but /health, /status, /privacy, and this guide all work, because they touch no storage. Mapping "which routes need which backing services" is a good first exercise in reading app.py.

Then run docker compose up --build: the same image in local dev mode, where localdev.py swaps every backing service for a stand-in on your own disk and /dev signs you in as any seeded account. Read its install() โ€” the handful of functions it replaces are exactly the seams between this app and the cloud.

๐Ÿ