The backend is one Python process. Flask's job is small and precise: turn an incoming HTTP request into a call to one of your functions, and turn whatever that function returns into an HTTP response. Everything else โ storage, auth, templates โ is libraries and your own code.
@app.route('/privacy')
def privacy():
return render_template('privacy.html', user=auth.current_user())
@app.route('/api/meals/<date_str>/<meal_id>', methods=['PUT'])
@auth.approved_required
def update_meal(date_str, meal_id):
...
The decorator registers the function in Flask's routing table.
<date_str> segments are extracted from the path and
passed as arguments โ but note they arrive as untrusted
strings; the first thing the handler does is validate them
(_valid_date rejects anything that isn't a canonical
YYYY-MM-DD). methods=['PUT'] means the same
path can route to different functions by verb.
Decorators stack, and each has its own enforcement point:
app.py โ@app.route('/api/estimate-fiber', methods=['POST'])
@limiter.limit('6 per minute') # enforced in a before-request hook โ
# runs FIRST, before any auth check
@auth.approved_required # wraps the view โ runs when it's invoked
def estimate_fiber():
...
The order of enforcement is a subtlety worth knowing:
decorators wrap bottom-up, but Flask-Limiter doesn't do its counting in
the wrapper โ it registers the limit and checks it in a hook that runs
before the view is called at all. So an anonymous client hammering this
endpoint sees 429 on the seventh call in a minute, not
401: the rate limiter shields the auth guard (and the
database read inside it), not the other way around.
approved_required (in auth.py) is ~20 lines
and worth reading now: it resolves the session to a fresh users-table
row, stashes it in g.user (Flask's per-request scratch
space), and short-circuits with a redirect or JSON 401/403 when the
check fails. That one decorator is the app's entire authorization
story.
Inside a handler, the global-looking request object is
actually request-scoped (Flask proxies it per thread). Everything the
client sent hangs off it:
| Accessor | What | Example in Ndiro |
|---|---|---|
request.args | Query string | request.args.get('month') in the meals API |
request.form | POSTed form fields | form.get('description') in _read_meal_form |
request.files | Uploaded files | request.files.get('photo') |
request.get_json() | JSON body | The text estimator's {'description': โฆ} |
Return values are flexible: a rendered template string, a
jsonify(โฆ) response, or a (body, status) tuple
like jsonify({'error': 'Description is required'}), 400.
App-wide error shapes live in @app.errorhandler functions โ
see the 429 and 413 handlers at the top of app.py.
Importing app.py runs the module top to bottom once:
config.py loads the environment and raises if
SECRET_KEY is unset (a misconfigured deploy should die at
boot, not limp), db.ensure_tables() creates any missing
DynamoDB tables, and the route table is built. Per-request work stays
out of import time; the auto-log worker thread is even started lazily on
first use because gunicorn forks after preload (the capstone chapter).
This is the part of the backend written specifically for someone with your instincts, because it looks wrong until you read why:
Dockerfile โ# SINGLE worker is load-bearing: the rate limiter is in-memory (memory://),
# so a second worker would hold a divergent limit state. Threads provide the
# concurrency; the 60s timeout leaves room for the AI calls' 20-25s reads.
CMD gunicorn --bind 0.0.0.0:$PORT --workers 1 --threads 8 --timeout 60 \
--no-control-socket --preload app:app
gunicorn is the production process manager: it speaks WSGI (the
Python standard interface between HTTP servers and apps โ python
app.py's dev server is for development only). Workers are
processes; threads share one process. Ndiro pins one worker,
eight threads, and three pieces of state silently depend on
that:
storage_uri='memory://'),db.py,Python's GIL doesn't hurt here because the workload is I/O-bound โ threads spend their time waiting on DynamoDB, S3, and OpenAI, and the GIL is released during I/O.
This is a deliberately single-node system, the way a
NameNode without HA is: correctness properties are pinned to "there
is exactly one of me". Scaling to two containers would be a real
project โ the limiter and photo cache move to Redis, the spool moves
to a real queue โ and at ~100 users that machinery would be pure
liability. The discipline is documenting the assumption where it's
made, so nobody adds --workers 4 in a well-meaning
"performance fix". That comment in the Dockerfile is doing the same
job as an invariant comment in a consensus implementation.
Any diff that adds cross-request, in-process state (a cache dict,
a counter, a background thread) is implicitly betting on the single
worker. That can be fine โ but the bet must be stated in a comment,
and it must survive the question "what happens under
--workers 2?"
Clone the repo, pip install -r requirements.txt, set
a throwaway SECRET_KEY in .env, and run
python app.py. Without AWS credentials most pages
fail โ but /health, /status,
/privacy, and this guide all work, because they touch
no storage. Mapping "which routes need which backing services" is a
good first exercise in reading app.py.
Then run docker compose up --build: the same image in
local dev mode, where localdev.py swaps every
backing service for a stand-in on your own disk and
/dev signs you in as any seeded account. Read its
install() โ the handful of functions it replaces are
exactly the seams between this app and the cloud.