# py-12-async-service — `f1api`: the flagship-1 service, fully async, load-tested with locust

py-10 built the shape (`/ask`, health vs readiness, a request id that sticks); this exercise adds
the piece that shape needs once it talks to Postgres for real — F1's pgvector store on cloud-09's
RDS instance — without a blocking call anywhere in the request path. You build a connection-pool
layer around an `asyncpg`-shaped `Pool`, a `/readyz` that reflects the pool's health instead of a
model's, and a shutdown that waits for in-flight requests to actually finish before it closes the
pool out from under them. The brief is the same text as the page
(`illustrated/7-mlops/exercise-async-service.html`).

## What to implement

`src/f1api/app.py` — `create_app(pool)`:

- `RequestIdMiddleware` — carried over unchanged from py-10 (same ASGI-middleware shape,
  same `ContextVar`); implement it exactly as before if you don't have that file handy.
- `POST /ask` — the answer is `f"hi there! you said: {body.prompt}"` (there is no real model
  here, the point is the pool); `async with pool.acquire() as conn:` and `await conn.execute(...)`
  an insert logging `(current_request_id(), body.prompt, answer)`.
- `GET /healthz` — always `{"status": "ok"}`, always 200, never touches `pool`.
- `GET /readyz` — `{"ready": pool.healthy}`, 503 when the pool is unhealthy.
- a `drain_and_close()` coroutine — wait for `pool.in_use == 0`, *then* `await pool.close()` —
  reachable both as `app.state.shutdown` (what the tests call directly) and via
  `@app.on_event("shutdown")` (what a real `uvicorn` shutdown runs).

`src/f1api/reqctx.py`, `schemas.py` and `pool.py` (the `FakePool` the tests run against) are
**provided** — do not edit them. `src/f1api/main.py` (the uvicorn entrypoint, real `asyncpg`) is
provided too and is never imported by the tests.

## Run it

```
cd exercises/py-12-async-service && uv sync && uv run pytest -q

# once the tests pass: talk to it for real (needs a reachable Postgres — cloud-09's RDS, or any local one)
DATABASE_URL=postgresql://... uv run uvicorn f1api.main:app --port 8000 &
curl -s -X POST localhost:8000/ask -H 'content-type: application/json' -d '{"prompt": "hi"}'
curl -si localhost:8000/readyz

# load test — watch p95 climb once concurrent users pass F1API_POOL_MAX
uv run locust -f locustfile.py --host http://127.0.0.1:8000

# the same bar the reference solution clears
uv run ruff check . && uv run mypy src
```

Done when `uv run pytest -q` prints **6 passed**. The untouched starter fails all six.

## The checks

- `test_no_sync_db_calls_or_blocking_sleep_in_request_path` — a static scan of `src/`: no
  `psycopg2` import, no blocking `time.sleep`.
- `test_pool_size_is_respected_under_concurrent_requests` — 9 concurrent `/ask` calls against a
  pool of 3 peak at exactly 3 connections checked out at once — proof `/ask` actually goes
  through `pool.acquire()`, and that nothing bypasses the cap.
- `test_100_concurrent_requests_complete_under_the_fake_pool` — 100 concurrent `/ask` calls
  against a pool of 8 all return 200.
- `test_graceful_shutdown_drains_inflight_requests` — 4 requests already holding a connection
  when `app.state.shutdown()` runs all still complete with 200, and the pool only closes once
  they have.
- `test_request_ids_preserved_under_concurrent_requests` — 4 concurrent requests, 4 distinct
  `X-Request-Id` headers; each response's header and body `request_id` still match its own
  request after an `await pool.acquire()` in the middle of the handler.
- `test_readyz_flips_to_503_when_pool_is_lost` — 200 while the pool is healthy, 503 with
  `{"ready": false}` after `pool.simulate_loss()`.

## Files

- [README.md](README.md) — this brief, offline
- [pyproject.toml](pyproject.toml) — deps, ruff and mypy config
- [locustfile.py](locustfile.py) — the load test
- [src/f1api/app.py](src/f1api/app.py) — the starter you edit
- [src/f1api/pool.py](src/f1api/pool.py) — `FakePool`, the asyncpg stand-in
- [tests/test_service.py](tests/test_service.py) — the checks

## If you get stuck

- **`test_pool_size_is_respected_under_concurrent_requests` reads a peak of 0** — `/ask` never
  called `pool.acquire()` at all; check the insert is inside the `async with` block, not after it.
- **`test_graceful_shutdown_drains_inflight_requests` raises `PoolClosedError`** — the shutdown
  handler called `pool.close()` without first waiting for `pool.in_use` to hit 0; the fake pool
  raises immediately instead of hanging the way the real one would.
- **`test_request_ids_preserved_under_concurrent_requests` is flaky or red** — a plausible bug is
  reading the request id from a module-level variable instead of `current_request_id()`; under
  real concurrency (this test's whole point) a later request overwrites it before an earlier one
  reads it back.
- **Reading a red row** — `uv run pytest -q -x --tb=short` stops at the first failure and shows
  the assertion that tripped; the message names the observed value, not the fix.
