Cloud 08 — Docker Compose for the flagship: api, pgvector, Langfuse

In Java there is no standard way to tell an orchestrator "don't start me until my database can actually accept a connection" — Spring Boot papers over it with its own retry loop inside HikariCP, or a shell script polls a port before java -jar runs. Docker Compose has a real answer: depends_on: postgres: condition: service_healthy. The trap is that depends_on: postgres on its own — no condition — looks like the same thing and is not: it means "start my container after the postgres container process exists," which can be a full second or more before Postgres is accepting connections. That gap is a real, reproducible race, and this guide measures it before showing the fix.

~90 min$0 Docker Composecloud-08

What this creates

Three local containers for the flagship: postgres (the pgvector/pgvector image — plain Postgres plus the vector extension), langfuse (self-hosted tracing, pointed at the same Postgres), and api (the flagship's own FastAPI service, built the same way as cloud-01's multi-stage Dockerfile). Nothing here touches AWS or costs anything — it is the local loop you run on every commit, before cloud-09 puts the same shape on Fargate. Pick a startup mode below; the log underneath is a real, measured run on this machine (Docker 28, Compose v2, an Apple-silicon Mac) — not a guess at what would happen.

Preconditions

Do it

1. compose.yaml — complete; three services, nothing else:

# compose.yaml
services:
  postgres:
    image: pgvector/pgvector:pg16
    restart: unless-stopped
    environment:
      POSTGRES_USER: f1
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set in .env}
      POSTGRES_DB: f1
    volumes:
      - pgdata:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U f1 -d f1"]
      interval: 2s
      timeout: 3s
      retries: 10
      start_period: 10s

  langfuse:
    image: langfuse/langfuse:2         # v2 on purpose: the last major that runs on Postgres alone (see note below)
    restart: unless-stopped
    depends_on:
      postgres:
        condition: service_healthy       # waits for pg_isready, not just "container exists"
    environment:
      DATABASE_URL: postgresql://f1:${POSTGRES_PASSWORD:?set in .env}@postgres:5432/f1
      NEXTAUTH_SECRET: ${LANGFUSE_NEXTAUTH_SECRET:?set in .env}
      SALT: ${LANGFUSE_SALT:?set in .env}
      NEXTAUTH_URL: http://localhost:3000
    ports:
      - "3000:3000"
    healthcheck:                          # /api/public/health is Langfuse's documented liveness path (v2 and later)
      test: ["CMD", "wget", "-q", "--spider", "http://localhost:3000/api/public/health"]
      interval: 5s
      timeout: 3s
      retries: 10
      start_period: 20s

  api:
    build: ./api
    restart: unless-stopped
    depends_on:
      postgres:
        condition: service_healthy
    # no depends_on: langfuse — the tracing SDK batches and retries in a background
    # thread; a trace call before langfuse answers is queued, not a crash. Gate startup only
    # on what a failed connection actually breaks: the database.
    environment:
      DATABASE_URL: postgresql://f1:${POSTGRES_PASSWORD:?set in .env}@postgres:5432/f1
      LANGFUSE_HOST: http://langfuse:3000
      LANGFUSE_PUBLIC_KEY: ${LANGFUSE_PUBLIC_KEY:-}
      LANGFUSE_SECRET_KEY: ${LANGFUSE_SECRET_KEY:-}
    ports:
      - "8000:8000"
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request as u; u.urlopen('http://localhost:8000/readyz', timeout=2)"]
      interval: 3s
      timeout: 3s
      retries: 5
      start_period: 10s

volumes:
  pgdata:
Why Langfuse v2 and not the current major. From v3 on, Langfuse's own docker-compose.yml is six services, not one: langfuse-web, langfuse-worker, Postgres, ClickHouse, Redis and MinIO, plus an ENCRYPTION_KEY and a dozen CLICKHOUSE_*/REDIS_*/LANGFUSE_S3_* variables. That is the right shape for a team's tracing server and the wrong shape for a 90-minute local loop, so this guide pins langfuse/langfuse:2, which still runs on the one Postgres you already have. The @observe SDK calls in the flagship are the same against either version. When you outgrow it, lift Langfuse's compose file into yours as-is rather than hand-porting it — the five extra containers are exactly what a bare depends_on gets wrong.

2. .env (git-ignored; .env.example in the repo instead):

# .env.example — copy to .env, then fill in real values; .env is git-ignored
POSTGRES_PASSWORD=
LANGFUSE_NEXTAUTH_SECRET=
LANGFUSE_SALT=
LANGFUSE_PUBLIC_KEY=
LANGFUSE_SECRET_KEY=

3. Makefile — three targets, nothing hidden inside them:

# Makefile
up:
	docker compose up -d --wait   # exits non-zero if any service never reaches healthy

test:
	docker compose ps --format '{{.Service}}: {{.Health}}'
	curl -sf http://localhost:8000/readyz
	curl -sf http://localhost:3000/api/public/health

down:
	docker compose down -v        # -v drops pgdata too — a clean slate, not a pause

Verify

Teardown

make down
# same as: docker compose down -v
# confirm nothing is left:
docker compose ps -a   # empty
docker volume ls | grep pgdata   # no output — the volume, and every row it held, is gone

After teardown: no container, no volume, no network — pgdata is deliberately dropped with -v so the next make up starts from an empty database, the same as a fresh clone would. Nothing here ever touched AWS, so there is no bill to check.

This is self-attestation — the site cannot see your Docker daemon, so checking the box and pressing the button is you telling The Path you actually ran it.

Takeaways: depends_on without a condition only orders container starts, not application readiness — the gap between the two is a real race, not a theoretical one, and this page measured it: 0 restarts gated, several ungated, and a worse number ungated the moment Postgres takes even a little longer to come up. A healthcheck on a container that has no dependents is still worth writing, because docker compose ps and --wait are the only things that can tell you "started" and "actually ready" apart. And depends_on is a startup-order tool, not a runtime one — once everything is up, an app's own /readyz is what should notice a downstream dependency going away later.