Skip to content

Docker Compose Waits for a Ready Database, Not a Running One

Adityo Guni Waluyo

depends_on only waits for a running container. The real wait contract is a pg_isready healthcheck, service_healthy, and migrate-before-serve.

TL;DR

Compose's depends_on only guarantees start order, not database readiness, so the app crashed while Postgres was still initializing. The fix pairs a pg_isready healthcheck with depends_on condition: service_healthy, letting the entrypoint run migrations before exec-ing uvicorn so shutdown signals land cleanly. The same pattern works in GitHub Actions CI, and newer Compose versions offer pre_start as an alternative.

Bootstrap night for DemandScope: the first docker compose up, and the screen filled with uvicorn logs failing to reach the database until the container gave up. What kept me staring: the Postgres service was clearly running. I opened compose.yaml, and there was depends_on pointing at the database service. Why didn’t it wait?

The answer lives in Compose’s definition of “waiting”. At startup, Compose does not wait for containers to become ready, only running [1]. Running means the main process has been started, not that Postgres finished initializing and port 5432 is accepting connections. Those few seconds are what made my app slam into a door that had not opened. Last night’s guess: depends_on surely takes care of this. The default turns out to be nothing more than start order.

Layer 1: a healthcheck that means something

The wait contract starts with a healthcheck block at the service level [4]. For Postgres, the documented readiness probe is pg_isready, and the official Compose example calls it through CMD-SHELL [1]. In DemandScope I wrote it like this:

services:
  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: demand_user
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set in .env}
      POSTGRES_DB: demand_db
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 5s
      timeout: 5s
      retries: 20

pg_isready fires every 5 seconds, and retries 20 is the patience limit: at most about a hundred seconds before the db is declared unhealthy. These numbers are not decoration. They decide how long the chain below is allowed to wait, and they sit in one place so the whole team reads the same contract.

Layers 2 and 3: service_healthy, then migrations before serving

In the app service, depends_on gets a condition, not just a list:

  app:
    build: .
    depends_on:
      db:
        condition: service_healthy
    entrypoint:
      - sh
      - -c
      - alembic upgrade head && exec uvicorn app.main:app --host "$$APP_BIND" --port 8000
    healthcheck:
      test: ["CMD", "python", "-c", "<GET /health from inside the container, exit 0 on 200>"]
      interval: 10s
      timeout: 5s
      retries: 12
      start_period: 15s

condition: service_healthy makes Compose hold the app’s startup until db passes its healthcheck [1]. Only then does the entrypoint run: alembic upgrade head first, the server afterwards.

The exec at the end of the chain is not decoration. A container’s main process is responsible for the processes it starts [5], and Linux ignores signals delivered to PID 1 unless it carries its own handler [7]. Without exec, the shell occupies PID 1; the SIGTERM from docker compose down stops at the shell, uvicorn never hears about it, and the container dies hard once the timeout expires. With exec, uvicorn replaces the shell as PID 1 and shutdown lands cleanly.

One honest footnote: the Dockerfile also carries a HEALTHCHECK instruction for container health [2]. I still put the healthcheck at the compose service level so one file governs the whole chain, from database to application.

The same pattern in CI

GitHub Actions can declare a Postgres service container for a test job [6], and the workflow file must live under .github/workflows per GitHub’s rules [3].

jobs:
  test:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_USER: demand_user
          POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set in .env}
          POSTGRES_DB: demand_db
        options: >-
          --health-cmd "pg_isready -U demand_user -d demand_db"
          --health-interval 5s
          --health-timeout 5s
          --health-retries 20
    env:
      DATABASE_URL: postgresql+psycopg://demand_user:<password-from-secret>@<runner-address>:5432/demand_db
    steps:
      - run: alembic upgrade head
      - run: pytest -q

The result: local and CI wait for the database under the exact same contract. A CI runner is always fresh, so without health options like these the test job can lose the race against Postgres initialization; with the pattern above there is no random sleep in scripts and no manual restart ritual.

One recent Compose change deserves its own line: since 5.3.0 there is pre_start, an init-container style step that takes over setup work like migrations with tidier semantics, for example a step that already succeeded is not repeated when the container restarts [8]. If your Compose version qualifies, that is a cleaner shape than migrations inside the entrypoint. I still pick the entrypoint for cross-version portability, but pre_start is recorded as the upgrade path, not as doctrine.

“Wait before you connect” finally became configuration you can test instead of a restart reflex. If DemandScope moves to pre_start later, the healthcheck chain underneath stays the same.

Sources

  1. [1] Control startup and shutdown order in Compose
  2. [2] Dockerfile reference
  3. [3] Workflow syntax for GitHub Actions
  4. [4] Compose services reference
  5. [5] Run multiple processes in a container
  6. [6] Creating PostgreSQL service containers (GitHub Actions)
  7. [7] docker container run reference
  8. [8] Use init containers in Compose

Related articles