Skip to content

Red Pipeline, Green Laptop: Migrations Before pytest

Adityo Guni Waluyo

CI called pytest on an empty database while production applied migrations first; one workflow line aligned both contracts.

TL;DR

A test suite that passed locally kept failing in CI with misleading database errors. The root cause was that production ran Alembic migrations before boot, while CI pointed pytest straight at an empty Postgres. Adding alembic upgrade head before tests in the workflow made starting state deterministic, so failures now point at actual code.

Red pipeline, green laptop

I stared at the red CI pipeline for DemandScope. The exact same test suite passed flawlessly on my laptop, but in continuous integration it was a wall of failures. The suite covers authentication end-to-end, registration, login, and profile, all executed against a real Postgres service in the pipeline. My initial guess was the usual one: the code is fine, the CI runner is just being weird.

CI failures like this feel personal because the laptop becomes the alibi. If it passes here, the code must be right, and everything downstream is somebody else’s config problem. The guess missed, and the error trail pointed the wrong way while it did. Before the fix, failures surfaced at the wrong layer: refused connections, unknown relations, missing tables. Nothing about those messages says a migration is missing, so debugging kept circling the code when the environment was the problem.

One more thing made it worse: the pipeline still reported test names as the failing unit. A suite that runs against a half-born database produces errors that look like logic bugs. The truth was structural. The production container applies database migrations before the application starts; CI called pytest directly on an empty database service. Two environments, two different contracts.

Schema first, tests second

Migrations are code, and the schema they produce is part of the API surface. Tables, indexes, and constraints decide which requests are even possible. The schema is part of the API surface. Testing without migrations means testing code against an imaginary schema, and Alembic exists exactly to keep that from drifting: the command alembic upgrade head runs every migration up to the newest revision, and Alembic tracks the database position along the way[10]. The two-line local sequence makes the failure reproducible:

  - run: pip install ".[dev]"
  - run: ruff check .
  # same contract as the container CMD: schema first, then tests
  - run: alembic upgrade head
  - run: pytest -v

pytest exit codes are the contract CI reads: 0 means all tests passed, 1 means some failed[11]. With the schema in place, a failing code path surfaces as a failed assertion at the code layer instead of a database connection error. Once both environments follow one order, the workflow only needs the same two steps before the test run:

# without migrations, pytest fires at an empty schema
alembic upgrade head
pytest -v

The distinction matters for maintenance. When a migration ships in the same pull request as the code that needs it, the workflow line guarantees the order everywhere: CI applies the schema before any assertion runs, and production does the same before the app boots. One contract, written once, enforced twice.

It is worth stating what the contract buys in practice. Before it existed, every red run began with an investigation of the environment: is the service up, did the image cache betray us, which table is missing. After it, the environment is boring by construction, and the search space collapses to the diff. A test suite only earns that trust when its starting state is deterministic, and for a database-backed suite deterministic means migrated.

The comment above the migration line is not decoration; it states the contract so nobody removes the step while tidying the workflow. Reliability went up not because tests were added but because the initial condition stopped lying. The migration line is cheap; the hours lost reading misleading error trails are what is expensive.

The laptop-versus-CI divide never comes from the tests themselves. It comes from everything around them: service versions, cached layers, and starting state. Starting state is the only one of those three that a two-line fix can make identical everywhere, which is why the schema contract is the first thing to write down when a database-backed suite joins the pipeline, and the last thing to debug afterwards.

Sources:

[10] Alembic Tutorial

[11] pytest: exit codes

Related articles