Skip to content

Hybrid Upgrade Playbook: npm In-Container, Image Fallback

Adityo Guni Waluyo

A two-phase upgrade playbook: npm inside the running container plus an automatic custom-image fallback, verified against the version the app actually serves.

TL;DR

After upgrading 9router, the endpoint still served the old version because the container ran code baked into the image, not the freshly installed npm package. The rewritten playbook uses a two-phase approach: cheap in-container installs first, then building an image from source, with served-version verification, health checks, and automatic rollback. Running serial: 1 keeps failures contained to one host.

That morning I ran the upgrade playbook for 9router on one of the fleet hosts. The npm install of 0.5.99 [8] finished without a single error, the container restarted cleanly. Then I opened the /api/version endpoint and the served version was still 0.5.95. The package inside the container had moved; the application had not.

My first guess was wrong twice. I suspected the npm cache, cleared it, ran the install again. Same result. After dissecting the package I found the cause: the 9router npm package does not ship .next/custom-server.js, and the running CMD executes the bundled code from the image. npm i -g only swaps the global package under node_modules, while the live process keeps executing the old code baked into the image. The package version rises; the served version does not.

So I rewrote the upgrade playbook as a two-phase hybrid: Phase A is fast, Phase B is heavy. The most important change was not the install method but the definition of "success" the playbook uses.

Verify the Version That Is Actually Served

Phase A still runs first, because for most releases an in-container npm install is far cheaper than an image build: docker exec runs npm i -g, the container restarts, done. The new part is the verification. The playbook no longer trusts the package manager's report. It calls /api/version repeatedly and only declares success when the served version matches the target. When it does not, the playbook does not fail; it falls through to Phase B.

The right definition of success turns out to be a philosophical matter. The package version is a claim; the served version is the fact users experience. One small endpoint turns a playbook from "probably updated" into "proven updated".

Phase B: Build from Source Instead of Waiting for Docker Hub

Phase B attacks the old root problem: images on Docker Hub often lag far behind npm releases. Instead of waiting, the playbook builds its own image from the source master on the control node. Three gates protect the build. A source-version guard compares the version in the source package.json with the target; on a mismatch the build aborts rather than producing an image of the wrong version. Once built, docker image save packs the image with all of its parent layers and tags into a tar archive [4], gzip compresses it, and the Ansible copy module ships it to the host [6]. On the host, docker image load reads the compressed archive and restores the image along with its tags [5]. If the image or archive for that version already exists from a previous run, every one of these steps is skipped. The playbook can be re-run at will without rebuilding.

The container swap follows a rename pattern: the old container is stopped, renamed into a spare, and the new one is started with identical flags, the same data bind mount, the same port, the same log rotation, and the unless-stopped restart policy [3]. That policy fits this case: the container comes back after a daemon restart but stays down when I stop it for maintenance.

The Health Gate with Automatic Rollback

The new container must pass a health gate before anything is declared done. Docker only considers a container "successfully started" after it has been up for ten seconds, a guard against restart loops [10]. My playbook is stricter than that: HTTP probes until a redirect or 200 shows up plus a served version matching the target, eight attempts, three seconds apart. On failure the spare container is renamed back and started. Rollback is part of the playbook's normal flow, not an emergency plan.

Two more guards close classic gaps. Never-downgrade: when the served version is somehow higher than the target, the playbook stops for that host through meta: end_host, which ends the play for that host without marking it failed [7]. The host simply skips to the sidelines; no silent version drop ever happens. And because registered variables are always set for every host, including the ones that fail or get skipped [2], the per-host report stays inspectable through the debug module without halting the playbook [9].

serial: 1, a Blast Radius of One Host

The whole run goes through serial: 1, the Ansible keyword for rolling updates that finishes the play for one batch of hosts before touching the next [1]. With one host per batch, the worst failure touches a single machine and the rest of the fleet stays safe.

The end-to-end test on one host proved the flow: from 0.5.95 to 0.5.99 with no other container touched. The full fleet rollout then ran on the same playbook, and a re-run with --limit for one slow host showed its idempotency. Two manual recipes I once wrote in two separate articles have become a single playbook that makes its own decisions. That is the difference between a playbook that admires the package version and one that checks the version actually being served.

Sumber: [1] Ansible strategies (serial) · [2] Ansible conditionals (register) · [3] docker run reference · [4] docker image save · [5] docker image load · [6] Ansible copy module · [7] Ansible meta (end_host) · [8] npm registry, 9router · [9] Ansible debug module · [10] Docker policy guide

Related articles