Skip to main content
Self-Hosting & Privacy

Docker Compose Healthchecks Explained

Compose's 'Up' status only means a process started. Real tuned interval/retries/start_period values across three homelab services, plus wiring depends_on: condition: service_healthy to close the gap.

Milan BuhaSeptember 29, 20267 min read
ShareXin
Docker Compose Healthchecks Explained

Docker Compose Healthchecks Explained

docker compose ps said every container was Up. The dashboard behind one of them had been returning 500s for forty minutes before anything noticed — a downstream backup job kept trying to hit it and kept failing, and nothing in Compose's own status line ever said so. Up was true the entire time. It was also useless.

TL;DR

  • Compose's default status (Up) only means the container's process started — it says nothing about whether the application inside is actually working.
  • A healthcheck: block gives Compose a real signal: test, interval, timeout, retries, and start_period, tuned per service.
  • start_period is the field people get wrong most — too short and a slow-booting app gets marked unhealthy while it's still starting; too long and a genuine crash-loop hides behind it.
  • Pair condition: service_healthy under a dependent service's depends_on so nothing starts against a container that's running but not ready — see our depends_on ordering guide for the restart side of this.
  • Compose does not restart an unhealthy container by itself — that's a job for restart: policy, not the healthcheck.

Why 'Up' Isn't 'Healthy'

By default, Docker only tracks one thing: is the container's main process still running. Per the Compose file healthcheck reference, without a healthcheck: block a container has no other state — it's either running or it's stopped, with nothing in between. A process can be alive and still be completely broken: stuck waiting on a dependency, wedged in a deadlock, or serving 500s on every request while its event loop technically keeps spinning.

That gap produces two failure modes:

  • Silently dead: the process is alive but the app inside has crashed into a bad state and never exits.
  • Slow-boot mistaken for dead: a service like a media indexer or a JVM app takes real time to become useful, and a naive check flags it unhealthy before it's had a fair chance.

A healthcheck: block turns "process running" into "application actually answering", and gives Compose (and anything that depends on this service) a signal it can act on.

The healthcheck Block, Field by Field

services:
  db:
    image: postgres:16
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U app"]
      interval: 5s
      timeout: 3s
      retries: 5
      start_period: 10s
  • test: the command Compose runs inside the container. CMD-SHELL runs it through the container's shell; a plain CMD array skips the shell and avoids surprises with quoting.
  • interval: how often the check runs once the container is past start_period.
  • timeout: how long a single check is allowed to run before it counts as a failure.
  • retries: consecutive failures needed before Compose marks the container unhealthy.
  • start_period: a grace window where failures don't count against retries — only a pass during this window can end it early.

Tuning It for Real Services (Not Toy Defaults)

Every healthcheck tutorial shows the Postgres example above and stops there. The values that actually matter are the ones that differ by service — a fast-failing API and a slow-booting media indexer should not share the same start_period. These are the values I run across three services on my own Proxmox homelab, picked after watching each one's real boot time in docker compose logs -f rather than copying a default:

Service Boot profile interval timeout retries start_period Why
Postgres (db) Fast, predictable 5s 3s 5 10s WAL recovery is rare and short on a clean shutdown; a short start_period catches a genuinely broken container fast.
Internal REST API Fast, but depends on db 10s 5s 3 15s Slightly longer start_period to absorb the API's own DB-connection retry loop on cold start.
Jellyfin (media indexer) Slow — library scan on first boot 30s 10s 3 120s A cold library scan on a large media volume can run past two minutes; a short start_period here just produces false unhealthy flaps.

KEY-STAT: 120s | start_period I run for a slow-booting media indexer before its failed checks count at all — 12x longer than the 10s I give Postgres, my homelab

Note

these numbers come from watching each container's own boot time, not a formula. The right way to set start_period is to run docker compose logs -f <service> once, note how long the app actually takes to become ready under normal conditions, then add roughly 30–50% margin — not to guess a round number and hope.

The start_period Judgment Call

start_period is the field most healthchecks get wrong, because it's the only one making a judgment call rather than measuring something concrete. Set it too short, and a legitimately slow-booting service — the media indexer above is the clearest case — gets marked unhealthy while it's still doing normal first-boot work, which then blocks anything with condition: service_healthy waiting on it. Set it too long, and a container that's genuinely stuck in a crash-loop sits inside the grace window for minutes before Compose says anything is wrong at all.

Warning

a start_period set to "whatever makes the flapping stop" instead of the service's real boot time is the single most common healthcheck mistake — it doesn't fix the readiness problem, it just hides it behind a longer grace window until the day the real boot time grows past that number too.

Wiring Healthchecks Into depends_on: condition: service_healthy

A healthcheck by itself only changes what docker compose ps reports. The payoff is connecting it to depends_on, so a dependent service doesn't start until the dependency is actually ready — not just started:

services:
  db:
    image: postgres:16
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U app"]
      interval: 5s
      timeout: 3s
      retries: 5
      start_period: 10s

  api:
    image: my-api:latest
    depends_on:
      db:
        condition: service_healthy

Without condition: service_healthy, per the Compose startup-order docs, depends_on only waits for db's container process to start — not for Postgres to finish recovery and accept connections. With it, api doesn't start at all until pg_isready reports success, which closes the exact gap that causes a first-boot migration or a first query to fail against a database that technically "started" seconds earlier.

What Happens When a Container Goes Unhealthy

Compose marking a container unhealthy does not restart it automatically — healthchecks and restart policies are two separate mechanisms. Diagnose an unhealthy container with:

docker compose ps
docker inspect --format='{{json .State.Health}}' <container> | python3 -m json.tool

The second command prints the actual failing output of the last few check attempts, which is almost always faster than guessing from logs alone. To make Compose actually recover from the unhealthy state, pair the healthcheck with a restart: policy — the specifics of which policy fits which failure mode are covered in our Docker Compose restart-policies guide; a healthcheck without a restart policy behind it just produces an accurate but permanent "broken" label.

Common Healthcheck Mistakes

  • Using curl or wget in a minimal image that doesn't have them — the check fails immediately with "command not found", which looks identical to a real failure in docker compose ps. Either install the tool deliberately or use a language-native check (Postgres's own pg_isready, a small script, or the app's existing /health endpoint via a tool that's already in the image).
  • Checking that a port is open instead of that the app is ready — a port can accept TCP connections before the application behind it can serve a real request, especially for services with a warm-up phase.
  • Retries set too low for a service under real load — a container that's briefly slow under load (not broken) gets flagged unhealthy and restarted, which is worse than the transient slowness it was reacting to. If you're still deciding which network mode a busy service should run under, our network_mode: host explainer covers a related source of connectivity flakiness.

FAQ

How do I check if a Docker container is healthy?

Run docker compose ps to see the status column, or docker inspect --format='{{json .State.Health}}' <container> for the full history of recent check attempts and their output.

What is start_period in a Docker healthcheck?

A grace window after the container starts during which failed checks don't count against retries. Only a passing check during this window can end it early. It exists so a slow-booting service isn't marked unhealthy before it's had a fair chance to finish starting.

Why does my container show as running but the app isn't working?

Because without a healthcheck: block, Docker only tracks whether the container's process is alive — not whether the application inside is actually responding correctly. Add a healthcheck that tests real readiness, not just that the process exists.

Does Docker Compose restart unhealthy containers automatically?

No. A healthcheck only changes the reported status. Automatic recovery requires a separate restart: policy on the service.

How does depends_on condition service_healthy work in Docker Compose?

It makes a dependent service wait until the dependency's own healthcheck reports healthy before starting — closing the gap left by the default depends_on behavior, which only waits for the dependency's container process to start.

The Bottom Line

Up was never a readiness signal — it's a process-alive signal that happens to look like one. A tuned healthcheck: block, sized to each service's real boot time rather than a copied default, turns that gap into a real status Compose (and anything depending on the service) can act on. If you're assembling the rest of a multi-service stack around this, our Docker Compose self-hosting guide covers the pieces this one assumes, and our guide to debugging a broken Compose stack is the next step once a healthcheck tells you something actually is wrong.

Related stories

More from Self-Hosting & Privacy

Stay in the loop

Get the latest articles delivered to your inbox. No spam, unsubscribe anytime.

Read next

network_mode: host in Docker Compose Explained

Home Assistant or Pi-hole goes into a docker-compose.yml, the stack comes up healthy, and device discovery still doesn't work -- no new devices show up. The forum fix is network_mode: host. What it actually does, when it's worth the isolation you give up, and when bridge is still right.

Continue Reading