DevOps & VPS2026-09-248 min read903 views

Running Docker Swarm on Hetzner: Why We Choose It Over Kubernetes for Small & Medium Stacks

Why we stopped paying the Kubernetes complexity tax for medium workloads, and how to run a resilient, zero-downtime cluster on low-cost European VPS.

Share:𝕏 PostLinkedIn
Running Docker Swarm on Hetzner: Why We Choose It Over Kubernetes for Small & Medium Stacks
Running Docker Swarm on Hetzner: Why We Choose It Over Kubernetes for Small & Medium Stacks — Legion Mind Architecture Dispatch

Key Architecture Takeaways

  • Eliminating single-point-of-failure restarts on single-node and multi-node VPS clusters.
  • Using Traefik label-based service discovery for dynamic zero-reload configuration.
  • Configuring rolling updates with health checks and rollback thresholds.
  • Achieving 99.99% uptime on cost-effective European dedicated hardware.

The Unspoken Cost of the Kubernetes Tax

Kubernetes is an incredible piece of software if you have 40 engineers, a dedicated platform operations team, and hundreds of microservices.

For almost everyone else, running managed Kubernetes (like EKS or GKE) is architectural overkill. Before deploying your very first customer container, you are already spending:

  • €70–€150/month just for the cloud control plane fee.
  • 2–4 GB of RAM per node consumed purely by kubelet, kube-proxy, and cluster management daemons.
  • Countless hours debugging YAML manifests, ingress controllers, CSI storage plugins, and cert-manager crashes.

Over the past three years, we have migrated dozens of web applications and internal tools to Docker Swarm on Hetzner Cloud. Here is why it works so well, how we architect clusters, and the exact configuration we use for zero-downtime rolling deployments.

Docker Swarm Clustered Topology on Hetzner
Figure 1: 3-Node Swarm manager/worker overlay architecture with Traefik ingress.

1. Why Docker Swarm Still Makes Sense in 2026

Docker Swarm is not dead—it is built directly into the Docker engine you already have installed. There is no separate binary to manage or complex certificate authority to maintain.

  • Zero Memory Overhead: Swarm’s management daemon consumes roughly 40MB of RAM. On a small 4GB or 8GB VPS, 95% of the memory remains available for your actual databases and applications.
  • Built-in Ingress Routing Mesh: Swarm includes native routing mesh and load balancing powered by Linux IPVS. Any incoming port published in Swarm mode automatically routes to active containers across any node in the cluster.
  • Encrypted Overlay Networks: A single flag (--opt encrypted) transparently encrypts all inter-node container communication using IPsec with zero manual configuration.
  • Predictable Hetzner Economics: Three CX32 Hetzner instances (8 vCPU, 16GB RAM each, in Falkenstein or Helsinki) cost around €36/month total. On AWS, comparable compute and networking easily surpass €250/month.

2. Ingress & Automated SSL with Traefik v3

Instead of manually maintaining Nginx configuration files and certbot cron jobs, we run Traefik as a global Swarm service on manager nodes. Traefik listens directly to the Docker socket, dynamically discovering new services through labels and issuing Let's Encrypt certificates automatically.

Here is a battle-tested traefik-stack.yml:

YAML
version: '3.8'

services:
  traefik:
    image: traefik:v3.1
    command:
      - "--providers.swarm=true"
      - "--providers.swarm.exposedbydefault=false"
      - "--entrypoints.web.address=:80"
      - "--entrypoints.websecure.address=:443"
      - "--entrypoints.web.http.redirections.entrypoint.to=websecure"
      - "--entrypoints.web.http.redirections.entrypoint.scheme=https"
      - "--certificatesresolvers.letsencrypt.acme.email=ops@legionmind.si"
      - "--certificatesresolvers.letsencrypt.acme.storage=/letsencrypt/acme.json"
      - "--certificatesresolvers.letsencrypt.acme.tlschallenge=true"
    ports:
      - target: 80
        published: 80
        mode: host
      - target: 443
        published: 443
        mode: host
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - traefik-certificates:/letsencrypt
    networks:
      - public-traefik-net
    deploy:
      mode: global
      placement:
        constraints:
          - node.role == manager
      restart_policy:
        condition: on-failure

networks:
  public-traefik-net:
    external: true

volumes:
  traefik-certificates:

Notice mode: host on ports 80 and 443. This preserves client real IP addresses in HTTP request headers (X-Forwarded-For), which is essential for rate limiting and fraud analysis.


3. The Secret to True Zero-Downtime Rolling Deploys

By default, updating a Docker service will tear down the old container and start the new one. During that 5-to-15 second boot window, incoming user requests return 502 Bad Gateway.

To prevent this, you must configure two settings in your application stack:

  1. update_config.order: start-first: Swarm spins up the new container first, waits for it to report healthy, and only then terminates the old container.
  2. Explicit container healthcheck: Swarm relies on Docker healthchecks to know when the new instance is ready to receive network sockets.
YAML
services:
  web-app:
    image: myregistry.com/app:v2.4
    networks:
      - public-traefik-net
    healthcheck:
      test: ["CMD-SHELL", "curl -f http://localhost:3000/api/health || exit 1"]
      interval: 5s
      timeout: 3s
      retries: 3
      start_period: 10s
    deploy:
      replicas: 3
      update_config:
        parallelism: 1
        delay: 5s
        order: start-first
        failure_action: rollback
      rollback_config:
        parallelism: 1
        order: stop-first
      labels:
        - "traefik.enable=true"
        - "traefik.http.routers.app.rule=Host(`app.legionmind.si`)"
        - "traefik.http.routers.app.entrypoints=websecure"
        - "traefik.http.routers.app.tls.certresolver=letsencrypt"
        - "traefik.http.services.app.loadbalancer.server.port=3000"

If the new container throws an exception or fails its health check during deployment, Swarm automatically aborts the rollout and rolls back to the previous healthy tag without dropping user traffic.

Sequential Container Cutover
Sequential task swap ensuring live connections are never dropped
Cluster Cost Comparison
Monthly cost comparison: Hetzner bare-metal vs AWS managed EKS

4. Disaster Recovery with Offsite BorgBackups

Running your own nodes means taking responsibility for stateful volume data. We pair every Swarm manager with an automated BorgBackup cron job targeting a Hetzner Storage Box.

Borg provides deduplication, client-side encryption (AES-256), and instant snapshot verification. Even with multiple database dumps and uploaded media volumes, incremental backups take under 30 seconds and occupy negligible disk space.

If a server catches fire or a cloud datacenter experiences an outage, a complete cluster can be restored on fresh hardware using our standard playbook in less than 30 minutes.

The Bottom Line

Don't let tech-industry hype dictate your infrastructure choices. If your stack fits comfortably on 2 to 5 nodes, Docker Swarm gives you 99.99% uptime, zero-downtime deployments, and complete architectural autonomy at a fraction of the cost.

Found this architecture guide useful?Share this dispatch with your engineering team
Share:𝕏 PostLinkedIn
Field Engineering Dispatches

Production Runbooks & Architecture Notes

Monthly technical dispatches covering Linux VPS hardening, strict DMARC deliverability, Next.js optimization, and cloud operations. Zero sales fluff.

Need direct help implementing this stack?

Our principal engineers audit and configure infrastructure with guaranteed SLAs.