Architecture

One daemon per node. An encrypted mesh between them. A single elected leader at the edge.

Components

Every node runs the same single Rust binary. Inside it, these subsystems cooperate:

ComponentRole
Daemon coreOwns startup, configuration, the shutdown coordinator, and state persistence across restarts.
Mesh layerWireGuard L3 overlay between nodes (e.g. 10.8.0.x). Keys are auto-generated if not configured.
Gossip membershipSWIM-style protocol: nodes discover and health-track each other and propagate state changes, converging in seconds.
Leader electionBully algorithm with configurable per-node priorities; exactly one leader holds the ingress entrypoint.
RegistryThe fleet's shared view of nodes, routes, and service state.
ProxyIngress engine running in one of three modes: L7 Direct, L4 SNI passthrough, or Managed Coolify.
DiscoveryWatches the Docker socket; labeled containers become routes automatically alongside static config routes.
Failover duplicatorReplicates marked workloads to peer nodes on failure via the container driver, with failback and cooldowns.
IPCUnix socket (/tmp/bridge.sock) the bridge CLI uses to talk to the running daemon.
DashboardEmbedded operations console (dark, Cloudflare-inspired) on 127.0.0.1:9090 by default, with a JSON API and WebSocket event stream.

Node lifecycle

Gossip moves every node through a small state machine:

Alive → Suspected → Dead → Recovered

A node that stops responding is first marked Suspected; continued silence marks it Dead and the fleet routes around it. When it comes back it is marked Recovered and rejoins routing. Because this spreads by gossip rather than central polling, membership converges in seconds without a single point of failure in the control plane itself.

The leader's role

The bully algorithm elects the highest-priority alive node as leader. The leader holds the fleet's ingress entrypoint: the Cloudflare Tunnel, the DNS failover target, or the floating IP, depending on the configured handoff tier. All nodes still serve traffic: the leader terminates ingress and routes requests across the mesh to whichever node hosts the target service.

Request flow

The three proxy modes

L7 Direct

BRIDGE acts as an HTTP reverse proxy per hostname. It terminates the connection, matches the hostname to a route, and forwards to the upstream over the mesh. This is the default mode and the one the examples use.

L4 SNI Passthrough

BRIDGE reads the TLS SNI hostname and forwards the raw stream without terminating TLS: certificates live on the backends. Use this when you need end-to-end TLS to your applications.

Managed Coolify

BRIDGE delegates ingress to an existing Coolify installation on the fleet instead of proxying itself, while still providing mesh, membership, election, and failover.

Ingress handoff tiers

Handoff decides how the public entrypoint survives a leader change. Configure it under [handoff]; see Configuration.

TierHow it worksWhen to use it
none No handoff; the entrypoint is fixed to one node. Single-node or standalone deployments.
dns Health-checked DNS failover moves the record to the new leader. Any provider with a supported DNS setup; simple and portable.
tunnel The leader holds the Cloudflare Tunnel; on re-election the new leader takes it over. No public ports or floating IPs available; nodes behind NAT.
floating_ip A floating IP is reassigned to the new leader. Providers that offer reassignable floating IPs.

Failure scenarios

Leader loss

Gossip marks the leader Suspected, then Dead. The remaining nodes run the bully election; the highest-priority alive node wins and performs the configured handoff: taking over the Cloudflare Tunnel or moving DNS / the floating IP. Traffic resumes against the new leader without manual intervention.

Backend or node failure

Dead targets drop out of the consistent-hash ring, so requests redistribute to the remaining targets of a service. For services marked for replication, the failover duplicator spawns the workload on a healthy peer via the container driver, subject to placement rules and cooldowns. When the original node recovers, failback is manual or automatic depending on the configured failback_mode.

Graceful shutdown

On shutdown the daemon's shutdown coordinator drains work in order and persists state, so a restart resumes from the last known fleet state rather than a cold start.