Writing

Building a Resilient Site-Aware Synchronization

How to leverage RabbitMQ Cluster along with Outbox Pattern

Summary

The setting is one I work in: several on-premise sites, each with its own local database, some on PostgreSQL and some on Oracle, under network constraints that ruled out any cloud-native design. Every site had to keep operating on its own, even through an outage lasting days, and the data still had to converge across sites once the link came back.

Classic replication tools such as GoldenGate or SymmetricDS were considered and set aside: they bring their own operational complexity, and the synchronisation conflicts have to be handled anyway. I chose an event-driven design instead. Each write goes into a local outbox table in the same transaction as the business data; a dispatcher publishes those events to a three-node RabbitMQ cluster with quorum queues; a fanout exchange delivers every event to every site, including the one that emitted it, which is how a site knows its event has propagated.

Two properties make it hold. Every event carries a globally unique identifier, so a consumer that sees it twice applies it once. And one site plays the golden source: it archives every event, which gives an audit trail, a retention window and the ability to replay. Conflicts go to a dead-letter queue for a human to look at. An optional confirmation event, sent back by each consuming site, makes replay detection and debugging easier still.

The article closes on a choice that is often made too quickly: clustering, federation or shovel to connect RabbitMQ brokers, each with different guarantees on ordering and delivery. The answer depends on the business need, not on habit. The repository that accompanies the article runs the cluster with Docker Compose and lets you take a broker down to see the rest keep going.

Key ideas

  • Sites that must survive days without network need autonomy first and consistency eventually; classic replication adds complexity without removing the conflicts.
  • An outbox written in the same transaction as the business data, then published to RabbitMQ, so that nothing is lost while the broker is unreachable.
  • A fanout exchange delivers each event to every site, including the emitter, which is how a site learns that its event has propagated.
  • A globally unique event id makes consumers idempotent; a golden-source site archives everything for audit and replay.
  • Clustering, federation or shovel between brokers is a business decision, not a habit.

Why I wrote this

In our case we manage several on-premise sites, each with its own local database, some on PostgreSQL and others on Oracle. A cloud-native solution was not viable under strict network constraints, and every site had to operate independently, without relying on real-time connectivity.

Companion repositories