Chapter 02 · Distributed Systems Foundations: Nodes, Networks, Failure, and State

Failure Domains: Process, Host, Rack, Zone, Region, Provider, and Human Operations

Map replicas onto host, rack, zone, region, provider, and human/change failure domains and prove why replication count is not the same as independent redundancy.

Beginner95–115 minutesFailure-domain placement simulatorVendor-neutral · no cloud account requiredLast reviewed: August 2026

Learning outcomes

AtlasMart proudly reports “three replicas.” A rack power failure removes two. A zone network event removes all three. Later, a perfectly healthy multi-zone deployment is taken down by one bad configuration pushed to every data node. The problem is not replication count; it is shared fate. Reliability claims must name the failure domain they are designed to tolerate.

01

Define process, host, rack, zone, region, provider, and human/operational failure domains as progressively different forms of shared fate.

02

Evaluate replica placement by the number of replicas lost under an injected host, rack, zone, region, or provider failure.

03

Explain why three replicas in one zone are not three independent zone-failure protections.

04

Treat configuration, credential, schema, automation, and operator mistakes as correlated failure domains that physical placement alone does not fix.

05

Connect topology-aware placement to latency, cost, recovery, capacity headroom, and change-management tradeoffs without prescribing universal multi-region or multi-provider designs.

Failure domain

A failure domain is a set of components that can fail together because they share a dependency, location, control plane, change, credential, resource, or operator action. Domains are architecture-specific. “Rack,” “zone,” and “region” are useful concepts, but their exact guarantees are provider/datacenter specific and must be verified rather than inferred from names.

1. Independence is the scarce resource

Replication creates multiple copies. Reliability improves only when those copies are sufficiently independent for the failure you care about. Three processes on one host survive one process crash but not host power loss. Three hosts in one rack may survive a host failure but not rack power/network loss. Three zones in one region may survive a zonal event but not every regional or provider-wide event. Copies in two providers may reduce some shared-infrastructure risks while introducing network, identity, consistency, operations, observability, and support complexity.

Domain Shared-fate examples A placement question
Process Crash, memory leak, bad thread/task state Are replicas actually different processes, and can one restart without killing all?
Host / VM Kernel panic, host power, hypervisor, local disk/controller Do replicas share the same physical/virtual host or attached failure dependency?
Rack / datacenter segment Top-of-rack switch, power distribution, cooling Are replicas spread across independent racks/segments where that guarantee exists?
Zone / facility group Facility power/network/control-plane incidents Does the platform document zones as meaningful isolation boundaries for this service?
Region Regional network/control-plane/disaster event Is regional survival required by the AtlasMart RTO/RPO and latency model?
Provider / organization Provider-wide control plane, account/identity, billing, policy Would multi-provider independence justify the large operational/data-consistency complexity?
Human/change Bad config, schema migration, credential revoke, automation bug, destructive command Can one action touch every replica/back-up at once? Are rollouts staged and fenced?

2. Topology-aware placement turns a requirement into a rule

Suppose the order partition has replication factor three. If AtlasMart requires survival of one zone failure, the placement policy must ensure the three replicas do not all share one zone, and the protocol must remain able to serve the required operation with the surviving replicas. Merely setting replication_factor=3 is incomplete because placement and acknowledgement/read rules determine what that number buys.

Topology-aware databases often expose rack/zone/region labels, placement constraints, awareness attributes, replica-selection policies, or managed-service regional settings. The names and guarantees vary. A correct design starts from the failure requirement and verifies that the product's placement mechanism and underlying infrastructure actually correspond to independent domains.

3. Run four placement layouts against correlated failures

The deterministic lab compares four three-replica layouts. host_diverse_same_zone spreads processes across hosts but keeps all copies in east-a. zone_diverse uses three zones in region east. region_diverse puts one copy in west. multi_provider places one copy in a second provider. The simulator then removes every node sharing a selected host, rack, zone, region, or provider value.

python · inject host/rack/zone/region/provider failures into replica placements
from dataclasses import dataclass@dataclass(frozen=True)class Place:    host: str    rack: str    zone: str    region: str    provider: strlayouts = {    "host_diverse_same_zone": {        "n1": Place("h1", "r1", "east-a", "east", "provider-a"),        "n2": Place("h2", "r1", "east-a", "east", "provider-a"),        "n3": Place("h3", "r2", "east-a", "east", "provider-a"),    },    "zone_diverse": {        "n1": Place("h1", "r1", "east-a", "east", "provider-a"),        "n2": Place("h2", "r2", "east-b", "east", "provider-a"),        "n3": Place("h3", "r3", "east-c", "east", "provider-a"),    },    "region_diverse": {        "n1": Place("h1", "r1", "east-a", "east", "provider-a"),        "n2": Place("h2", "r2", "east-b", "east", "provider-a"),        "n3": Place("h9", "r9", "west-a", "west", "provider-a"),    },    "multi_provider": {        "n1": Place("h1", "r1", "east-a", "east", "provider-a"),        "n2": Place("h2", "r2", "east-b", "east", "provider-a"),        "n3": Place("h7", "r7", "north-a", "north", "provider-b"),    },}faults = [    ("host", "h1"),    ("rack", "r1"),    ("zone", "east-a"),    ("region", "east"),    ("provider", "provider-a"),]def survivors(layout, dimension, value):    return [n for n,p in layout.items() if getattr(p, dimension) != value]for name, layout in layouts.items():    print(f"layout={name}")    for node, place in layout.items():        print(f"  {node} host={place.host} rack={place.rack} zone={place.zone} region={place.region} provider={place.provider}")    for dimension, value in faults:        live = survivors(layout, dimension, value)        print(f"  inject {dimension}={value:<10s} survivors={live} lost={3-len(live)}")print("human_failure=bad config selector role=data reaches n1,n2,n3 in every layout -> survivors=[]")print("lesson=replica count is not failure independence; placement and change blast radius must match the failure you claim to tolerate")

Verified deterministic output

text · replica-survival evidence by failure domain
layout=host_diverse_same_zone  n1 host=h1 rack=r1 zone=east-a region=east provider=provider-a  n2 host=h2 rack=r1 zone=east-a region=east provider=provider-a  n3 host=h3 rack=r2 zone=east-a region=east provider=provider-a  inject host=h1         survivors=['n2', 'n3'] lost=1  inject rack=r1         survivors=['n3'] lost=2  inject zone=east-a     survivors=[] lost=3  inject region=east       survivors=[] lost=3  inject provider=provider-a survivors=[] lost=3layout=zone_diverse  n1 host=h1 rack=r1 zone=east-a region=east provider=provider-a  n2 host=h2 rack=r2 zone=east-b region=east provider=provider-a  n3 host=h3 rack=r3 zone=east-c region=east provider=provider-a  inject host=h1         survivors=['n2', 'n3'] lost=1  inject rack=r1         survivors=['n2', 'n3'] lost=1  inject zone=east-a     survivors=['n2', 'n3'] lost=1  inject region=east       survivors=[] lost=3  inject provider=provider-a survivors=[] lost=3layout=region_diverse  n1 host=h1 rack=r1 zone=east-a region=east provider=provider-a  n2 host=h2 rack=r2 zone=east-b region=east provider=provider-a  n3 host=h9 rack=r9 zone=west-a region=west provider=provider-a  inject host=h1         survivors=['n2', 'n3'] lost=1  inject rack=r1         survivors=['n2', 'n3'] lost=1  inject zone=east-a     survivors=['n2', 'n3'] lost=1  inject region=east       survivors=['n3'] lost=2  inject provider=provider-a survivors=[] lost=3layout=multi_provider  n1 host=h1 rack=r1 zone=east-a region=east provider=provider-a  n2 host=h2 rack=r2 zone=east-b region=east provider=provider-a  n3 host=h7 rack=r7 zone=north-a region=north provider=provider-b  inject host=h1         survivors=['n2', 'n3'] lost=1  inject rack=r1         survivors=['n2', 'n3'] lost=1  inject zone=east-a     survivors=['n2', 'n3'] lost=1  inject region=east       survivors=['n3'] lost=2  inject provider=provider-a survivors=['n3'] lost=2human_failure=bad config selector role=data reaches n1,n2,n3 in every layout -> survivors=[]lesson=replica count is not failure independence; placement and change blast radius must match the failure you claim to tolerate

Notice the specificity. Zone-diverse placement loses only one replica when east-a fails, but still loses all three when the entire east region fails. Region-diverse placement keeps one copy after region east is lost, but whether that one copy can safely serve writes is a separate protocol question. Multi-provider placement keeps one copy after provider-a fails, but that does not make multi-provider architecture a default recommendation.

4. Deliberately wrong approach: count replicas, ignore shared fate

A team sees three database containers and writes “HA: 3 replicas” in the architecture document. All three containers run on one host or one zone because the scheduler had no anti-affinity/topology rule. A single infrastructure event removes every copy. The root cause is not “replication failed”; replication did exactly what was configured. The architecture failed to connect the replication count to an independence requirement.

The repair is to write the reliability claim in testable form: “for order data, one zone may be unavailable and the service must still preserve the documented write/read invariant within the SLO.” Then enforce topology constraints, maintain enough capacity in surviving domains, and game-day the failure. If the requirement also includes regional disaster, design and test that separately; do not assume a zone-aware configuration implies region survival.

5. Human operations create a cross-topology failure domain

The final line of the lab applies one bad configuration selector to role=data. It reaches n1, n2, and n3 in every layout. Physical diversity does not help because the correlated dependency is the deployment mechanism. The same is true for a bad schema migration, expired shared certificate, deleted account, compromised administrator, destructive automation, or corrupted software version distributed everywhere.

Reduce change blast radius with staged/canary rollout, health gates, per-zone/region sequencing, automated rollback where safe, peer review for destructive operations, immutable/audited infrastructure definitions, independent backup credentials/accounts, and tested break-glass procedures. The exact controls depend on the system. The key idea is to ask “what single action can still touch all copies?”

Backup is a different independence problem.

Live replicas intentionally copy legitimate writes and can also copy accidental deletes, corruption, or bad application behavior. Backup/restore independence, immutability, retention, and credential separation are treated in Chapter 22. Do not label replication as backup.

6. Wider failure tolerance has latency and cost consequences

Spreading synchronous acknowledgements across distant failure domains can increase write latency because the request must wait for network round trips and durable work outside the local zone/region. Asynchronous remote copies can reduce request latency but introduce a data-loss/freshness window if the primary region fails before replication catches up. Multi-region and multi-provider systems also multiply networking, egress, observability, security, deployment, capacity, and incident-response complexity.

Therefore topology is an optimization problem constrained by business RPO/RTO, consistency/invariants, user latency, regulation/data residency, and operational skill. “More regions” is not a free reliability toggle. Sometimes a single-region multi-zone system plus tested independent backup satisfies the requirement more safely than an under-operated active-active global design.

7. Production judgment and bridge to SLOs

Every redundancy claim should name the unit of failure: process, host, rack, zone, region, provider, account/control plane, or human change. Record where replicas live, how many are required for each operation, what capacity remains after failure, what latency changes in degraded topology, and how failover/recovery is tested.

Lesson 5 turns these architecture claims into measurable service objectives. “Survives a zone failure” is still vague until AtlasMart states which reads/writes remain available, what latency/freshness is acceptable, and how much data/time loss is allowed during recovery.

Verification checklist

  • The same replication factor produces different survivability under different placements.
  • The host-diverse layout still loses all copies under one zone failure.
  • The zone-diverse layout survives one modeled zone failure but not a whole-region failure.
  • The simulator treats provider failure separately from region failure.
  • A human/configuration failure can remove all replicas regardless of physical placement.
  • The lesson does not claim multi-region or multi-provider is universally better.

Check your understanding

  1. Why is replica count alone not a reliability guarantee?
  2. What is a failure domain?
  3. What additional requirement is needed after “spread replicas across zones”?
  4. Why can a bad configuration defeat perfect physical placement?
  5. Why might synchronous cross-region replication hurt a latency objective?
Review the answers

The replicas may share the same host/rack/zone/provider/control plane or be governed by acknowledgement rules that do not tolerate the desired failure.

It is a set of components that can fail together because they share some dependency, location, control plane, change, or operator action.

You must state and test which operation/invariant must remain available after a zone loss, with enough surviving capacity and correct protocol behavior.

The shared failure dependency is the change mechanism itself; one command or credential can affect every geographically separated copy.

The request may have to wait for longer network round trips and remote durable work before acknowledgement, increasing normal and tail latency.

Authoritative references

  • AWS Fault Isolation Boundaries — Official cloud architecture material illustrating zones, regions, control/data planes, and isolation-boundary reasoning; vendor-specific details are examples, not universal definitions.
  • Google SRE: Addressing Cascading Failures — Operational guidance on correlated failures, overload, retries, and safe rollouts/capacity.
  • Google SRE: Testing for Reliability — Guidance for testing failure behavior and reliability assumptions rather than relying only on architecture diagrams.
  • NIST SP 800-34 Rev. 1 — Contingency-planning guidance connecting recovery strategy, alternate resources, testing, and organizational failure planning.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.