The final design is credible only if its failure tests and redesign triggers are measurable.
Capstone: Defend a Distributed Data Architecture Under Partitions, Growth, Incidents, and Cost Constraints
Defend the final distributed-data architecture against partitions, hot keys, stale reads, node loss, compaction debt, duplicate events, credential compromise, bad deployments, restores, traffic growth, and cost pressure.
Defend a complete AtlasMart distributed-data architecture against a panel of correctness, availability, recovery, security, scale, operations, and cost scenarios.
Express every defense as an acceptance criterion with observable evidence, a safe failure test, and a rollback or degraded-mode path.
Identify the assumptions that make the current design valid and the measurable thresholds that would force redesign.
Finish the course with a decision packet that another engineer can audit, challenge, reproduce, and operate.
1. The review panel does not accept architecture slogans
The final AtlasMart architecture must answer a hostile but realistic review: What happens during a region partition? Which side may write? What if one key becomes 100× hotter? How stale may search be? What survives node loss? What happens when compaction debt consumes headroom? Can duplicate events charge a card twice? What can a stolen service credential administer? Can a bad deployment roll back? Has a restore actually been tested? What breaks at 5× growth? What happens if managed-service cost doubles?
A defensible design answers each question with an acceptance criterion, an observable signal, a safe test, and a rollback/degraded path. “The provider handles it,” “we use eventual consistency,” or “we use Kubernetes” are not answers.
2. Final AtlasMart architecture and failure boundaries
| Domain | Role/design | Key correctness or SLO | Failure/repair strategy |
|---|---|---|---|
| Orders/payments/inventory | Relational authoritative core | Invariant-preserving transaction boundary; idempotent commands | Single write authority under partition; tested backups/PITR |
| Catalog | Document authoritative aggregates | Versioned updates; bounded document growth | Restore + CDC replay; search independently rebuildable |
| Sessions/rate limits | Key-value ephemeral | Atomic key operations; defined TTL/eviction policy | Reauthenticate/reconstruct; never sole business truth |
| Telemetry | Wide-column/time-oriented authoritative history | Bounded partitions; ingest/backpressure SLO | Repair/rebalance with disk/headroom alarms |
| Fraud | Derived graph | Freshness SLO; traceable source evidence | Rebuild/reconcile; fail to manual review when stale |
| Search/vector | Derived retrieval infrastructure | Recall/relevance + freshness SLO | Rebuild from source stream; exact/degraded fallback where feasible |
| Synchronization | Outbox/CDC, at-least-once | Stable IDs, ordering scope, idempotent consumers | Replay + reconciliation; no end-to-end exactly-once claim |
3. Deliberately wrong approach: one global active-active store with no invariant map
A tempting capstone answer is “put everything in one globally distributed database and enable multi-region writes.” That avoids a diagram but not the tradeoffs. If checkout allows conflicting writes during a partition without an invariant-preserving strategy, stock can oversell. If search relevance, graph traversal, time-oriented telemetry and ad hoc financial reporting all share one generic model, each may pay unnecessary complexity or performance cost.
The repair is not automatically more products. It is explicit boundaries. Keep workloads together when one system meets them cleanly; split only when a different optimization target is material. For every split, pay the synchronization and operations cost consciously.
4. AtlasMart capstone lab: scenario defense
Python 3.13+ standard library only. This is a deterministic architecture-review simulator; it does not inject failures into real services.
from dataclasses import dataclass
@dataclass(frozen=True)
class Scenario:
name: str
evidence: str
pass_condition: bool
redesign_trigger: str
architecture = {
"orders_payments_inventory":"relational-authoritative",
"catalog":"document-authoritative",
"sessions_rate_limits":"key-value-ephemeral",
"telemetry":"wide-column-authoritative",
"fraud":"graph-derived",
"search":"search-derived",
"sync":"outbox+CDC-at-least-once",
}
scenarios=[
Scenario("region partition","checkout keeps one write authority; remote region degrades",True,"business requires active-active invariant-preserving checkout"),
Scenario("hot key","celebrity product cache is replicated/coalesced and origin rate-limited",True,"single key still exceeds safe origin/cache fanout"),
Scenario("stale read","search exposes source_version and falls back for correctness-critical lookup",True,"freshness SLO repeatedly breached"),
Scenario("node loss","capacity model retains survivor headroom",True,"survivor disk or throughput > policy threshold"),
Scenario("compaction debt","telemetry alerts on backlog/disk headroom and throttles ingest",True,"backlog growth outruns recovery window"),
Scenario("duplicate events","consumers dedup by stable event id and version",True,"side effects cannot be made idempotent"),
Scenario("credential compromise","service identity cannot administer cluster; break-glass is separate",True,"provider/service cannot express required privilege boundary"),
Scenario("bad deployment","progressive rollout + rollback + schema compatibility gate",True,"rollback cannot restore old reader/writer compatibility"),
Scenario("restore","isolated restore passes count/hash/invariant checks before cutover",True,"measured RTO/RPO misses objective"),
Scenario("traffic growth","capacity/partition model re-evaluated at 2x and 5x demand",True,"hot partitions or cost exceed approved envelope"),
Scenario("cost increase","TCO sensitivity includes egress, replicas, labor and exit",True,"cost range exceeds value or viable alternative crosses threshold"),
]
print("FINAL ATLASMART ARCHITECTURE")
for k,v in architecture.items(): print(k,"->",v)
print("\nREVIEW PANEL")
passed=0
for s in scenarios:
print(f"{s.name:22} PASS={s.pass_condition} | evidence={s.evidence}")
print(" redesign if:",s.redesign_trigger)
passed += int(s.pass_condition)
print("passed",passed,"of",len(scenarios))
assumptions=[
"checkout can use one authoritative write region during partition",
"search/fraud may be eventually consistent and rebuildable",
"telemetry queries are partition-local/time-bounded",
"session loss is survivable through reauthentication",
"event delivery is at-least-once, not magically exactly-once",
]
print("\nASSUMPTIONS THAT MUST STAY TRUE")
for a in assumptions: print("-",a)
print("course-complete artifact: architecture + tests + evidence + redesign triggers, not a product shopping list")
The program prints the final store roles and eleven review scenarios. Each scenario includes evidence and a redesign trigger. The goal is not the printed PASS label itself; it is the discipline of stating what must be proven and what change would invalidate the architecture.
5. Measurable acceptance criteria and redesign triggers
Translate every assumption into a monitor or periodic test. Examples: checkout p99 under failover; zero duplicate payment side effects in retry tests; search freshness below its SLO percentile; telemetry partitions below the planned size/throughput envelope; survivor disk utilization below the recovery threshold; restore RPO/RTO inside objectives; least-privilege policy with audited break-glass; reconciliation drift below a defined rate; and TCO within an approved low/base/high envelope.
Write redesign triggers beside them. If active-active checkout becomes a business requirement, the single-authority partition strategy must be replaced. If search becomes a regulatory decision source, its current eventual/rebuildable classification may be insufficient. If telemetry access becomes cross-device ad hoc analytics, the partition model needs another system or projection. If a managed service loses a required region, guarantee, feature or acceptable cost, rerun the selection matrix.
6. Course completion: keep the evidence alive
This course does not end with a universal technology stack. It ends with a method: start from workload and invariants; reason about distributed failure; make time, replication, partitioning and storage mechanics observable; choose the weakest consistency that still preserves correctness; make derived data rebuildable; test recovery; benchmark realistic distributions; price the full operating model; and record the assumptions that make the design valid.
The durable capstone artifact is a versioned architecture decision packet: access-pattern notebook, selection matrices, ownership map, migration and rollback plans, synchronization contracts, SLO/RPO/RTO tables, security and governance controls, backup/restore evidence, benchmark/capacity results, TCO ranges, failure-injection results, incident runbooks and redesign triggers. Another engineer should be able to reproduce the reasoning without trusting the original designer’s intuition.
Check your understanding
- What makes a capstone architecture “defensible”?
- Why must redesign triggers be written before an incident?
- How should the architecture handle a region partition for checkout?
- What does passing a restore test prove?
- What is the final course artifact?
Review the answers
1. Every major choice is tied to workload/invariant evidence, failure assumptions, measurable SLO/RPO/RTO criteria, tests, operational ownership, cost ranges, and a rollback or degraded mode.
2. They prevent teams from treating assumptions as permanent truths and make scaling, regulation, cost, or correctness thresholds actionable.
3. According to its stated invariant and authority model—for this capstone, preserve one authoritative write side and degrade the other rather than allow conflicting stock/payment decisions.
4. That a particular backup, procedure, dependency set and validation worked under the tested conditions; it does not prove every future restore or application behavior automatically.
5. A living decision packet: workload/invariant inventory, selection matrix, ownership/data-flow diagrams, migration plan, SLO/RPO/RTOs, security/recovery controls, capacity/TCO model, failure tests, evidence, and redesign triggers.
References
Course-wide foundational and current implementation anchors. Product references are examples, not mandatory dependencies:
- PostgreSQL 18 — Transactions — Current official transaction semantics used as one relational implementation anchor.
- PostgreSQL 18 — Logical Replication — Current official snapshot-plus-change replication model relevant to migration and synchronization.
- Debezium 3.6 release series — Latest stable 3.6 line; 3.6.1.Final was released 2026-08-04. Optional CDC implementation anchor only.
- Apache Cassandra downloads — Current official release page listing Cassandra 5.0.9 as latest GA on 2026-08-07.
- Spanner: Google’s Globally-Distributed Database — Primary research showing how global transactions depend on explicit timing/replication assumptions.
- In Search of an Understandable Consensus Algorithm (Raft) — Primary consensus reference revisited for control-plane agreement assumptions.
- Paxos Made Simple — Primary reference for quorum-intersection consensus reasoning.
- Sagas — Original work on long-lived transactions and compensating actions, revisited for distributed workflow design.