The final design is credible only if its failure tests and redesign triggers are measurable.

Capstone: Defend a Distributed Data Architecture Under Partitions, Growth, Incidents, and Cost Constraints

Defend the final distributed-data architecture against partitions, hot keys, stale reads, node loss, compaction debt, duplicate events, credential compromise, bad deployments, restores, traffic growth, and cost pressure.

Advanced170–220 minutesCourse capstone defensePython 3.13+ · standard librarySafe scenario simulationLast reviewed: August 2026
01

Defend a complete AtlasMart distributed-data architecture against a panel of correctness, availability, recovery, security, scale, operations, and cost scenarios.

02

Express every defense as an acceptance criterion with observable evidence, a safe failure test, and a rollback or degraded-mode path.

03

Identify the assumptions that make the current design valid and the measurable thresholds that would force redesign.

04

Finish the course with a decision packet that another engineer can audit, challenge, reproduce, and operate.

1. The review panel does not accept architecture slogans

The final AtlasMart architecture must answer a hostile but realistic review: What happens during a region partition? Which side may write? What if one key becomes 100× hotter? How stale may search be? What survives node loss? What happens when compaction debt consumes headroom? Can duplicate events charge a card twice? What can a stolen service credential administer? Can a bad deployment roll back? Has a restore actually been tested? What breaks at 5× growth? What happens if managed-service cost doubles?

A defensible design answers each question with an acceptance criterion, an observable signal, a safe test, and a rollback/degraded path. “The provider handles it,” “we use eventual consistency,” or “we use Kubernetes” are not answers.

2. Final AtlasMart architecture and failure boundaries

Domain Role/design Key correctness or SLO Failure/repair strategy
Orders/payments/inventory Relational authoritative core Invariant-preserving transaction boundary; idempotent commands Single write authority under partition; tested backups/PITR
Catalog Document authoritative aggregates Versioned updates; bounded document growth Restore + CDC replay; search independently rebuildable
Sessions/rate limits Key-value ephemeral Atomic key operations; defined TTL/eviction policy Reauthenticate/reconstruct; never sole business truth
Telemetry Wide-column/time-oriented authoritative history Bounded partitions; ingest/backpressure SLO Repair/rebalance with disk/headroom alarms
Fraud Derived graph Freshness SLO; traceable source evidence Rebuild/reconcile; fail to manual review when stale
Search/vector Derived retrieval infrastructure Recall/relevance + freshness SLO Rebuild from source stream; exact/degraded fallback where feasible
Synchronization Outbox/CDC, at-least-once Stable IDs, ordering scope, idempotent consumers Replay + reconciliation; no end-to-end exactly-once claim

3. Deliberately wrong approach: one global active-active store with no invariant map

A tempting capstone answer is “put everything in one globally distributed database and enable multi-region writes.” That avoids a diagram but not the tradeoffs. If checkout allows conflicting writes during a partition without an invariant-preserving strategy, stock can oversell. If search relevance, graph traversal, time-oriented telemetry and ad hoc financial reporting all share one generic model, each may pay unnecessary complexity or performance cost.

The repair is not automatically more products. It is explicit boundaries. Keep workloads together when one system meets them cleanly; split only when a different optimization target is material. For every split, pay the synchronization and operations cost consciously.

4. AtlasMart capstone lab: scenario defense

Mandatory lab environment

Python 3.13+ standard library only. This is a deterministic architecture-review simulator; it does not inject failures into real services.

python · AtlasMart deterministic simulation
from dataclasses import dataclass

@dataclass(frozen=True)
class Scenario:
    name: str
    evidence: str
    pass_condition: bool
    redesign_trigger: str

architecture = {
    "orders_payments_inventory":"relational-authoritative",
    "catalog":"document-authoritative",
    "sessions_rate_limits":"key-value-ephemeral",
    "telemetry":"wide-column-authoritative",
    "fraud":"graph-derived",
    "search":"search-derived",
    "sync":"outbox+CDC-at-least-once",
}

scenarios=[
Scenario("region partition","checkout keeps one write authority; remote region degrades",True,"business requires active-active invariant-preserving checkout"),
Scenario("hot key","celebrity product cache is replicated/coalesced and origin rate-limited",True,"single key still exceeds safe origin/cache fanout"),
Scenario("stale read","search exposes source_version and falls back for correctness-critical lookup",True,"freshness SLO repeatedly breached"),
Scenario("node loss","capacity model retains survivor headroom",True,"survivor disk or throughput > policy threshold"),
Scenario("compaction debt","telemetry alerts on backlog/disk headroom and throttles ingest",True,"backlog growth outruns recovery window"),
Scenario("duplicate events","consumers dedup by stable event id and version",True,"side effects cannot be made idempotent"),
Scenario("credential compromise","service identity cannot administer cluster; break-glass is separate",True,"provider/service cannot express required privilege boundary"),
Scenario("bad deployment","progressive rollout + rollback + schema compatibility gate",True,"rollback cannot restore old reader/writer compatibility"),
Scenario("restore","isolated restore passes count/hash/invariant checks before cutover",True,"measured RTO/RPO misses objective"),
Scenario("traffic growth","capacity/partition model re-evaluated at 2x and 5x demand",True,"hot partitions or cost exceed approved envelope"),
Scenario("cost increase","TCO sensitivity includes egress, replicas, labor and exit",True,"cost range exceeds value or viable alternative crosses threshold"),
]

print("FINAL ATLASMART ARCHITECTURE")
for k,v in architecture.items(): print(k,"->",v)
print("\nREVIEW PANEL")
passed=0
for s in scenarios:
    print(f"{s.name:22} PASS={s.pass_condition} | evidence={s.evidence}")
    print("  redesign if:",s.redesign_trigger)
    passed += int(s.pass_condition)
print("passed",passed,"of",len(scenarios))

assumptions=[
"checkout can use one authoritative write region during partition",
"search/fraud may be eventually consistent and rebuildable",
"telemetry queries are partition-local/time-bounded",
"session loss is survivable through reauthentication",
"event delivery is at-least-once, not magically exactly-once",
]
print("\nASSUMPTIONS THAT MUST STAY TRUE")
for a in assumptions: print("-",a)
print("course-complete artifact: architecture + tests + evidence + redesign triggers, not a product shopping list")
Expected evidence

The program prints the final store roles and eleven review scenarios. Each scenario includes evidence and a redesign trigger. The goal is not the printed PASS label itself; it is the discipline of stating what must be proven and what change would invalidate the architecture.

5. Measurable acceptance criteria and redesign triggers

Translate every assumption into a monitor or periodic test. Examples: checkout p99 under failover; zero duplicate payment side effects in retry tests; search freshness below its SLO percentile; telemetry partitions below the planned size/throughput envelope; survivor disk utilization below the recovery threshold; restore RPO/RTO inside objectives; least-privilege policy with audited break-glass; reconciliation drift below a defined rate; and TCO within an approved low/base/high envelope.

Write redesign triggers beside them. If active-active checkout becomes a business requirement, the single-authority partition strategy must be replaced. If search becomes a regulatory decision source, its current eventual/rebuildable classification may be insufficient. If telemetry access becomes cross-device ad hoc analytics, the partition model needs another system or projection. If a managed service loses a required region, guarantee, feature or acceptable cost, rerun the selection matrix.

6. Course completion: keep the evidence alive

This course does not end with a universal technology stack. It ends with a method: start from workload and invariants; reason about distributed failure; make time, replication, partitioning and storage mechanics observable; choose the weakest consistency that still preserves correctness; make derived data rebuildable; test recovery; benchmark realistic distributions; price the full operating model; and record the assumptions that make the design valid.

The durable capstone artifact is a versioned architecture decision packet: access-pattern notebook, selection matrices, ownership map, migration and rollback plans, synchronization contracts, SLO/RPO/RTO tables, security and governance controls, backup/restore evidence, benchmark/capacity results, TCO ranges, failure-injection results, incident runbooks and redesign triggers. Another engineer should be able to reproduce the reasoning without trusting the original designer’s intuition.

Check your understanding

  1. What makes a capstone architecture “defensible”?
  2. Why must redesign triggers be written before an incident?
  3. How should the architecture handle a region partition for checkout?
  4. What does passing a restore test prove?
  5. What is the final course artifact?
Review the answers

1. Every major choice is tied to workload/invariant evidence, failure assumptions, measurable SLO/RPO/RTO criteria, tests, operational ownership, cost ranges, and a rollback or degraded mode.

2. They prevent teams from treating assumptions as permanent truths and make scaling, regulation, cost, or correctness thresholds actionable.

3. According to its stated invariant and authority model—for this capstone, preserve one authoritative write side and degrade the other rather than allow conflicting stock/payment decisions.

4. That a particular backup, procedure, dependency set and validation worked under the tested conditions; it does not prove every future restore or application behavior automatically.

5. A living decision packet: workload/invariant inventory, selection matrix, ownership/data-flow diagrams, migration plan, SLO/RPO/RTOs, security/recovery controls, capacity/TCO model, failure tests, evidence, and redesign triggers.

References

Course-wide foundational and current implementation anchors. Product references are examples, not mandatory dependencies:

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.