Chapter 01 · Why NoSQL Exists: Workloads, Scale, Flexibility, and Polyglot Persistence
Build a Decision Notebook: Workload, Data Volume, Consistency, Availability, Recovery, and Cost Requirements
Create the reusable AtlasMart decision notebook with measurable workload, latency, consistency, availability, recovery, security, skill, and cost requirements before choosing any database technology.
Learning outcomes
The last lesson of Chapter 01 creates the artifact that every later chapter will refine: the AtlasMart distributed-data decision notebook. A decision notebook is not a vendor scorecard. It is a set of measurable assumptions and hard constraints recorded before a technology is chosen, so later experiments can prove or falsify the architecture.
Record read/write paths, volumes, key distributions, latency percentiles, consistency/invariant requirements, retention, and growth.
Define service-level objectives (SLOs), recovery point objective (RPO), and recovery time objective (RTO) in operationally testable terms.
Capture security, tenant isolation, data classification, operational skills, licensing, and cost as first-class database requirements.
Use family-level design questions as a filter without pretending they are a benchmark or automatic product selector.
Create a stable AtlasMart baseline that Chapters 02–25 can update as failure, replication, partitioning, storage, and cloud mechanisms are learned.
Architecture arguments become tractable when someone can point to a number or invariant that would change the decision. “MongoDB is flexible” or “Cassandra scales” is not falsifiable. “Catalog p95 under 80 ms from the target region at 12,000 reads/s with 55% annual growth and a documented hot-product skew” can be tested.
1. Translate business language into measurable workload statements
A database requirement should identify operation, population, distribution, objective, and failure scope. “Checkout must be fast” becomes “for successful checkout writes under declared payload/concurrency conditions, p95 application-observed latency must stay below the target while preserving payment idempotency and inventory rules.” “Search can be stale” becomes “99% of committed catalog changes become query-visible in the search index within the freshness SLO; checkout never trusts search for authorization or current inventory.”
The notebook intentionally starts with rough planning numbers. They are not measurements. Every estimate should carry an owner and later be replaced with telemetry, load-test evidence, or business forecasts.
| Notebook dimension | AtlasMart question | Evidence to collect later |
|---|---|---|
| Read/write rate | What are steady and peak operations/s by path? | Production telemetry or load model |
| Key/value distribution | Are tenants/products uniform or Zipf/celebrity-skewed? | Key-frequency histogram/top-N keys |
| Latency | Which p50/p95/p99 target from which client region? | End-to-end histogram with retries/timeouts included |
| Consistency/invariants | Which facts must never conflict or double-apply? | Concurrency tests, histories, domain rules |
| Durability/recovery | What data loss window is tolerable and how fast must restore complete? | RPO/RTO exercise and restore evidence |
| Security/governance | What tenant/PII/admin boundaries exist? | Threat model, access tests, audit/retention policy |
| Cost/skills | What can the team operate during incidents? | On-call capability, TCO model, licensing/cloud constraints |
2. SLO, RPO, and RTO describe different failure expectations
A service-level indicator (SLI) is a measured property such as successful request latency, error rate, or freshness. An SLO is a target for that indicator over a defined window. For databases, avoid one global “99.9% available” number: reads, writes, freshness, and durability can fail differently.
Recovery point objective (RPO) is the maximum tolerable data-loss window after a recovery scenario. Recovery time objective (RTO) is the target time to restore the required service capability. Replication may improve availability and reduce some failure recovery time, but it does not replace independent backup because accidental deletion, corruption, or malicious change can replicate too.
Record p95/p99 rather than only averages. Averages can hide hot shards, remote-region requests, retries, pauses, and queueing. Later benchmarking chapters will require the full load model and will discuss coordinated omission.
3. Build and validate the AtlasMart notebook
Save the script as atlasmart_decision_notebook.py.
It prints a JSON baseline and verifies that several hard fields
are present. The values are planning assumptions for the
fictional course domain; they are not vendor sizing
recommendations.
import jsonnotebook={ 'domain':'AtlasMart', 'workloads':{ 'checkout':{'writes_per_s_peak':450,'read_write_ratio':'2:1','p95_ms_target':180,'invariants':['one payment capture per idempotency key','inventory must not go below reserved safety floor']}, 'catalog':{'reads_per_s_peak':12000,'writes_per_s_peak':35,'p95_ms_target':80,'shape':'heterogeneous product attributes'}, 'search':{'reads_per_s_peak':9000,'freshness_slo_seconds':30,'source_of_truth':False}, 'sessions':{'reads_per_s_peak':18000,'writes_per_s_peak':10000,'ttl_minutes':45,'reconstructable':True} }, 'data':{'orders_year_1_gb':700,'catalog_gb':90,'search_index_multiplier_estimate':1.6,'annual_growth_percent':55,'key_skew_risk':'high for celebrity products'}, 'reliability':{'checkout_availability_target':'99.95%','catalog_availability_target':'99.9%','order_rpo':'<= 1 minute','order_rto':'<= 30 minutes','search_rpo':'rebuildable from source + event log','regional_failure_test':'quarterly'}, 'security':{'tenant_isolation':'authorization + tenant-scoped keys; key alone is not auth','pii':['customer name','email','shipping address'],'encryption':'in transit + at rest','admin_access':'separate privileged role'}, 'operations':{'team_on_call':'24x7 rotation','skills':['SQL','Python','Linux','containers'],'managed_cloud_required':False,'paid_lab_required':False}, 'cost':{'budget_model':'record compute + storage + replicas + backup + egress + operator time','currency':'USD','hard_ceiling_defined':False}}print('AtlasMart Lesson 5 — decision notebook validation')print(json.dumps(notebook, indent=2))required=[ ('checkout p95', notebook['workloads']['checkout'].get('p95_ms_target')), ('catalog peak reads', notebook['workloads']['catalog'].get('reads_per_s_peak')), ('order RPO', notebook['reliability'].get('order_rpo')), ('order RTO', notebook['reliability'].get('order_rto')), ('key skew risk', notebook['data'].get('key_skew_risk')), ('tenant isolation', notebook['security'].get('tenant_isolation')),]missing=[name for name,value in required if value in (None,'')]print('\nValidation:', 'PASS' if not missing else 'FAIL '+str(missing))# Family-level suitability questions — not a product recommendation or benchmark.questions={ 'relational':['Do multi-record invariants/transactions dominate?','Are joins/ad-hoc queries central?'], 'document':['Can one aggregate own most reads/writes?','Do attributes vary while validation remains explicit?'], 'key_value':['Is access primarily by key?','Is state reconstructable or durability explicitly configured?'], 'wide_column':['Can queries be designed around bounded partition keys?','Is high write throughput/time ordering central?'], 'graph':['Are multi-hop relationships the primary query primitive?'], 'search_vector':['Is this a derived retrieval view with explicit freshness/rebuild semantics?']}print('\nQuestions to answer before naming a product:')for family,qs in questions.items(): print('-',family+':', ' | '.join(qs))print('\nStop condition: do not select a technology until hard invariants, access patterns, latency percentiles, failure behavior, RPO/RTO, security, growth/skew, skills, and cost boundaries are written down.')
Verified output
AtlasMart Lesson 5 — decision notebook validation{ "domain": "AtlasMart", "workloads": { "checkout": { "writes_per_s_peak": 450, "read_write_ratio": "2:1", "p95_ms_target": 180, "invariants": [ "one payment capture per idempotency key", "inventory must not go below reserved safety floor" ] }, "catalog": { "reads_per_s_peak": 12000, "writes_per_s_peak": 35, "p95_ms_target": 80, "shape": "heterogeneous product attributes" }, "search": { "reads_per_s_peak": 9000, "freshness_slo_seconds": 30, "source_of_truth": false }, "sessions": { "reads_per_s_peak": 18000, "writes_per_s_peak": 10000, "ttl_minutes": 45, "reconstructable": true } }, "data": { "orders_year_1_gb": 700, "catalog_gb": 90, "search_index_multiplier_estimate": 1.6, "annual_growth_percent": 55, "key_skew_risk": "high for celebrity products" }, "reliability": { "checkout_availability_target": "99.95%", "catalog_availability_target": "99.9%", "order_rpo": "<= 1 minute", "order_rto": "<= 30 minutes", "search_rpo": "rebuildable from source + event log", "regional_failure_test": "quarterly" }, "security": { "tenant_isolation": "authorization + tenant-scoped keys; key alone is not auth", "pii": [ "customer name", "email", "shipping address" ], "encryption": "in transit + at rest", "admin_access": "separate privileged role" }, "operations": { "team_on_call": "24x7 rotation", "skills": [ "SQL", "Python", "Linux", "containers" ], "managed_cloud_required": false, "paid_lab_required": false }, "cost": { "budget_model": "record compute + storage + replicas + backup + egress + operator time", "currency": "USD", "hard_ceiling_defined": false }}Validation: PASSQuestions to answer before naming a product:- relational: Do multi-record invariants/transactions dominate? | Are joins/ad-hoc queries central?- document: Can one aggregate own most reads/writes? | Do attributes vary while validation remains explicit?- key_value: Is access primarily by key? | Is state reconstructable or durability explicitly configured?- wide_column: Can queries be designed around bounded partition keys? | Is high write throughput/time ordering central?- graph: Are multi-hop relationships the primary query primitive?- search_vector: Is this a derived retrieval view with explicit freshness/rebuild semantics?Stop condition: do not select a technology until hard invariants, access patterns, latency percentiles, failure behavior, RPO/RTO, security, growth/skew, skills, and cost boundaries are written down.
The final family-level questions intentionally do not produce a winner. They tell the team what additional evidence a family would require. A graph database may be attractive only if multi-hop relationships are the primary query primitive; a key-value store only if key access dominates and durability/reconstruction are explicit; a search/vector system only if it is treated as a maintained retrieval view with a freshness and rebuild contract.
4. Deliberately wrong notebook: adjectives without thresholds
A bad decision record contains statements such as “needs massive scale,” “must be highly available,” “schema changes often,” and “latency sensitive.” None defines a test. Under that document, every vendor can claim a fit and no later incident can show that the original assumption was wrong.
The repair is to write falsifiable thresholds and scopes. Which path peaks at 12,000 reads/s? Which one peaks at 450 writes/s? Does “available” mean reads continue during one zone loss, or writes too? Can search lag 30 seconds? Is one minute of order data loss acceptable after regional disaster? Which tenant keys are expected to dominate traffic? Which data must be deleted from backups after retention expiry? The notebook must expose tradeoffs, not hide them.
5. Add ownership and change control to every assumption
Requirements age. Traffic grows, laws change, teams gain or lose expertise, products change licenses, and managed services alter quotas. Therefore each important notebook entry should eventually include: owner, evidence date, confidence, measurement method, and review trigger. A technology decision that was correct at 100 GB and one region can become wrong at 20 TB and five regions.
Version discipline also belongs here. This course is vendor-neutral, but any later product experiment must record exact database version, edition/license, driver, topology, replication factor, consistency level, container/image, OS, and optional modules. The lab baseline itself records current upstream Python/SQLite release status separately from the older versions available in the generation environment, preventing “worked on my machine” from being mistaken for a universal product guarantee.
6. Chapter 01 decision record
AtlasMart does not choose a NoSQL product at the end of Chapter 01. It chooses a disciplined process. Checkout/payment/inventory remain candidates for strong transactional ownership. Catalog documents are a candidate for aggregate-oriented flexible modeling. Sessions are a candidate for key-oriented ephemeral state. Search is a derived retrieval view. Fraud relationships may justify a graph projection. Telemetry may justify a time/partition-oriented model. Each candidate can be rejected later if mechanism-level evidence does not meet the notebook.
| Candidate path | Current hypothesis | What can falsify it later |
|---|---|---|
| Checkout/payment/inventory | Strong transactional source of record | If partitioned/global requirements can preserve invariants with lower operational cost in another model |
| Catalog | Aggregate/document-friendly | If cross-product joins or rigid shared constraints dominate more than flexible attributes/local reads |
| Sessions/idempotency/rate limits | Key-oriented | If multi-key transactional semantics or durable relational reporting become dominant |
| Search/vector | Derived view | If source database native retrieval meets SLOs more cheaply and with less synchronization |
| Fraud relationships | Graph candidate | If traversals are shallow/rare enough that relational or search projections are simpler |
| Telemetry/activity | Wide-column/time-oriented candidate | If volume/query ranges fit a simpler relational/time-series path without partition pain |
7. Production judgment and bridge to Chapter 02
A database selection should survive three questions: what failure are we willing to expose to users, what coordination are we willing to pay for, and what operational system can our team actually recover? The notebook makes those tradeoffs explicit before product demos bias the discussion.
Chapter 02 now introduces the distributed-systems substrate underneath these choices: processes, nodes, clusters, partitions, replicas, coordinators, client routing, partial failure, timeouts, retries, failure domains, and SLOs. The AtlasMart notebook will be reused there to decide which failures matter for each path.
Verification checklist
- The notebook contains numeric or bounded targets instead of only adjectives.
- Search is explicitly non-authoritative and rebuildable.
- Order RPO and RTO are separate fields.
- Tenant isolation includes authorization; a tenant key alone is not treated as access control.
- The family questions do not automatically select or rank a vendor.
- Costs include replicas, backup, egress, and operator time rather than only storage price.
Check your understanding
- What makes a requirement falsifiable?
- How are SLO, RPO, and RTO different?
- Why must key distribution be recorded in addition to total request rate?
- Why is operational skill part of database architecture rather than a staffing footnote?
- What is the correct outcome of Chapter 01: a product choice or a testable decision model?
Review the answers
It has a defined scope, measurement method, and threshold or invariant that evidence can satisfy or violate.
An SLO targets normal service behavior over time; RPO bounds tolerable lost data after recovery; RTO bounds the time to restore required service capability.
Averages can hide celebrity/hot keys. One key or partition can overload while total cluster throughput remains below capacity.
Complex replication, repair, backup, and incident procedures are part of the system. A design the on-call team cannot diagnose or restore does not meet reliability requirements.
A testable decision model. Product choices come later, after mechanisms and workload evidence narrow the options.
Authoritative references
- Google SRE: Service Level Objectives — Official SRE guidance for SLIs, SLOs, percentiles, and user-centered reliability objectives.
- NIST SP 800-34 Rev. 1 — Contingency-planning guidance useful for grounding recovery objectives, testing, and recovery planning.
- Dynamo: Amazon’s Highly Available Key-value Store — Primary example of a database architecture tied to explicit application availability and consistency tradeoffs.
- Bigtable: A Distributed Storage System for Structured Data — Primary example of requirements-driven distributed storage design across varied workloads.
- Python downloads — Official release source used to re-check current Python stable versions at generation time.
- SQLite release history — Official source used to re-check SQLite release status at generation time.