Chapter 01 · Why NoSQL Exists: Workloads, Scale, Flexibility, and Polyglot Persistence
Operational vs Analytical Workloads and Why One Database Rarely Optimizes Every Access Pattern
Compare OLTP, analytical scans, search, graph traversal, and other access patterns by observable candidate work so one database is not assumed to optimize every physical path.
Learning outcomes
AtlasMart has four teams arguing about a shared “database platform.” Checkout needs point reads and short transactions. Finance wants large revenue scans. Search needs token and vector retrieval. Fraud investigators traverse relationships among accounts, cards, addresses, and devices. One product may support all four features, but supporting an operation is not the same as optimizing its physical access path.
Distinguish online transaction processing (OLTP) request serving from analytical scans, search, graph traversal, telemetry, and derived views.
Reason about work in terms of rows/keys/edges/postings examined rather than product marketing categories.
Explain how indexes and materialized access structures trade write cost and freshness for read efficiency.
Recognize when one engine remains operationally sensible even if another specialized engine could make one query faster.
Build an access-pattern inventory that later chapters can map to document, key-value, wide-column, graph, and search models.
1. Workload shape is more precise than “OLTP versus big data”
Online transaction processing (OLTP) describes interactive operational work dominated by relatively small requests—fetch one order, reserve inventory, update a payment state—where latency and concurrency matter. Analytical work often scans or aggregates many records to compute a result. Search turns text into token/posting structures. Graph workloads treat relationships as first-class traversal edges. Telemetry often appends time-ordered events and later queries bounded time/key ranges.
These shapes conflict physically. A row-oriented index can make one-key lookup cheap; a columnar layout can reduce I/O for scans over a few attributes; an inverted index precomputes token-to-document mappings; adjacency structures make graph neighborhoods local. Maintaining all of them means extra storage, extra writes, more failure modes, and freshness boundaries.
| Access pattern | Useful physical idea | Cost moved elsewhere |
|---|---|---|
| Get order by ID | Primary/key index | Index memory/storage and update work |
| Aggregate revenue over millions of orders | Sequential/column-oriented scan, partitions, pre-aggregation | Batch compute, materialized-view freshness, duplicate storage |
| Search product text | Inverted index | Tokenization/analyzer decisions and index update lag |
| Traverse related devices/accounts | Adjacency/graph index | Edge maintenance, supernodes, partitioning complexity |
| Read recent telemetry by device/time | Partition + clustering/time ordering | Partition-size bounds, compaction/retention work |
2. “One database can do it” is not a complete cost model
A mature relational database can support JSON, full-text search, graph-like recursive queries, materialized views, and analytical extensions. A document database can add secondary indexes and aggregation. A search engine can aggregate. The architectural question is not feature availability; it is whether the dominant path has predictable cost under the data volume, concurrency, failure, and operational constraints.
Keeping one database can be the right decision because every additional store adds replication, backup, security, incident response, schema evolution, and synchronization work. Conversely, forcing a high-volume search or graph workload through an unsuitable layout can impose repeated scans, large joins, or unbounded fan-out. Polyglot persistence is justified only when the saved request cost or isolation benefit exceeds the new synchronization/operations cost.
3. Observable work: count candidates before timing anything
The lab intentionally avoids wall-clock timings. Microbenchmarks are easy to distort with cache state, interpreter warmup, data locality, storage, and coordinated omission. Instead it counts conceptual work under fixed structures: one dictionary probe versus 50,000 scanned orders; 10,000 product-name candidates versus 100 postings for a token; 20,000 edges versus a 20-edge adjacency list.
These counts do not predict production latency. They make one causal mechanism visible: an access structure lets the query examine a smaller candidate set, but that structure had to be built and must be kept current. Later performance lessons add real distributions, client concurrency, storage, replication, and failure conditions.
4. Run the workload-shape simulator
from collections import defaultdictprint('AtlasMart Lesson 3 — compare access-path work, not product labels')products=[{'id':i,'name':f'product {i} blue desk lamp' if i%100==0 else f'product {i}'} for i in range(10000)]orders=[{'id':i,'customer':i%1000,'total':(i%200)*100} for i in range(50000)]edges=[(i%1000, (i*17 + i//1000 + 11)%1000) for i in range(20000)]# Point lookup with an index-like dictionary vs scan.order_index={o['id']:o for o in orders}point_steps=1scan_steps=0found=Nonefor o in orders: scan_steps+=1 if o['id']==49999: found=o; breakprint('\nOrder point lookup conceptual work: index steps=',point_steps,'scan steps=',scan_steps)# Analytical aggregation necessarily touches the target data set in this simple model.agg_steps=0; revenue=0for o in orders: agg_steps+=1; revenue+=o['total']print('Full revenue aggregation rows examined=',agg_steps,'result=',revenue)# Search: precompute inverted view instead of scanning all product names per query.inverted=defaultdict(set)for p in products: for token in p['name'].split(): inverted[token].add(p['id'])scan_search_steps=len(products)indexed_candidates=sorted(inverted['lamp'])print('Search for lamp: scan candidates=',scan_search_steps,'inverted-postings=',len(indexed_candidates))# Graph: adjacency makes neighborhood traversal proportional to local degree instead of all edges.adj=defaultdict(list)for a,b in edges: adj[a].append(b)node=42print('Graph neighbors for node 42: edge-scan work=',len(edges),'adjacency work=',len(adj[node]),'neighbors=',adj[node][:8])print('\nDecision: OLTP point access, analytical scans, text search, and graph traversal reward different physical access paths. One engine may implement several, but each path still has maintenance, memory, freshness, and operational costs.')
Verified deterministic output
AtlasMart Lesson 3 — compare access-path work, not product labelsOrder point lookup conceptual work: index steps= 1 scan steps= 50000Full revenue aggregation rows examined= 50000 result= 497500000Search for lamp: scan candidates= 10000 inverted-postings= 100Graph neighbors for node 42: edge-scan work= 20000 adjacency work= 20 neighbors= [725, 726, 727, 728, 729, 730, 731, 732]Decision: OLTP point access, analytical scans, text search, and graph traversal reward different physical access paths. One engine may implement several, but each path still has maintenance, memory, freshness, and operational costs.
The order index is an in-memory dictionary, not a database B-tree; the inverted index is a simple token map, not Lucene; the adjacency list is a Python structure, not a graph engine. That is deliberate. The lab isolates candidate-set size so the learner can see why products build specialized storage/index structures before learning any vendor syntax.
5. Deliberately wrong approach: optimize the rare query and punish the common path
Suppose AtlasMart denormalizes every order into a search-oriented representation because finance occasionally asks a broad query. Checkout now pays the write/index maintenance cost on every transaction, and the representation may weaken relational invariants. Or suppose the team keeps a fully normalized row model as the only representation and implements product search by scanning descriptions on every request. Both designs optimize the wrong path.
The repair is to rank access patterns by frequency, latency objective, correctness requirement, freshness tolerance, and data volume. A rare report can run asynchronously or use a derived analytical view. A product search index can be derived from the catalog system of record. Checkout can retain stronger transactional ownership. The architecture becomes a set of explicit serving paths rather than one overloaded schema.
6. Derived views change correctness questions
Once the same fact appears in more than one store or index, the copies are no longer automatically identical. A search result can lag the catalog; a cache can serve an expired value; a read model can miss an event; an analytical snapshot can represent yesterday. Those are acceptable only when the business requirement defines a freshness window and a rebuild/reconciliation strategy.
This is why “search is eventually consistent” is still too vague. Which field? What maximum age? What happens after the projector is down for 20 minutes? Can users purchase a product that search still displays after deletion? Is the search index authoritative for authorization? The safe answer to the last question is generally no: a derived retrieval view should not silently become the security boundary.
7. Production judgment and bridge to polyglot ownership
Keep one system when its access paths meet the SLOs with acceptable operational complexity. Add a specialized store only when you can name the owner of the source data, the exact derived purpose, the synchronization path, acceptable lag, replay/rebuild procedure, failure behavior, security model, and decommission/rollback plan.
Lesson 4 makes that cost concrete. AtlasMart will intentionally create a broken dual write between a relational system of record and a derived view, then repair it with an atomic outbox and replayable projector. That turns “polyglot persistence” from a collection of database logos into an ownership and failure-recovery design.
Verification checklist
- No timing result is presented as production performance.
- Each candidate count is tied to an explicit access structure.
- Index/search/graph structures are described as maintained views with write/storage/freshness costs.
- The lesson includes the operational argument for staying with one database when specialization is not worth it.
- You can list AtlasMart’s major access patterns before naming any database product.
Check your understanding
- Why can a database support a query yet still be a poor serving choice for that query?
- What cost does an inverted index move from read time to write/index-maintenance time?
- Why are operation counts useful before wall-clock benchmarks?
- When is one general-purpose database operationally superior to several specialized stores?
- What new correctness property appears when a search index becomes a derived view?
Review the answers
Feature support says an operation is possible; it does not guarantee bounded fan-out, candidate count, tail latency, or acceptable resource use at the target scale.
Tokenization and postings are precomputed and must be stored/updated, so reads can examine a small candidate set while writes pay maintenance and the view may lag.
They expose algorithmic/access-path work without conflating the result with cache, hardware, interpreter, network, or storage effects.
When its paths meet SLOs and the saved specialization benefit is smaller than the added synchronization, security, backup, monitoring, and on-call complexity.
Freshness/convergence becomes explicit: the index can disagree with its source temporarily or after missed updates, so lag, replay, rebuild, and authorization boundaries must be designed.
Authoritative references
- Bigtable: A Distributed Storage System for Structured Data — Primary example of storage designed for diverse structured-data access patterns at scale.
- Dynamo: Amazon’s Highly Available Key-value Store — Primary example of a system intentionally optimized for a narrower key-value workload and availability target.
- Google SRE: Service Level Objectives — Official framework for turning vague performance expectations into measurable objectives.
- SQLite query planner overview — Official relational example showing how indexes change candidate access and plan choices.