Compare partition-local and globally routable AtlasMart secondary indexes by fan-out, maintenance, hot-key risk, freshness, and repair rather than by product naming.
Local vs Global Secondary Indexes in Distributed Databases
Use the same catalog workload to show why local indexes preserve base-partition locality while global indexes create a separately partitioned distributed access path with independent scaling and failure behavior.
Learning outcomes
Use the same catalog workload to show why local indexes preserve base-partition locality while global indexes create a separately partitioned distributed access path with independent scaling and failure behavior.
Define local versus global secondary indexing in terms of partition scope and routing.
Explain why global indexes have independent distribution, write fan-out, consistency, and repair concerns.
Identify hot global index keys and scatter-gather local-index queries.
Choose index scope based on known routing keys, query breadth, freshness, and operational complexity.
Mandatory work uses Python 3.13+ standard library only in one local process. No Elasticsearch, OpenSearch, cloud service, Docker image, paid feature, network manipulation, or destructive failure injection is required. Current optional reference snapshots are Elasticsearch 9.5.2 (released August 20, 2026; default distribution under Elastic License 2.0) and OpenSearch 3.8.0 (released August 4, 2026; Apache License 2.0). Product-specific refresh, index, ANN, clustering, quota, security, and licensing semantics are examples—not universal database guarantees.
1. “Local” and “global” describe index scope
In a sharded database, a local secondary index is scoped to the same base partition or shard that owns the indexed record. If the request already knows that base partition key, the local index can provide another ordering or predicate without leaving that ownership boundary. A global secondary index is partitioned according to its own index key and can answer an alternate query across base partitions. These are conceptual terms; products use different names and semantics. Amazon DynamoDB is a useful labeled example: its Local Secondary Index shares the base partition key, while its Global Secondary Index can use a different partition key and is stored in its own partition space.
2. Routing determines the real query cost
AtlasMart asks for all products in category
footwear without knowing product IDs. If category
indexes exist only inside each product shard, the coordinator
must scatter the query to every relevant shard and merge
results. That may be acceptable for a tiny shard count or an
admin query, but query fan-out grows operationally expensive as
shards, regions, and replicas increase. A global index turns
category into a first-class routable key, so the read no longer
has to discover category matches on every base shard. The cost
shifts to the write path and to operating the index's own
partitioning.
3. AtlasMart lab: make the mechanism observable
Save the following as
lesson3_local_global_indexes.py and run it with
python lesson3_local_global_indexes.py. It uses
only deterministic in-memory data and mutates no external
service.
from collections import defaultdict
import hashlib
N_SHARDS = 3
base = [{} for _ in range(N_SHARDS)]
local_category = [defaultdict(set) for _ in range(N_SHARDS)]
global_category = defaultdict(set)
def shard_for(product_id):
h = int(hashlib.sha256(product_id.encode()).hexdigest(), 16)
return h % N_SHARDS
def put(product):
sid = shard_for(product["id"])
old = base[sid].get(product["id"])
if old:
local_category[sid][old["category"]].discard(product["id"])
global_category[old["category"]].discard(product["id"])
base[sid][product["id"]] = dict(product)
local_category[sid][product["category"]].add(product["id"])
global_category[product["category"]].add(product["id"])
return sid
for p in [
{"id":"p1","category":"footwear"}, {"id":"p2","category":"footwear"},
{"id":"p3","category":"audio"}, {"id":"p4","category":"outdoor"},
{"id":"p5","category":"audio"}, {"id":"p6","category":"footwear"},
]: put(p)
# If only category is known, local indexes require scatter to all base shards.
local_results = set(); touched = 0
for sid in range(N_SHARDS):
touched += 1
local_results |= local_category[sid]["footwear"]
print("local category results:", sorted(local_results), "shards touched:", touched)
print("global category results:", sorted(global_category["footwear"]), "index partitions touched: 1")
# If base partition is already known, local index stays local.
sid = shard_for("p1")
print("known base shard for p1:", sid)
print("local footwear entries on that shard:", sorted(local_category[sid]["footwear"]))
# Global low-cardinality key becomes a hot index partition.
for i in range(80): put({"id":f"sale-{i}","category":"sale"})
for i in range(10): put({"id":f"new-{i}","category":"new"})
print("global index key sizes:", {k:len(v) for k,v in sorted(global_category.items())})
# Deliberately broken global maintenance: base/local update succeeds; global stays stale.
sid = shard_for("p3")
old = dict(base[sid]["p3"])
local_category[sid][old["category"]].discard("p3")
base[sid]["p3"] = {"id":"p3","category":"clearance"}
local_category[sid]["clearance"].add("p3")
print("base/local says clearance:", "p3" in local_category[sid]["clearance"])
print("stale global still says audio:", "p3" in global_category["audio"])
# Reconcile global view from every authoritative base shard.
rebuilt = defaultdict(set)
for shard in base:
for p in shard.values(): rebuilt[p["category"]].add(p["id"])
global_category = rebuilt
print("repaired global clearance:", "p3" in global_category["clearance"])
print("repaired global audio excludes p3:", "p3" not in global_category["audio"])
Expected evidence: a category-only query through local indexes touches all three base shards, while the modeled global category index gives one routable access path; when the base shard for p1 is already known, the local index remains partition-local; the low-cardinality “sale” key becomes much larger than other global keys; and the deliberately stale global index is repaired by rebuilding from authoritative base partitions.
4. A global index is another distributed system
A global index needs partition ownership, routing metadata,
replica/durability policy, hot-key handling, backlog/lag
metrics, repair/rebuild procedures, and capacity. If the base
write is acknowledged before the global index catches up, the
index is eventually consistent for some interval. If index
maintenance is synchronous, write availability or latency may
depend on index health. A low-cardinality or skewed index
key—such as sale—can concentrate updates even if
base product IDs are perfectly distributed. This is why “add a
GSI” or equivalent is an architecture decision, not just DDL.
5. Product semantics must not be generalized
DynamoDB's current documentation says Global Secondary Index queries are eventually consistent, while Local Secondary Index queries can use eventual or strong consistency. That is a DynamoDB contract, not a universal law of “global” and “local” indexing. Other systems may maintain global indexes transactionally, asynchronously, through log-based projections, or not support them at all. Always document the concrete product/version: write acknowledgement, read consistency, projection attributes, online build behavior, failure/recovery semantics, and any quotas or billing model.
6. Repair and rebuild are part of the design
The lab intentionally updates p3's authoritative base shard and its local index while leaving the global view stale. A query through that global view returns a plausible false result until reconciliation. Production repair can use change-log replay, versioned index records, checksums/cardinality comparisons, sampled source hydration, or full rebuilds. Rebuild capacity must account for scanning all base partitions and for concurrent changes. If the global index supports customer-facing discovery, define a freshness service-level objective rather than merely saying it is “eventual.”
7. Production judgment
Prefer local indexes when the dominant query already supplies the base partition key and the extra predicate/order should remain within that ownership boundary. Consider a global index when the alternate key must route across all base partitions frequently enough to justify independent index maintenance. Measure scatter count, write fan-out, index key skew, stale-result windows, p95/p99 latency, backlog, rebuild throughput, and tenant isolation. If the desired alternate query is rich text, faceting, ranking, or vector similarity, a dedicated search structure may be more appropriate than stretching a database secondary index beyond its design envelope.
Wrong approach: call a global index “one lookup” and ignore its partitioning
The architecture review says a global index makes category lookup O(1), so no one models index-key skew or write fan-out. A promotion sends most products through the same “sale” key and that index partition becomes the bottleneck. The repair is to inspect index-key distribution, capacity/failure domains, and product-specific partitioning rules; where appropriate, bucket/salt the global key and accept the resulting read fan-out, or choose a search/materialized-view architecture better suited to the query.
Verification, cleanup, and production checklist
Verification is the deterministic program output plus the
conceptual checks below. Cleanup is deleting the local
lesson3_local_global_indexes.py file; the lab
creates no sockets, databases, containers, credentials, indexes,
or cloud resources. In production, additionally record
authoritative-versus-derived ownership, source/index versions,
refresh or projection lag, p95/p99 read/write latency, index
size and write amplification, shard/index-key skew, rebuild
throughput, ANN recall where applicable, tenant/authorization
tests, backup/rebuild evidence, current security advisories, and
edition/license constraints before relying on a product-specific
feature.
Check your understanding
- When is a local secondary index naturally efficient?
- Why can a category-only query scatter with local indexes?
- Why is a global index a distributed system of its own?
- Does “global index” imply eventual consistency in every database?
- What new hotspot can a global index create?
Review the answers
1. When the query already knows the base partition key, so the alternate predicate/order stays within that partition.
2. Category is not the base routing key, so every base shard may need to be asked for its local matches.
3. It has independent partitioning, routing, maintenance, capacity, lag, repair, and failure behavior.
4. No. That is product-specific; DynamoDB GSI eventual consistency is an example, not a universal rule.
5. A popular or low-cardinality alternate key can concentrate index writes even when base keys are balanced.
References
Foundational statements use primary research or standards where appropriate. Version-sensitive implementation examples use current official documentation and remain explicitly scoped to the cited product/version.
- DynamoDB — Secondary indexes — Current explicit comparison of local and global secondary indexes.
- DynamoDB — Secondary-index guidelines — Current index scope, projection, and efficiency guidance.
- PostgreSQL 18 — Indexes — Contrast with non-sharded relational index maintenance concepts.
- Elasticsearch — Refresh visibility — Example of explicitly documented derived-index visibility behavior.