Choose AtlasMart search boundaries from workload evidence, then prove synchronization, freshness, reconciliation, failure isolation, security, and rebuild behavior for any dedicated derived index.

Choose Database-Native Search vs Dedicated Search/Vector Infrastructure

Compare database-native secondary/search features with dedicated search/vector infrastructure by query richness, freshness, scale, ranking quality, vector needs, failure isolation, operational duplication, and synchronization cost.

Intermediate–Advanced105–135 minutesArchitecture + reconciliation labPython 3.13+ · standard libraryElasticsearch 9.5.2 / OpenSearch 3.8 optional referencesLast reviewed: August 2026

Learning outcomes

Compare database-native secondary/search features with dedicated search/vector infrastructure by query richness, freshness, scale, ranking quality, vector needs, failure isolation, operational duplication, and synchronization cost.

01

Choose database-native versus dedicated search/vector infrastructure from explicit workload evidence.

02

Define authority, projection freshness, synchronization, replay, reconciliation, and rebuild boundaries.

03

Account for ranking/vector quality, failure isolation, security, cost, and operational duplication.

04

Design migration and rollback so search infrastructure can be introduced or replaced without losing business truth.

Implementation snapshot

Mandatory work uses Python 3.13+ standard library only in one local process. No Elasticsearch, OpenSearch, cloud service, Docker image, paid feature, network manipulation, or destructive failure injection is required. Current optional reference snapshots are Elasticsearch 9.5.2 (released August 20, 2026; default distribution under Elastic License 2.0) and OpenSearch 3.8.0 (released August 4, 2026; Apache License 2.0). Product-specific refresh, index, ANN, clustering, quota, security, and licensing semantics are examples—not universal database guarantees.

1. The boundary is an architecture choice, not a product popularity contest

AtlasMart can satisfy some alternate queries with database-native indexes or built-in text/vector features. A dedicated search service becomes attractive when the workload needs rich analyzers, relevance tuning, faceting, many alternate fields, high query concurrency, large vector ANN structures, independent scaling, or specialized search operations. The decision should start from access patterns and invariants. If the only need is exact lookup by email or status with immediate transactional visibility, a native index may be simpler and safer. If customer discovery needs typo tolerance, lexical relevance, facets, semantic retrieval, ranking experiments, and independent search capacity, dedicated infrastructure may justify the synchronization burden.

2. Compare the systems across the whole lifecycle

Evaluate query richness, freshness, data scale, ranking quality, vector workload, write amplification, failure isolation, security model, backup/rebuild, operator skill, and cost. Native search can reduce duplicated infrastructure and synchronization paths, but may couple heavy search to transactional resources or provide fewer ranking/vector controls. Dedicated search can isolate query load and evolve independently, but every indexed fact becomes duplicated state whose freshness and lifecycle must be governed.

3. AtlasMart lab: make the mechanism observable

Save the following as lesson5_search_boundary.py and run it with python lesson5_search_boundary.py. It uses only deterministic in-memory data and mutates no external service.

python · AtlasMart deterministic simulation
from collections import deque

workloads = {
 "admin_exact_lookup": {"rich_text":0,"vector":0,"freshness":5,"scale":2,"ranking":0,"ops_tolerance":1},
 "storefront_discovery": {"rich_text":5,"vector":2,"freshness":3,"scale":5,"ranking":5,"ops_tolerance":4},
 "semantic_related": {"rich_text":2,"vector":5,"freshness":2,"scale":4,"ranking":4,"ops_tolerance":4},
}

def choose(w):
    # Transparent teaching heuristic, not a vendor benchmark.
    dedicated = 2*w["rich_text"] + 3*w["vector"] + w["scale"] + 2*w["ranking"] + w["ops_tolerance"]
    native = 3*w["freshness"] + 2*(5-w["rich_text"]) + 2*(5-w["vector"])
    return ("dedicated-derived" if dedicated > native else "database-native"), dedicated, native

for name,w in workloads.items():
    print(name, "->", choose(w))

# AtlasMart source of truth and a dedicated derived search document.
source = {"p1":{"version":1,"title":"Trail Runner","price":120,"stock":8}}
search = {"p1":{"version":1,"title":"Trail Runner","price":120,"stock":8}}
outbox = deque()

# Authority commits version 2; projection event is durable but not applied yet.
source["p1"] = {"version":2,"title":"Trail Runner","price":99,"stock":8}
outbox.append(("p1", dict(source["p1"])))
print("source version/price:", source["p1"]["version"], source["p1"]["price"])
print("search stale version/price:", search["p1"]["version"], search["p1"]["price"])
print("freshness gap versions:", source["p1"]["version"] - search["p1"]["version"])

# Replay projection idempotently.
while outbox:
    pid,event = outbox.popleft()
    if event["version"] >= search.get(pid,{}).get("version",0):
        search[pid] = event
print("after replay search version/price:", search["p1"]["version"], search["p1"]["price"])

# Deliberately damage the derived index; source remains authoritative.
del search["p1"]
print("derived search missing p1:", "p1" not in search)
print("source still has p1:", "p1" in source)

# Full deterministic reconciliation/rebuild.
for pid,doc in source.items(): search[pid] = dict(doc)
extra = set(search) - set(source)
missing = set(source) - set(search)
wrong_version = {pid for pid in source if search[pid]["version"] != source[pid]["version"]}
print("rebuild missing/extra/wrong_version:", sorted(missing), sorted(extra), sorted(wrong_version))
Expected evidence

Expected evidence: the transparent decision heuristic chooses database-native access for exact/fresh admin lookup and dedicated derived infrastructure for richer discovery/vector workloads; an authoritative version-2 price change is temporarily stale in the search projection; replay closes the version gap; deleting the derived search document does not delete the authoritative product; and full rebuild/reconciliation returns no missing, extra, or wrong-version records.

4. Make authority and synchronization explicit

A robust polyglot boundary says “the catalog database owns product price and stock; the search index is a derived retrieval view.” A transaction commits the source fact and a durable change stream/outbox/CDC mechanism makes that change replayable. The projector writes idempotently using a monotonic source version or equivalent ordering token. Search responses can carry that version for diagnostics. If the index is stale, a correctness-sensitive checkout hydrates authoritative price/stock from the source rather than trusting an old search document. This preserves the retrieval benefit without silently transferring business authority.

5. Reconciliation is what turns eventual synchronization into an operable system

Retries do not prove every change was applied exactly once or in order. AtlasMart needs a reconciliation process that compares source and derived state using counts, keys, versions, checksums, sampled field comparisons, or domain invariants. Full rebuild should be rehearsed from a known source snapshot plus catch-up stream. Measure source-to-index lag, oldest pending event age, projection error rate, missing/stale document count, refresh/segment/merge pressure, vector build backlog, and query success/latency. A derived index that cannot be rebuilt safely is closer to a hidden system of record than the architecture diagram admits.

6. Security and tenant isolation cross the boundary too

Search copies often contain denormalized customer, product, entitlement, or behavioral fields. Authorization must not rely on the source database protecting data that has already been copied into another system. Enforce tenant/visibility filters inside search requests, minimize indexed sensitive fields, secure ingestion and administrative APIs, encrypt transport/storage as appropriate, and test facet/aggregation leakage. For vector search, embeddings may encode sensitive information or enable inference, so retention and access rules apply to vectors as data—not merely to the original text.

7. Migration, rollback, and cost

Introduce a dedicated index through shadow ingestion, backfill, dual-read comparison, freshness/error dashboards, and gradual traffic cutover. Keep a rollback path to native/database queries while the source remains authoritative. Cost modeling includes search nodes, replicas, hot/warm storage, vector memory, ingest/merge CPU, network/egress, backup snapshots, observability, operator time, and duplicate retained data. Elasticsearch 9.5.2 and OpenSearch 3.8.0 are current examples with different licensing and operational ecosystems; neither is automatically the right answer. The correct boundary is the one whose measured retrieval benefit exceeds synchronization and operating cost while preserving business invariants.

8. Bridge to access-pattern-driven modeling

This chapter treated secondary indexes, inverted search, global index structures, and ANN as alternate representations created to make specific reads cheaper. Chapter 15 generalizes that idea: denormalization and materialized views are precomputed work. The same questions return—who owns the fact, which query is being optimized, how many copies change on a write, what staleness is acceptable, how are projections rebuilt, and what evidence proves the copies have converged? Search is therefore not a separate design universe; it is one specialized case of access-pattern-driven data modeling.

Wrong approach: let the dedicated search index quietly become the source of truth

Failure injection / diagnosis

The storefront starts writing product price edits directly to search because “that is where users read them.” Back-office and checkout still use the catalog database, so two authorities emerge. A search rebuild can now destroy newer edits, and a search outage blocks business updates. The repair is to restore one authoritative write path, capture changes through replayable CDC/outbox events, hydrate correctness-sensitive reads from authority when necessary, reconcile source/index versions, and prove that the search index can be deleted and rebuilt without losing business facts.

Verification, cleanup, and production checklist

Verification is the deterministic program output plus the conceptual checks below. Cleanup is deleting the local lesson5_search_boundary.py file; the lab creates no sockets, databases, containers, credentials, indexes, or cloud resources. In production, additionally record authoritative-versus-derived ownership, source/index versions, refresh or projection lag, p95/p99 read/write latency, index size and write amplification, shard/index-key skew, rebuild throughput, ANN recall where applicable, tenant/authorization tests, backup/rebuild evidence, current security advisories, and edition/license constraints before relying on a product-specific feature.

Check your understanding

  1. When is database-native search/indexing often preferable?
  2. What must be explicit when search is a derived system?
  3. Why hydrate price or stock from the authority at checkout?
  4. What does a successful rebuild prove?
  5. What costs are easy to miss with dedicated search/vector infrastructure?
Review the answers

1. When queries are relatively simple, freshness/transactional coupling matters, and dedicated ranking/vector capabilities do not justify another operational system.

2. The authoritative source, event/change path, ordering/version semantics, freshness SLO, replay, reconciliation, and rebuild procedure.

3. Search may be intentionally stale; correctness-sensitive invariants should use the system that owns those facts.

4. That the derived index can be reconstructed from authoritative data and catch-up changes without becoming a hidden source of truth.

5. Duplicate storage, replicas, vector memory, ingest/merge CPU, network, backups, observability, operator time, and synchronization/reconciliation work.

References

Foundational statements use primary research or standards where appropriate. Version-sensitive implementation examples use current official documentation and remain explicitly scoped to the cited product/version.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.