Design storefront search around relevance, facets, autocomplete, and authorization—not telemetry defaults.
Application/Product Search: Catalog Modeling, Synonyms, Relevance Signals, Autocomplete, Facets, and Personalization Boundaries
Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.
Learning outcomes
Model product/catalog documents from storefront query and facet requirements rather than reusing telemetry schemas.
Separate synonym expansion, autocomplete, faceting, business relevance signals, and personalization because they fail and evolve differently.
Measure relevance and tail latency with representative queries instead of inferring search quality from cluster health.
Identify privacy, tenant, and authorization boundaries that personalization must not cross.
Explain why product search and log/security/observability workloads deserve different mappings, lifecycle, and operational ownership.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart problem: one engine, four incompatible definitions of “good”
AtlasMart can store product documents, application logs, security events, and traces in Elasticsearch or OpenSearch. That does not make them one workload. A shopper expects a relevant result in tens or hundreds of milliseconds, stable facets, forgiving synonyms, and predictable autocomplete. A security investigator may accept a slower interactive query if it spans months of normalized events. Log ingestion values sustained write throughput and retention economics. Observability requires cross-signal correlation and freshness. The architecture must begin with these user-visible contracts, not with “we already have a search cluster.”
For product search, the search contract is the combination of document model, analyzers, exact-filter fields, query template, ranking signals, facet definitions, and relevance judgments. Changing a synonym rule or autocomplete analyzer can change user-visible ranking just as materially as changing application code.
2. Catalog model: searchable text is not the same thing as a facet
AtlasMart keeps stable identity and filtering fields as
keyword, full-text names/descriptions as
text, numeric price and availability as typed
fields, and variant structures only when the query requires
per-variant relationships. A facet such as brand or category
must aggregate on an exact-value field; running a terms
aggregation over analyzed product prose is both semantically
wrong and operationally expensive.
| Requirement | Mapping/query mechanism | Why it is separate |
|---|---|---|
| Search “waterproof hiking boot” |
text with a tested analyzer; multi-field
exact form where needed
|
Analysis and BM25 decide lexical matching. |
| Filter brand/category/tenant | keyword in filter context |
Exact values and authorization constraints must not be fuzzed. |
| Price slider |
numeric scaled_float or integer cents
|
Range semantics and sorting are numeric, not text. |
| Autocomplete | bounded prefix/edge-ngram strategy on a dedicated field | Index-time expansion and query-time cost need their own budget. |
| Facets | terms/range aggregations over controlled fields | Facet correctness depends on modeling and cardinality. |
| Popularity/freshness | declared business signal with bounded influence | Business relevance must not drown textual relevance. |
3. Portable AtlasMart product-search fixture
All Chapter 28 labs keep the existing local endpoints and
security assumptions: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD and the copied CA file
atlasmart-es-http-ca; OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's
demo certificate trust bypass (-k) is acceptable
only for this disposable local lab, never production. The shared
Docker network remains atlasmart-search. Lab
indices use one primary and zero replicas so a single-node
workstation can complete the exercises; production redundancy
decisions are deliberately separate.
PUT atlasmart-product-search-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0,
"analysis": {
"filter": {
"atlasmart_synonyms": {
"type": "synonym_graph",
"synonyms": [
"sneakers, trainers",
"waterproof, water resistant"
]
}
},
"analyzer": {
"atlasmart_search": {
"tokenizer": "standard",
"filter": ["lowercase", "atlasmart_synonyms"]
}
}
}
},
"mappings": {
"dynamic": "strict",
"properties": {
"sku": {"type":"keyword"},
"tenant_id": {"type":"keyword"},
"name": {"type":"text", "fields":{"raw":{"type":"keyword"}}},
"name_suggest":{"type":"text"},
"description":{"type":"text"},
"brand": {"type":"keyword"},
"category": {"type":"keyword"},
"price": {"type":"scaled_float", "scaling_factor":100},
"available": {"type":"boolean"},
"popularity": {"type":"float"},
"updated_at": {"type":"date"}
}
}
}
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1001"}}
{"sku":"P-1001","tenant_id":"tenant-a","name":"Waterproof Trail Boot","name_suggest":"waterproof trail boot","description":"Lightweight hiking boot for wet trails","brand":"NorthPeak","category":"boots","price":129.00,"available":true,"popularity":0.83,"updated_at":"2026-09-01T00:00:00Z"}
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1002"}}
{"sku":"P-1002","tenant_id":"tenant-a","name":"City Trainer","name_suggest":"city trainer","description":"Everyday sneaker with cushioned sole","brand":"MetroRun","category":"shoes","price":79.00,"available":true,"popularity":0.72,"updated_at":"2026-09-02T00:00:00Z"}
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1003"}}
{"sku":"P-1003","tenant_id":"tenant-a","name":"Alpine Shell Jacket","name_suggest":"alpine shell jacket","description":"Water resistant shell for hiking","brand":"NorthPeak","category":"jackets","price":159.00,"available":true,"popularity":0.55,"updated_at":"2026-09-03T00:00:00Z"}
The inline synonym list is intentionally tiny and portable. Production synonym lifecycle differs between products and distributions; do not mistake this lab choice for a recommendation to hard-code large synonym sets into mappings.
4. Synonyms: linguistic equivalence is a versioned relevance decision
Synonym expansion changes the token stream and therefore recall and ranking. “Sneakers” and “trainers” may be equivalent for one locale and not another. Search-time synonyms are attractive because they can change without reindexing when the platform and analyzer design support safe reload, but the exact management APIs and reload behavior differ between Elastic and OpenSearch. The portable rule is simpler: version the synonym source, test the analyzer output, keep a lexical baseline, and treat a synonym change as a relevance release.
A dangerous shortcut is to add every merchandising association as a synonym. “Boots → winter” is not linguistic equivalence; it is a ranking or recommendation rule and should be measured as such.
5. Autocomplete: prefix convenience can become an index-size and CPU tax
Autocomplete is a separate workload. Edge n-grams move work to indexing and increase terms; prefix queries move work to search; specialized suggestion fields have their own semantics. Choose from a measured interaction contract: minimum prefix length, maximum suggestions, typo behavior, language, latency budget, and whether suggestions must respect tenant/availability filters.
POST atlasmart-product-search-v1/_search
{
"size": 5,
"_source": ["sku","name","brand","category"],
"query": {
"bool": {
"filter": [
{"term": {"tenant_id": "tenant-a"}},
{"term": {"available": true}}
],
"must": [
{"match_phrase_prefix": {"name_suggest": {"query": "water", "max_expansions": 25}}}
]
}
}
}
Expected invariant: only authorized, available tenant-a products are candidates. The exact hit order is not a cross-platform contract; compare relevance using fixed judgments instead of copying raw scores between engines.
6. Facets and business signals: keep filtering semantics visible
POST atlasmart-product-search-v1/_search
{
"size": 10,
"query": {
"bool": {
"filter": [
{"term":{"tenant_id":"tenant-a"}},
{"term":{"available":true}}
],
"must": [
{"multi_match": {
"query":"waterproof hiking",
"fields":["name^3","description"],
"analyzer":"atlasmart_search"
}}
]
}
},
"aggs": {
"brands":{"terms":{"field":"brand","size":10}},
"categories":{"terms":{"field":"category","size":10}}
}
}
Facet counts answer “how many matching documents are in each bucket under this query/filter context?” They are not global catalog counts. If the UI needs disjunctive faceting, define that explicitly and test it; do not improvise by removing filters at random.
Popularity, freshness, margin, inventory, and personalization can influence ranking, but each should have a bounded role and an evaluation set. Raw BM25 scores and business-feature values are on different scales, so adding them blindly is not a stable scoring contract.
7. Personalization boundary: identity is not just another feature
Personalization introduces sensitive user/tenant context. A product-search service should first apply authorization and catalog eligibility, then personalize within the authorized candidate set. Never retrieve cross-tenant documents and “filter them out later” in application code. Keep the source of personalization features, retention policy, consent basis, and failure fallback explicit. If the personalization service is unavailable, the product should degrade to a tested non-personalized ranking rather than fail open or expose another user’s features.
8. Mini lab: prove relevance and workload shape
- Run the lexical/facet query above with five representative AtlasMart product queries.
- Record top-5 IDs, facet counts, request latency, and the exact query/mapping version.
- Add one synonym that should increase recall. Re-run the same judgments; verify the intended query improves without unrelated regressions.
- Run 100 autocomplete requests with a fixed prefix list and record p50/p95/p99 from the client. Do not publish synthetic numbers as production capacity.
- In parallel, run a small log-ingest loop from Lesson 2 and observe whether product p99 changes. This is the first noisy-neighbor signal, not yet proof of causality.
Acceptance: the team can explain which fields are searchable, filterable, aggregatable, and personalized; which metrics demonstrate storefront health; and why this index should not inherit telemetry retention or security policy.
Check your understanding
- Why should brand/category facets use exact-value fields?
- Why is a synonym change a release?
- Does a green cluster prove product search is healthy?
- Where should tenant authorization be applied?
- What is the bridge to Lesson 2?
Review the answers
1. Because facets group exact values; analyzed prose changes tokens and therefore bucket semantics.
2. It changes token expansion, recall, and ranking, so it needs versioning and relevance regression tests.
3. No. User-visible relevance, p95/p99 latency, error rate, facet correctness, and freshness can fail while cluster health remains green.
4. Before or as part of retrieval, not as a best-effort post-filter after unauthorized candidates have been retrieved.
5. Product search optimizes relevance and interactive latency; log analytics instead prioritizes structured event contracts, sustained ingest, cardinality control, and retention.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch 9.5.3 release notes
- Elastic Common Schema 9.5 reference
- ECS getting started and normalization
- ECS log fields
- Elastic Observability fields and object schemas
- Elastic ECS-formatted application logs
- OpenSearch 3.8 version history
- OpenSearch Security Analytics overview
- OpenSearch Security Analytics detectors
- OpenSearch Security Analytics access control
- OpenSearch APM configuration
- OpenSearch Trace Analytics
- OpenTelemetry logs data model
- OpenTelemetry service semantic conventions