Design storefront search around relevance, facets, autocomplete, and authorization—not telemetry defaults.

Application/Product Search: Catalog Modeling, Synonyms, Relevance Signals, Autocomplete, Facets, and Personalization Boundaries

Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.

Intermediate → Advanced135–180 minutesProduct-search modeling & relevance lab · Chapter 28 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · ECS 9.5.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Model product/catalog documents from storefront query and facet requirements rather than reusing telemetry schemas.

02

Separate synonym expansion, autocomplete, faceting, business relevance signals, and personalization because they fail and evolve differently.

03

Measure relevance and tail latency with representative queries instead of inferring search quality from cluster health.

04

Identify privacy, tenant, and authorization boundaries that personalization must not cross.

05

Explain why product search and log/security/observability workloads deserve different mappings, lifecycle, and operational ownership.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The current Elastic Common Schema reference is ECS 9.5.0. The mandatory AtlasMart labs use free/local HTTP APIs and deterministic fixtures. The generation environment did not execute live clusters, so numeric latency, throughput, shard growth, and storage values shown as acceptance criteria are measurement instructions—not fabricated captured results.

1. AtlasMart problem: one engine, four incompatible definitions of “good”

AtlasMart can store product documents, application logs, security events, and traces in Elasticsearch or OpenSearch. That does not make them one workload. A shopper expects a relevant result in tens or hundreds of milliseconds, stable facets, forgiving synonyms, and predictable autocomplete. A security investigator may accept a slower interactive query if it spans months of normalized events. Log ingestion values sustained write throughput and retention economics. Observability requires cross-signal correlation and freshness. The architecture must begin with these user-visible contracts, not with “we already have a search cluster.”

For product search, the search contract is the combination of document model, analyzers, exact-filter fields, query template, ranking signals, facet definitions, and relevance judgments. Changing a synonym rule or autocomplete analyzer can change user-visible ranking just as materially as changing application code.

2. Catalog model: searchable text is not the same thing as a facet

AtlasMart keeps stable identity and filtering fields as keyword, full-text names/descriptions as text, numeric price and availability as typed fields, and variant structures only when the query requires per-variant relationships. A facet such as brand or category must aggregate on an exact-value field; running a terms aggregation over analyzed product prose is both semantically wrong and operationally expensive.

Requirement Mapping/query mechanism Why it is separate
Search “waterproof hiking boot” text with a tested analyzer; multi-field exact form where needed Analysis and BM25 decide lexical matching.
Filter brand/category/tenant keyword in filter context Exact values and authorization constraints must not be fuzzed.
Price slider numeric scaled_float or integer cents Range semantics and sorting are numeric, not text.
Autocomplete bounded prefix/edge-ngram strategy on a dedicated field Index-time expansion and query-time cost need their own budget.
Facets terms/range aggregations over controlled fields Facet correctness depends on modeling and cardinality.
Popularity/freshness declared business signal with bounded influence Business relevance must not drown textual relevance.

3. Portable AtlasMart product-search fixture

All Chapter 28 labs keep the existing local endpoints and security assumptions: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab, never production. The shared Docker network remains atlasmart-search. Lab indices use one primary and zero replicas so a single-node workstation can complete the exercises; production redundancy decisions are deliberately separate.

Create the shared product-search lab index
PUT atlasmart-product-search-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "analysis": {
      "filter": {
        "atlasmart_synonyms": {
          "type": "synonym_graph",
          "synonyms": [
            "sneakers, trainers",
            "waterproof, water resistant"
          ]
        }
      },
      "analyzer": {
        "atlasmart_search": {
          "tokenizer": "standard",
          "filter": ["lowercase", "atlasmart_synonyms"]
        }
      }
    }
  },
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "sku":       {"type":"keyword"},
      "tenant_id": {"type":"keyword"},
      "name":      {"type":"text", "fields":{"raw":{"type":"keyword"}}},
      "name_suggest":{"type":"text"},
      "description":{"type":"text"},
      "brand":     {"type":"keyword"},
      "category":  {"type":"keyword"},
      "price":     {"type":"scaled_float", "scaling_factor":100},
      "available": {"type":"boolean"},
      "popularity": {"type":"float"},
      "updated_at": {"type":"date"}
    }
  }
}
Index a compact product fixture
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1001"}}
{"sku":"P-1001","tenant_id":"tenant-a","name":"Waterproof Trail Boot","name_suggest":"waterproof trail boot","description":"Lightweight hiking boot for wet trails","brand":"NorthPeak","category":"boots","price":129.00,"available":true,"popularity":0.83,"updated_at":"2026-09-01T00:00:00Z"}
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1002"}}
{"sku":"P-1002","tenant_id":"tenant-a","name":"City Trainer","name_suggest":"city trainer","description":"Everyday sneaker with cushioned sole","brand":"MetroRun","category":"shoes","price":79.00,"available":true,"popularity":0.72,"updated_at":"2026-09-02T00:00:00Z"}
{"index":{"_index":"atlasmart-product-search-v1","_id":"P-1003"}}
{"sku":"P-1003","tenant_id":"tenant-a","name":"Alpine Shell Jacket","name_suggest":"alpine shell jacket","description":"Water resistant shell for hiking","brand":"NorthPeak","category":"jackets","price":159.00,"available":true,"popularity":0.55,"updated_at":"2026-09-03T00:00:00Z"}

The inline synonym list is intentionally tiny and portable. Production synonym lifecycle differs between products and distributions; do not mistake this lab choice for a recommendation to hard-code large synonym sets into mappings.

4. Synonyms: linguistic equivalence is a versioned relevance decision

Synonym expansion changes the token stream and therefore recall and ranking. “Sneakers” and “trainers” may be equivalent for one locale and not another. Search-time synonyms are attractive because they can change without reindexing when the platform and analyzer design support safe reload, but the exact management APIs and reload behavior differ between Elastic and OpenSearch. The portable rule is simpler: version the synonym source, test the analyzer output, keep a lexical baseline, and treat a synonym change as a relevance release.

A dangerous shortcut is to add every merchandising association as a synonym. “Boots → winter” is not linguistic equivalence; it is a ranking or recommendation rule and should be measured as such.

5. Autocomplete: prefix convenience can become an index-size and CPU tax

Autocomplete is a separate workload. Edge n-grams move work to indexing and increase terms; prefix queries move work to search; specialized suggestion fields have their own semantics. Choose from a measured interaction contract: minimum prefix length, maximum suggestions, typo behavior, language, latency budget, and whether suggestions must respect tenant/availability filters.

A bounded prefix-style search pattern
POST atlasmart-product-search-v1/_search
{
  "size": 5,
  "_source": ["sku","name","brand","category"],
  "query": {
    "bool": {
      "filter": [
        {"term": {"tenant_id": "tenant-a"}},
        {"term": {"available": true}}
      ],
      "must": [
        {"match_phrase_prefix": {"name_suggest": {"query": "water", "max_expansions": 25}}}
      ]
    }
  }
}

Expected invariant: only authorized, available tenant-a products are candidates. The exact hit order is not a cross-platform contract; compare relevance using fixed judgments instead of copying raw scores between engines.

6. Facets and business signals: keep filtering semantics visible

Search plus brand/category facets
POST atlasmart-product-search-v1/_search
{
  "size": 10,
  "query": {
    "bool": {
      "filter": [
        {"term":{"tenant_id":"tenant-a"}},
        {"term":{"available":true}}
      ],
      "must": [
        {"multi_match": {
          "query":"waterproof hiking",
          "fields":["name^3","description"],
          "analyzer":"atlasmart_search"
        }}
      ]
    }
  },
  "aggs": {
    "brands":{"terms":{"field":"brand","size":10}},
    "categories":{"terms":{"field":"category","size":10}}
  }
}

Facet counts answer “how many matching documents are in each bucket under this query/filter context?” They are not global catalog counts. If the UI needs disjunctive faceting, define that explicitly and test it; do not improvise by removing filters at random.

Popularity, freshness, margin, inventory, and personalization can influence ranking, but each should have a bounded role and an evaluation set. Raw BM25 scores and business-feature values are on different scales, so adding them blindly is not a stable scoring contract.

7. Personalization boundary: identity is not just another feature

Personalization introduces sensitive user/tenant context. A product-search service should first apply authorization and catalog eligibility, then personalize within the authorized candidate set. Never retrieve cross-tenant documents and “filter them out later” in application code. Keep the source of personalization features, retention policy, consent basis, and failure fallback explicit. If the personalization service is unavailable, the product should degrade to a tested non-personalized ranking rather than fail open or expose another user’s features.

Wrong approach. Put product search, raw logs, security events, and traces into one dynamic mega-index so every team can “search everything.” The result is mapping conflict risk, uncontrolled field growth, conflicting retention rules, excessive privileges, and noisy-neighbor contention. Repair: model each workload from its own query/retention/security contract, then decide where physical consolidation is safe.

8. Mini lab: prove relevance and workload shape

  1. Run the lexical/facet query above with five representative AtlasMart product queries.
  2. Record top-5 IDs, facet counts, request latency, and the exact query/mapping version.
  3. Add one synonym that should increase recall. Re-run the same judgments; verify the intended query improves without unrelated regressions.
  4. Run 100 autocomplete requests with a fixed prefix list and record p50/p95/p99 from the client. Do not publish synthetic numbers as production capacity.
  5. In parallel, run a small log-ingest loop from Lesson 2 and observe whether product p99 changes. This is the first noisy-neighbor signal, not yet proof of causality.

Acceptance: the team can explain which fields are searchable, filterable, aggregatable, and personalized; which metrics demonstrate storefront health; and why this index should not inherit telemetry retention or security policy.

Check your understanding

  1. Why should brand/category facets use exact-value fields?
  2. Why is a synonym change a release?
  3. Does a green cluster prove product search is healthy?
  4. Where should tenant authorization be applied?
  5. What is the bridge to Lesson 2?
Review the answers

1. Because facets group exact values; analyzed prose changes tokens and therefore bucket semantics.

2. It changes token expansion, recall, and ranking, so it needs versioning and relevance regression tests.

3. No. User-visible relevance, p95/p99 latency, error rate, facet correctness, and freshness can fail while cluster health remains green.

4. Before or as part of retrieval, not as a best-effort post-filter after unauthorized candidates have been retrieved.

5. Product search optimizes relevance and interactive latency; log analytics instead prioritizes structured event contracts, sustained ingest, cardinality control, and retention.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.