Choose workload isolation from SLO, security, lifecycle, ownership, and noisy-neighbor evidence.

Choose Index/Shard/Lifecycle/Search Strategies per Workload Instead of Running Every Use Case in One Undifferentiated Cluster

Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.

Intermediate → Advanced160–210 minutesMixed-workload architecture lab · Chapter 28 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · ECS 9.5.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Translate product search, logs, security analytics, and observability requirements into separate mapping, shard, lifecycle, query, and security plans.

02

Distinguish logical isolation from physical workload isolation and choose the boundary using measured noisy-neighbor risk.

03

Build a mixed-load experiment that records per-workload p95/p99 latency, ingest freshness, errors/rejections, and storage growth.

04

Allocate operational ownership and cost without hiding shared-cluster contention.

05

Create an evidence package that can support consolidation, split-cluster, or migration decisions and bridge into Chapter 29.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The current Elastic Common Schema reference is ECS 9.5.0. The mandatory AtlasMart labs use free/local HTTP APIs and deterministic fixtures. The generation environment did not execute live clusters, so numeric latency, throughput, shard growth, and storage values shown as acceptance criteria are measurement instructions—not fabricated captured results.

1. AtlasMart architecture review: “one cluster” is a deployment choice, not a data model

The four Chapter 28 workloads can run on one local lab cluster. Production should not assume they belong together. Consolidation can reduce operational overhead and improve resource utilization; isolation can protect latency, credentials, retention, upgrade cadence, and incident blast radius. The decision needs measured workload dimensions and failure semantics.

Start with SLOs and ownership. A product-search team may page on p99 latency and relevance regressions. A logging team may own ingest freshness and storage growth. Security may require restricted access, immutable evidence, and long retention. Observability may need high ingest burst tolerance and independent health visibility precisely when the application cluster is degraded.

2. Four-workload architecture matrix

Workload Lab target Query shape Lifecycle emphasis Primary security boundary
Product search atlasmart-product-search-v1 low-latency ranked retrieval + facets versioned catalog/reindex, not age-only deletion shop/API role + tenant/catalog filters
Logs atlasmart-logs-app data stream time filters, text search, service aggregations rollover + finite retention operations readers/writers
Security atlasmart-security-events-v1 detections + investigative lookback longer evidence retention + snapshot policy security-only roles/detector actions
Observability logs/spans/metrics indices trace correlation + time-series aggregations signal-specific retention/sampling SRE/observability roles

3. Shard and index strategy: measure before multiplying shards

Small lab indices use one primary because extra shards would add coordination overhead without parallelism benefit. Production shard count follows measured data volume, query concurrency, recovery time, and node topology from Chapters 13 and 30—not a universal “50 GB shard” rule. Product search may prefer a compact index optimized for query fan-out; telemetry may create time-partitioned backing indices through data streams; long-retention security data may use tiers or dedicated capacity.

Routing can isolate tenant access or reduce fan-out in some designs, but it can also create hotspots. Do not choose custom routing merely because tenant IDs exist; measure tenant skew and recovery implications.

4. Logical isolation controls before physical separation

Control What it isolates What it does not isolate
Separate indices/data streams mappings, settings, lifecycle targets CPU/heap/disk/network on shared nodes
Roles/API keys authorized actions/data resource contention
Data tiers/node attributes storage classes and allocation coordinator/cluster-state coupling
Query limits/backpressure some abusive requests all noisy-neighbor effects
Separate clusters failure/upgrade/resource domain cross-cluster operational complexity/cost

Physical separation is warranted when shared failure modes violate an SLO/security requirement, when maintenance windows conflict, or when the cost of resource contention exceeds the operational cost of another cluster.

5. Mixed-load harness: capture evidence, do not invent it

All Chapter 28 labs keep the existing local endpoints and security assumptions: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab, never production. The shared Docker network remains atlasmart-search. Lab indices use one primary and zero replicas so a single-node workstation can complete the exercises; production redundancy decisions are deliberately separate.

Python 3 standard-library mixed-load skeleton
# Run only against the disposable local lab.
# Example: python mixed_load.py https://localhost:9200 elastic "$ELASTIC_PASSWORD" atlasmart-es-http-ca
import base64, json, ssl, sys, time, statistics, urllib.request
from concurrent.futures import ThreadPoolExecutor, as_completed

endpoint, user, password, cafile = sys.argv[1:5]
ctx = ssl.create_default_context(cafile=cafile)
auth = base64.b64encode(f"{user}:{password}".encode()).decode()

CASES = [
  ("product", "/atlasmart-product-search-v1/_search", {"size":5,"query":{"match":{"name":"waterproof hiking"}}}),
  ("logs", "/atlasmart-logs-app/_search", {"size":0,"query":{"range":{"@timestamp":{"gte":"now-1h"}}},"aggs":{"services":{"terms":{"field":"service.name","size":10}}}}),
  ("security", "/atlasmart-security-events-v1/_search", {"size":0,"query":{"term":{"event.outcome":"failure"}}}),
  ("traces", "/atlasmart-otel-spans-v1/_search", {"size":0,"aggs":{"p99":{"percentiles":{"field":"duration_ms","percents":[99]}}}}),
]

def one(case):
    name, path, body = case
    req = urllib.request.Request(endpoint+path, data=json.dumps(body).encode(), method="POST",
        headers={"Authorization":"Basic "+auth,"Content-Type":"application/json"})
    t0=time.perf_counter()
    try:
        with urllib.request.urlopen(req, context=ctx, timeout=15) as r:
            r.read(); status=r.status
    except Exception as e:
        status=getattr(e,'code',0) or 0
    return name, (time.perf_counter()-t0)*1000, status

samples=[]
with ThreadPoolExecutor(max_workers=4) as ex:
    futs=[ex.submit(one, CASES[i % len(CASES)]) for i in range(200)]
    for f in as_completed(futs): samples.append(f.result())

for name in sorted({x[0] for x in samples}):
    vals=sorted(x[1] for x in samples if x[0]==name and 200 <= x[2] < 300)
    errs=sum(1 for x in samples if x[0]==name and not (200 <= x[2] < 300))
    if vals:
        def pct(p): return vals[min(len(vals)-1, int(round((p/100)*(len(vals)-1))))]
        print(name, "n=",len(vals),"errors=",errs,"p50_ms=",round(pct(50),2),"p95_ms=",round(pct(95),2),"p99_ms=",round(pct(99),2))
    else:
        print(name, "no successful samples; errors=", errs)

For OpenSearch’s local demo certificates, create a trusted local CA context if possible. If you deliberately use an unverified TLS context for the disposable OpenSearch lab, label it as unsafe and never copy that pattern into production. Run each workload alone first, then the same workload under mixed concurrency. Keep dataset, cache state, concurrency, query set, and measurement window fixed.

6. What to record

Evidence Product Logs Security Observability
Tail latency p95/p99 search p95/p99 investigation query p95/p99 investigation/detection query p95/p99 trace/metric queries
Freshness catalog update → searchable event time → searchable security event → searchable/detectable span/log/metric → visible
Write pressure catalog bulk/reindex events/s, bulk rejections events/s + detector load telemetry ingest + sampling
Quality judged relevance/facet correctness schema/query completeness false-positive/negative review correlation coverage/service-map completeness
Storage catalog index bytes bytes/day × retention bytes/day × longer retention per-signal bytes/day × sampling/retention
Security shop/API least privilege ops read/write restricted security roles SRE/telemetry roles

7. Decision patterns

Consolidate when datasets are modest, SLOs are compatible, security boundaries are enforceable, maintenance cadence is shared, and mixed-load tests preserve headroom. Split when one workload repeatedly violates another’s latency/freshness target, when security/compliance requires a separate blast radius, or when lifecycle/storage economics need incompatible hardware/tier policies.

A hybrid design is common: product search in one serving cluster; logs/observability in another telemetry cluster; security data in a restricted analytics cluster or logically isolated domain with independent retention. Managed services can change the available isolation units and cost model, so repeat the reasoning for the exact service rather than copying a self-managed topology.

Wrong approach. Observe contention in one mixed run and immediately split every workload into separate clusters. A single test can be confounded by cache warmness, shard layout, background merges, or an unrealistic dataset. Repair: establish per-workload baselines, reproduce the contention, identify the resource bottleneck, test a lower-cost mitigation, and split only when the SLO/security/ownership evidence supports it.

8. Architecture decision record for AtlasMart

Complete one row per workload: owner, SLOs, ingest rate, storage/retention, peak query pattern, security role, backup/RPO/RTO, current cluster, contention evidence, and chosen isolation. The output is more valuable than a diagram because it preserves the assumptions behind the topology.

Before production approval, repeat the mixed-load test with realistic document counts, shard placement, replica count, TLS/auth, background merges, snapshots/lifecycle jobs, and failover headroom. Tail latency and recovery behavior under node loss matter more than the empty-lab average.

9. Bridge to Chapter 29: portability must be workload-specific too

Chapter 29 starts with Assess API, Mapping, Query, Security, Plugin, Snapshot, and Client Compatibility Before Migration. The Chapter 28 matrix becomes its migration inventory. Product relevance, ECS-oriented logs, OpenSearch Security Analytics detectors, Elastic Security rules, APM/service-map pipelines, ILM/ISM policies, and managed-service integrations do not migrate as one generic “index.” Each workload needs its own compatibility and rollback evidence.

Check your understanding

  1. What does separate indices guarantee?
  2. When is a separate cluster justified?
  3. Why run each workload alone before mixed load?
  4. What must never be compared as if identical across workloads?
  5. What evidence carries into migration planning?
Review the answers

1. Mapping/settings/lifecycle boundaries, but not CPU, heap, disk, network, coordinator, or cluster-state isolation on shared nodes.

2. When measured shared failure/resource effects, security/compliance, lifecycle/hardware needs, or maintenance ownership make a shared failure domain unacceptable.

3. It establishes a baseline so mixed-load degradation can be attributed rather than guessed.

4. Raw latency/throughput targets without accounting for each workload’s user contract, query shape, freshness, retention, and security requirements.

5. Per-workload mappings, queries, analyzers, lifecycle, security, plugins/UI workflows, baselines, relevance/detection/correlation tests, and rollback requirements.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.