Choose workload isolation from SLO, security, lifecycle, ownership, and noisy-neighbor evidence.
Choose Index/Shard/Lifecycle/Search Strategies per Workload Instead of Running Every Use Case in One Undifferentiated Cluster
Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.
Learning outcomes
Translate product search, logs, security analytics, and observability requirements into separate mapping, shard, lifecycle, query, and security plans.
Distinguish logical isolation from physical workload isolation and choose the boundary using measured noisy-neighbor risk.
Build a mixed-load experiment that records per-workload p95/p99 latency, ingest freshness, errors/rejections, and storage growth.
Allocate operational ownership and cost without hiding shared-cluster contention.
Create an evidence package that can support consolidation, split-cluster, or migration decisions and bridge into Chapter 29.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. AtlasMart architecture review: “one cluster” is a deployment choice, not a data model
The four Chapter 28 workloads can run on one local lab cluster. Production should not assume they belong together. Consolidation can reduce operational overhead and improve resource utilization; isolation can protect latency, credentials, retention, upgrade cadence, and incident blast radius. The decision needs measured workload dimensions and failure semantics.
Start with SLOs and ownership. A product-search team may page on p99 latency and relevance regressions. A logging team may own ingest freshness and storage growth. Security may require restricted access, immutable evidence, and long retention. Observability may need high ingest burst tolerance and independent health visibility precisely when the application cluster is degraded.
2. Four-workload architecture matrix
| Workload | Lab target | Query shape | Lifecycle emphasis | Primary security boundary |
|---|---|---|---|---|
| Product search | atlasmart-product-search-v1 |
low-latency ranked retrieval + facets | versioned catalog/reindex, not age-only deletion | shop/API role + tenant/catalog filters |
| Logs | atlasmart-logs-app data stream |
time filters, text search, service aggregations | rollover + finite retention | operations readers/writers |
| Security | atlasmart-security-events-v1 |
detections + investigative lookback | longer evidence retention + snapshot policy | security-only roles/detector actions |
| Observability | logs/spans/metrics indices | trace correlation + time-series aggregations | signal-specific retention/sampling | SRE/observability roles |
3. Shard and index strategy: measure before multiplying shards
Small lab indices use one primary because extra shards would add coordination overhead without parallelism benefit. Production shard count follows measured data volume, query concurrency, recovery time, and node topology from Chapters 13 and 30—not a universal “50 GB shard” rule. Product search may prefer a compact index optimized for query fan-out; telemetry may create time-partitioned backing indices through data streams; long-retention security data may use tiers or dedicated capacity.
Routing can isolate tenant access or reduce fan-out in some designs, but it can also create hotspots. Do not choose custom routing merely because tenant IDs exist; measure tenant skew and recovery implications.
4. Logical isolation controls before physical separation
| Control | What it isolates | What it does not isolate |
|---|---|---|
| Separate indices/data streams | mappings, settings, lifecycle targets | CPU/heap/disk/network on shared nodes |
| Roles/API keys | authorized actions/data | resource contention |
| Data tiers/node attributes | storage classes and allocation | coordinator/cluster-state coupling |
| Query limits/backpressure | some abusive requests | all noisy-neighbor effects |
| Separate clusters | failure/upgrade/resource domain | cross-cluster operational complexity/cost |
Physical separation is warranted when shared failure modes violate an SLO/security requirement, when maintenance windows conflict, or when the cost of resource contention exceeds the operational cost of another cluster.
5. Mixed-load harness: capture evidence, do not invent it
All Chapter 28 labs keep the existing local endpoints and
security assumptions: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD and the copied CA file
atlasmart-es-http-ca; OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's
demo certificate trust bypass (-k) is acceptable
only for this disposable local lab, never production. The shared
Docker network remains atlasmart-search. Lab
indices use one primary and zero replicas so a single-node
workstation can complete the exercises; production redundancy
decisions are deliberately separate.
# Run only against the disposable local lab.
# Example: python mixed_load.py https://localhost:9200 elastic "$ELASTIC_PASSWORD" atlasmart-es-http-ca
import base64, json, ssl, sys, time, statistics, urllib.request
from concurrent.futures import ThreadPoolExecutor, as_completed
endpoint, user, password, cafile = sys.argv[1:5]
ctx = ssl.create_default_context(cafile=cafile)
auth = base64.b64encode(f"{user}:{password}".encode()).decode()
CASES = [
("product", "/atlasmart-product-search-v1/_search", {"size":5,"query":{"match":{"name":"waterproof hiking"}}}),
("logs", "/atlasmart-logs-app/_search", {"size":0,"query":{"range":{"@timestamp":{"gte":"now-1h"}}},"aggs":{"services":{"terms":{"field":"service.name","size":10}}}}),
("security", "/atlasmart-security-events-v1/_search", {"size":0,"query":{"term":{"event.outcome":"failure"}}}),
("traces", "/atlasmart-otel-spans-v1/_search", {"size":0,"aggs":{"p99":{"percentiles":{"field":"duration_ms","percents":[99]}}}}),
]
def one(case):
name, path, body = case
req = urllib.request.Request(endpoint+path, data=json.dumps(body).encode(), method="POST",
headers={"Authorization":"Basic "+auth,"Content-Type":"application/json"})
t0=time.perf_counter()
try:
with urllib.request.urlopen(req, context=ctx, timeout=15) as r:
r.read(); status=r.status
except Exception as e:
status=getattr(e,'code',0) or 0
return name, (time.perf_counter()-t0)*1000, status
samples=[]
with ThreadPoolExecutor(max_workers=4) as ex:
futs=[ex.submit(one, CASES[i % len(CASES)]) for i in range(200)]
for f in as_completed(futs): samples.append(f.result())
for name in sorted({x[0] for x in samples}):
vals=sorted(x[1] for x in samples if x[0]==name and 200 <= x[2] < 300)
errs=sum(1 for x in samples if x[0]==name and not (200 <= x[2] < 300))
if vals:
def pct(p): return vals[min(len(vals)-1, int(round((p/100)*(len(vals)-1))))]
print(name, "n=",len(vals),"errors=",errs,"p50_ms=",round(pct(50),2),"p95_ms=",round(pct(95),2),"p99_ms=",round(pct(99),2))
else:
print(name, "no successful samples; errors=", errs)
For OpenSearch’s local demo certificates, create a trusted local CA context if possible. If you deliberately use an unverified TLS context for the disposable OpenSearch lab, label it as unsafe and never copy that pattern into production. Run each workload alone first, then the same workload under mixed concurrency. Keep dataset, cache state, concurrency, query set, and measurement window fixed.
6. What to record
| Evidence | Product | Logs | Security | Observability |
|---|---|---|---|---|
| Tail latency | p95/p99 search | p95/p99 investigation query | p95/p99 investigation/detection query | p95/p99 trace/metric queries |
| Freshness | catalog update → searchable | event time → searchable | security event → searchable/detectable | span/log/metric → visible |
| Write pressure | catalog bulk/reindex | events/s, bulk rejections | events/s + detector load | telemetry ingest + sampling |
| Quality | judged relevance/facet correctness | schema/query completeness | false-positive/negative review | correlation coverage/service-map completeness |
| Storage | catalog index bytes | bytes/day × retention | bytes/day × longer retention | per-signal bytes/day × sampling/retention |
| Security | shop/API least privilege | ops read/write | restricted security roles | SRE/telemetry roles |
7. Decision patterns
Consolidate when datasets are modest, SLOs are compatible, security boundaries are enforceable, maintenance cadence is shared, and mixed-load tests preserve headroom. Split when one workload repeatedly violates another’s latency/freshness target, when security/compliance requires a separate blast radius, or when lifecycle/storage economics need incompatible hardware/tier policies.
A hybrid design is common: product search in one serving cluster; logs/observability in another telemetry cluster; security data in a restricted analytics cluster or logically isolated domain with independent retention. Managed services can change the available isolation units and cost model, so repeat the reasoning for the exact service rather than copying a self-managed topology.
8. Architecture decision record for AtlasMart
Complete one row per workload: owner, SLOs, ingest rate, storage/retention, peak query pattern, security role, backup/RPO/RTO, current cluster, contention evidence, and chosen isolation. The output is more valuable than a diagram because it preserves the assumptions behind the topology.
Before production approval, repeat the mixed-load test with realistic document counts, shard placement, replica count, TLS/auth, background merges, snapshots/lifecycle jobs, and failover headroom. Tail latency and recovery behavior under node loss matter more than the empty-lab average.
9. Bridge to Chapter 29: portability must be workload-specific too
Chapter 29 starts with Assess API, Mapping, Query, Security, Plugin, Snapshot, and Client Compatibility Before Migration. The Chapter 28 matrix becomes its migration inventory. Product relevance, ECS-oriented logs, OpenSearch Security Analytics detectors, Elastic Security rules, APM/service-map pipelines, ILM/ISM policies, and managed-service integrations do not migrate as one generic “index.” Each workload needs its own compatibility and rollback evidence.
Check your understanding
- What does separate indices guarantee?
- When is a separate cluster justified?
- Why run each workload alone before mixed load?
- What must never be compared as if identical across workloads?
- What evidence carries into migration planning?
Review the answers
1. Mapping/settings/lifecycle boundaries, but not CPU, heap, disk, network, coordinator, or cluster-state isolation on shared nodes.
2. When measured shared failure/resource effects, security/compliance, lifecycle/hardware needs, or maintenance ownership make a shared failure domain unacceptable.
3. It establishes a baseline so mixed-load degradation can be attributed rather than guessed.
4. Raw latency/throughput targets without accounting for each workload’s user contract, query shape, freshness, retention, and security requirements.
5. Per-workload mappings, queries, analyzers, lifecycle, security, plugins/UI workflows, baselines, relevance/detection/correlation tests, and rollback requirements.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References and current-version checks
- Elastic Stack 9.5.3 release
- Elasticsearch 9.5.3 release notes
- Elastic Common Schema 9.5 reference
- ECS getting started and normalization
- ECS log fields
- Elastic Observability fields and object schemas
- Elastic ECS-formatted application logs
- OpenSearch 3.8 version history
- OpenSearch Security Analytics overview
- OpenSearch Security Analytics detectors
- OpenSearch Security Analytics access control
- OpenSearch APM configuration
- OpenSearch Trace Analytics
- OpenTelemetry logs data model
- OpenTelemetry service semantic conventions