Tune indexing by balancing throughput, freshness, durability, merge debt, and backpressure.

Indexing Tuning: Bulk Size, Refresh, Replicas, Pipelines, Merge, Translog/Durability, and Backpressure

Build capacity and performance plans from measured workload dimensions and tail behavior rather than generic shard, heap, bulk, or hardware rules.

Intermediate → Advanced165–220 minutesIndexing throughput/freshness experiment · Chapter 30 · Lesson 02Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · free/local benchmark pathLast reviewed: September 2026

Learning outcomes

01

Benchmark bulk request size and client concurrency to find a stable plateau without using a universal batch number.

02

Explain refresh, replicas, ingest pipelines, merge work, translog durability, and indexing backpressure as separate mechanisms.

03

Recognize per-item bulk failures and 429/rejection signals as capacity evidence rather than retry noise to hide.

04

Change indexing settings without silently violating freshness, durability, or availability requirements.

05

Run one controlled AtlasMart indexing experiment and prove both performance and recovery to the original state.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned performance baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The established local endpoints remain Elasticsearch at https://localhost:9200 and OpenSearch at https://localhost:9201 on the atlasmart-search Docker network. The generation environment did not execute live clusters, so this chapter never invents throughput, p95/p99, GC, disk, vector-recall, or cost results: numeric fields shown in report templates are deliberately blank/null until measured.

1. AtlasMart problem: the “bigger bulk is faster” tuning loop

AtlasMart's importer is slow, so a developer increases each bulk request until throughput initially rises. Then memory pressure, long request times, merge debt, and retries appear. The error is treating bulk size as an isolated constant. Bulk requests amortize network and parsing overhead, but each request also occupies coordinating/primary/replica resources and can multiply memory pressure when many clients run concurrently.

2. Discover bulk size and concurrency empirically

Keep the document fixture, shard layout, refresh policy, and hardware fixed. Increase bulk size geometrically—small to larger—until achieved indexing throughput stops improving materially or tail latency/resource pressure worsens. Then choose the smaller stable point, not the largest request accepted. Repeat with client concurrency because one huge bulk and many medium bulks have different queueing/memory behavior.

Finite bulk/concurrency experiment matrix
bulk_docs = [100, 200, 400, 800]        # example experiment points, not recommendations
clients   = [1, 2, 4, 8]                 # stop before unsafe saturation
for each pair:
  warm up with the same fixture shape
  run for the same measurement window
  record achieved docs/s and MB/s
  record p95/p99 bulk latency
  count item failures and HTTP 429/rejections
  record indexing lag, CPU, heap/GC, disk, merge and queue evidence
  STOP if swap/OOM risk, sustained rejection, or unrelated workload SLO breach appears

3. Refresh controls visibility—not durability

A refresh opens new searchable Lucene segments; it does not mean “flush everything safely to disk.” Shorter refresh intervals improve write-to-search freshness but create more frequent segment/searcher work and can increase future merge pressure. Longer intervals can improve indexing throughput when the freshness SLO allows it. In both current products, forcing refresh=true repeatedly is intentionally expensive; wait_for waits for a refresh rather than forcing one.

Compare the same load at two refresh policies
PUT atlasmart-perf-v1/_settings
{"index":{"refresh_interval":"1s"}}
# run the fixed benchmark window; record write→search visibility and all resource evidence

PUT atlasmart-perf-v1/_settings
{"index":{"refresh_interval":"10s"}}
# rerun exactly the same offered load and query mix

# restore the original policy after the experiment
PUT atlasmart-perf-v1/_settings
{"index":{"refresh_interval":"1s"}}

The correct choice is whichever meets both the freshness SLO and capacity/headroom objective. “10s was faster” is not sufficient if AtlasMart requires products to appear in search within 2 seconds.

4. Replicas trade write work for availability and read options

Each replica must receive indexed changes, so replicas add write work and recovery/storage cost. They also provide redundant shard copies and can add search-serving copies. During an initial bulk load from a durable source that can be replayed, temporarily using zero replicas can be a legitimate controlled optimization. It is not a general production tuning recommendation, and replicas are not backups: correlated deletion/corruption can propagate to all copies, so snapshots remain a separate recovery control.

Wrong approach: remove replicas and weaken durability to make a benchmark “win,” then compare that number with a resilient production configuration. Repair: benchmark the availability/durability configuration you intend to operate, and if you test a faster initial-load mode, label it as a different risk profile with an explicit restoration gate.

5. Ingest pipelines move CPU into the write path

Grok-like parsing, script processors, enrichment, inference, or other processors consume CPU/memory and may access additional data structures. Benchmark with the actual pipeline enabled; otherwise the capacity test measures a different product. Record processor failures and dead-letter/failure-path behavior as correctness signals, not only aggregate docs/s.

6. Merges are deferred work, not free throughput

Lucene writes immutable segments and later merges them. A short indexing test can report high throughput while merge debt accumulates and disk utilization rises. Continue the run long enough to reach steady behavior, record segment/merge statistics, and include a post-load observation window. Do not disable normal merge safety controls merely to improve a short chart.

7. Translog/durability is a risk contract

The translog supports recovery of acknowledged operations between Lucene commits. Do not conflate refresh with fsync/durability. Product/version settings around request durability, flush thresholds, and recovery must be treated as data-loss contracts. Changing them requires a stated RPO and failure test, not a throughput preference.

8. Backpressure tells the client to slow down

Elasticsearch exposes indexing pressure and write-thread-pool rejections; OpenSearch exposes indexing pressure and can add shard indexing backpressure in supported configurations. When the server rejects work, a client should apply bounded exponential backoff with jitter and a retry budget, while still inspecting each Bulk API item. Infinite immediate retries convert overload into more overload.

Inspect write pressure and item-level failures
GET _cat/thread_pool/write?v&h=node_name,active,queue,rejected,completed
GET _nodes/stats/indexing_pressure,thread_pool,jvm,fs

POST _bulk
{"index":{"_index":"atlasmart-perf-v1","_id":"..."}}
{"...":"..."}
# parse every item.status and item.error; HTTP 200 for the bulk envelope does not mean every item succeeded

9. Controlled AtlasMart indexing lab

All Chapter 30 mandatory exercises preserve the course's established free/local security boundary: Elasticsearch 9.5.3 at https://localhost:9200 authenticated with ELASTIC_PASSWORD and the copied CA atlasmart-es-http-ca; OpenSearch 3.8.0 at https://localhost:9201 authenticated with OPENSEARCH_INITIAL_ADMIN_PASSWORD. The shared Docker network remains atlasmart-search. OpenSearch's -k examples are for the disposable demo certificate only and are not production TLS guidance. The default local fixture uses one primary and zero replicas because it is a workstation lab; any availability or node-failure conclusion must be tested in a disposable multi-node variant rather than inferred from this single-node setup.

AtlasMart performance fixture mapping
PUT atlasmart-perf-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "refresh_interval": "1s"
  },
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "sku":        {"type":"keyword"},
      "tenant_id":  {"type":"keyword"},
      "name":       {"type":"text", "fields":{"raw":{"type":"keyword"}}},
      "category":   {"type":"keyword"},
      "brand":      {"type":"keyword"},
      "price":      {"type":"scaled_float", "scaling_factor":100},
      "available":  {"type":"boolean"},
      "event_time": {"type":"date"},
      "popularity": {"type":"float"},
      "embedding":  {"type":"dense_vector", "dims":8, "index":true, "similarity":"cosine"}
    }
  }
}
Cross-product mapping note. The Elasticsearch dense_vector example above is intentionally product-specific. For OpenSearch, create an equivalent knn_vector mapping using the Chapter 25 pinned engine/method settings rather than copying this JSON. The benchmark contract is equivalent dimensions, source vectors, filters, candidate count, and judgments—not identical mapping syntax.
Capture server-side evidence before and after each run
GET _cluster/health
GET _cat/nodes?v&h=name,cpu,heap.percent,ram.percent,disk.used_percent,node.role,master
GET _cat/thread_pool/search,write?v&h=node_name,name,active,queue,rejected,completed
GET _nodes/stats/jvm,process,os,fs,indices,thread_pool,indexing_pressure
GET atlasmart-perf-v1/_stats?level=shards
GET _cat/recovery/atlasmart-perf-v1?v
GET atlasmart-perf-v1/_segments
  1. Create identical fresh indices on both products with the established one-primary/zero-replica workstation topology.
  2. Seed or generate the same deterministic product shape; do not mix datasets between runs.
  3. Run two load levels using the same bulk size/concurrency baseline and record tails, throughput, indexing lag, item failures, CPU/heap/GC/disk, and merges.
  4. Choose one variable—bulk size, concurrency, or refresh interval—and repeat.
  5. Accept the change only if it improves the target without breaking freshness, relevance, recovery, or error/rejection budgets.
  6. Restore the original setting and prove the fixture returns to its baseline state.
Benchmark report: fill only measured values
{
  "run_id": "atlasmart-perf-YYYYMMDD-NN",
  "platform": {"product": null, "version": null, "topology": null},
  "dataset": {"documents": null, "source_bytes": null, "indexed_bytes": null, "warm_state": null},
  "offered_load": {"search_qps": null, "write_ops_s": null, "clients": null},
  "results": {
    "achieved_search_qps": null,
    "achieved_write_ops_s": null,
    "latency_ms": {"p50": null, "p95": null, "p99": null},
    "indexing_lag_ms": {"p95": null, "p99": null},
    "errors": null,
    "rejections": null,
    "relevance": {"ndcg_at_10": null, "vector_recall_at_10": null},
    "recovery_seconds": null
  },
  "resources": {"cpu": null, "heap": null, "gc": null, "disk_io": null, "network": null},
  "cost": {"currency": null, "window_cost": null, "cost_per_1k_queries": null},
  "notes": {"warmup": null, "cache_state": null, "one_change_from_baseline": null}
}

10. Production judgment

Indexing tuning is successful when the cluster sustains required write rate with bounded indexing lag, acceptable search tails, no hidden merge/recovery collapse, and the intended durability/availability contract. Throughput by itself is not the objective.

Check your understanding

  1. Why should bulk size be discovered rather than copied from a blog?
  2. What does refresh control?
  3. When can zero replicas be a legitimate experiment?
  4. What should a client do with 429/rejections?
  5. What stays fixed in Lesson 3?
Review the answers

1. Its optimum depends on document size, ingest work, shard layout, hardware, concurrency, and the current product/version.

2. Search visibility of recent changes; it is distinct from translog durability and flush/commit behavior.

3. For a controlled initial load from a durable replayable source, with explicit risk acceptance and a gate to restore redundancy.

4. Treat them as overload evidence, slow down with bounded backoff/jitter, preserve a retry budget, and inspect item-level failures.

5. The measured workload and indexing baseline while search-side query shape, filtering, caches, routing, shard fan-out, precomputation, and aggregation strategy are evaluated.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.