Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals
Diagnose an Indexing/Storage Problem Using Segment Counts, Merge Pressure, Disk I/O, and Refresh Behavior
Use segment, refresh, merge, translog, node and application-latency evidence to diagnose an AtlasMart indexing/storage regression and verify a causal repair.
Learning outcomes
AtlasMart reports a production symptom: indexing throughput fell, search p99 doubled, disk utilization rose, and operators see hundreds of segments. The wrong response is to force merge immediately. A storage diagnosis must establish a timeline and causal chain across refresh rate, segment creation, update/delete churn, merge pressure, translog/flush behavior, disk I/O and cache state.
Build an evidence timeline from application SLOs and index/node storage metrics before changing settings.
Distinguish excessive refresh/small-segment pressure from merge debt, translog/flush pressure and cold-cache I/O.
Use segment, docs.deleted, merge, refresh, flush, translog and filesystem statistics together.
Run one-variable experiments and reject repairs that only improve averages while p99 or recovery risk worsens.
Produce a reversible AtlasMart storage runbook with acceptance, rollback and cleanup criteria.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual runtime with
GET _nodes/jvm instead of hard-coding a JDK
patch. Labs use one primary and zero replicas unless a step
explicitly says otherwise. No moving latest tags,
no manual editing of Lucene files, and no production force
merge are used.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Start from the user-visible regression, not from a favorite metric
Collect the time window in which AtlasMart search/indexing degraded. Record p50/p95/p99 latency, indexing rate, errors/rejections, freshness lag and deployment/configuration changes. Only then align server metrics. Segment count alone is not a diagnosis; a small busy index can have many harmless young segments, while a few huge segments can still be I/O-bound.
The goal is a falsifiable hypothesis such as “a refresh-per-document change created small-segment/merge pressure that saturated storage, raising write and search tail latency.”
GET _cluster/health
GET _cat/indices/atlasmart-products*?v&h=health,status,index,pri,rep,docs.count,docs.deleted,store.size,pri.store.size
GET _cat/segments/atlasmart-products*?v
GET atlasmart-products*/_stats/docs,store,indexing,search,refresh,flush,merge,segments,translog?level=shards
GET _nodes/stats/fs,indices,jvm,process,thread_pool
GET _nodes/hot_threads
2. Read symptom patterns, then verify competing explanations
| Observed pattern | Hypothesis to test | Disconfirming evidence |
|---|---|---|
| Refresh count/rate jumps + many small segments | Forced/too-frequent refresh creates merge work | Refresh unchanged; segment sizes/count stable. |
| docs.deleted/store grow with update rate | Update/delete churn awaits merge reclamation | No churn; deleted docs flat. |
| merge current/bytes rise with disk await and p99 | Merge/storage contention | Disk idle and CPU/query profile dominates instead. |
| Translog grows; recovery/flush behavior changes | Flush threshold/durability/write pattern changed | Translog/flush counters normal for same write rate. |
| First queries slow after restart, then improve | Cold filesystem cache | Repeated warm runs remain equally slow. |
| Search slow but merge/storage normal | Query/relevance/heap/thread-pool cause outside this chapter | Profile, GC/rejections/hot threads point elsewhere. |
3. Build a controlled AtlasMart failure fixture
In a disposable lab only, compare two write modes against the
same dataset: A uses bulk writes and a
normal/explicit refresh cadence; B deliberately
uses refresh=true per write for a bounded number of
documents. Capture segment/refresh/merge metrics and latency.
This safely demonstrates mechanism without filling disks or
killing host processes.
Do not expect exact numeric outcomes across machines. The pass condition is directional evidence on your environment plus a clear causal explanation.
Index A: atlasmart-storage-good
1 primary, 0 replicas
bulk 1,000 deterministic documents
refresh only at test boundary
Index B: atlasmart-storage-bad
same mapping + documents
index each document with refresh=true
keep document count bounded
For both capture:
_stats/refresh,merge,segments,docs,store
_segments
wall-clock ingest time
fixed-query p50/p95/p99 after indexing
host disk/CPU evidence
Cleanup both disposable indices after comparison.
Do not intentionally fill the filesystem, disable safety mechanisms, or force merge a production/hot index to create a “better demo.” Controlled request patterns are enough to expose segment/refresh behavior.
4. One-variable repair: remove forced refresh and retest
If the evidence supports the refresh hypothesis, change only
that variable first: remove per-write forced refresh, batch
writes, and use the product requirement to choose normal
periodic refresh or wait_for where a caller truly
needs search visibility. Rerun the identical fixture and query
set.
Verify that the repair preserves correctness: expected documents are searchable within the agreed freshness bound, no writes are lost, and application behavior that depended on immediate search is either redesigned or explicitly waits for visibility.
Before -> after comparison:
document count/checksum: must match
search relevance fixture: must pass
freshness SLO: must pass
indexing docs/s: record
p50/p95/p99 search + indexing latency: record
refresh total/rate: record
segment count/size distribution: record
merge bytes/time/current: record
disk latency/utilization: record
translog/flush stats: record
Reject the change if correctness or tail-latency/recovery objectives regress.
5. Force merge is not the repair for a bad hot-index refresh policy
A force merge might temporarily reduce segment count, but if the application continues forcing refreshes, small segments begin accumulating again. Worse, the force merge itself adds heavy I/O and temporary disk demand. Fix the cause first.
If a rolled-over read-only index has a lifecycle reason for force merge, schedule and observe that operation separately, as demonstrated in Lesson 3. Keep the hot write index on normal background merge policy unless benchmark evidence and vendor guidance justify a change.
6. Flush/translog evidence prevents a second category mistake
If operators see a large translog and manually flush repeatedly, ask why. Automatic flush normally manages translog growth. A manual flush might shorten replay state but can add I/O; it does not make new documents searchable and does not remove the need to diagnose storage pressure.
Likewise, a refresh does not guarantee a Lucene commit. Keep the five distinct states from Lesson 2 in the incident runbook: acknowledged, GET-visible, search-visible, translog-persisted per policy, Lucene-committed.
7. Cold-cache and storage checks
If the incident followed node restart or shard relocation, compare cold and warmed query behavior before blaming merges. If remote storage or a degraded device shows elevated tail latency, correlate it with node filesystem and host/storage telemetry. A merge can expose storage weakness, but the root cause may be device/network latency rather than “too many segments.”
Managed-service users may not have host-level
iostat access. Use provider metrics and the
product’s node/index statistics, and document observability gaps
rather than inventing values.
8. AtlasMart storage runbook and rollback
- Freeze the evidence window: application p95/p99, throughput, freshness, errors, deployment/config changes.
- Capture index/node stats before changing anything.
- State one causal hypothesis and one disconfirming observation.
- Reproduce safely on a disposable bounded fixture when possible.
- Change one variable; preserve mapping/query/data fixture.
- Validate document counts, search fixtures, freshness and tail latency.
- Keep rollback command/config and owner.
-
Delete only disposable lab indices; never manually remove
files under
path.data.
Production judgment
Storage incidents cross layers: application refresh behavior, Lucene segment lifecycle, translog/flush, OS cache, device latency, merge/recovery and query workload. Good diagnosis narrows those layers with evidence. The next chapter deliberately continues upward into JVM heap, caches, circuit breakers, thread pools and backpressure so storage and runtime saturation are not conflated.
Check your understanding
- Why is segment count alone insufficient for diagnosis?
- What is a strong refresh-pressure hypothesis?
- Why not force merge the hot index as first repair?
- How do you distinguish cold cache from persistent storage/query problems?
- What must every tuning experiment preserve?
Review the answers
1. It lacks workload, segment-size, refresh, merge, disk and latency context; many segments can be normal or pathological depending on evidence.
2. A documented increase in forced refreshes aligns with more/smaller segments, more merge work/storage pressure and degraded foreground latency under the same workload.
3. It treats a symptom, adds heavy I/O/temporary disk demand, and new writes can immediately recreate small segments.
4. Run controlled repeated query sets and correlate latency with device reads/page-cache warmup and server metrics.
5. Data/query correctness, freshness requirements, tail-latency/error objectives, recovery safety and a reversible change record.
Summary and next step
You can now connect application-visible latency to segment creation, merge debt, deleted docs, translog/flush behavior, filesystem cache and storage I/O without mixing the mechanisms. Chapter 15 adds JVM heap, caches, circuit breakers, thread pools, queues and backpressure to the same evidence-first diagnostic method.
Authoritative references
- Elastic index segments API — Low-level Lucene segment metadata for index shards.
- Elastic index stats API — Refresh, flush, merge, segment, translog, docs and store statistics.
- Elastic translog settings — Flush as Lucene commit plus new translog generation and request durability semantics.
- Elastic tune for indexing speed — Filesystem-cache and indexing guidance.
- Elastic tune for search speed — Filesystem cache, storage latency and local-versus-remote storage guidance.
- OpenSearch Index Segments API — Segment committed/search flags and Lucene writer-version evidence.
- OpenSearch Index Stats API — Refresh, flush, merge, segments, translog and deleted-document statistics.
- OpenSearch Refresh API — Search-visibility semantics and refresh cost guidance.
- OpenSearch Flush API — Flush and transaction-log lifecycle semantics.
- OpenSearch Force Merge API — Merge/deleted-document behavior and temporary disk-space risk.
- Elastic CAT segments — Human-readable segment inspection; application code should use structured APIs.
- OpenSearch CAT segments — Human-readable segment-level state.
- Elastic hot threads API — CPU/blocking evidence during diagnosis.
- OpenSearch hot threads API — Node hot-thread evidence for runtime/storage correlation.