Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals

Diagnose an Indexing/Storage Problem Using Segment Counts, Merge Pressure, Disk I/O, and Refresh Behavior

Use segment, refresh, merge, translog, node and application-latency evidence to diagnose an AtlasMart indexing/storage regression and verify a causal repair.

Intermediate115–150 minutesLucene storage internals & evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart reports a production symptom: indexing throughput fell, search p99 doubled, disk utilization rose, and operators see hundreds of segments. The wrong response is to force merge immediately. A storage diagnosis must establish a timeline and causal chain across refresh rate, segment creation, update/delete churn, merge pressure, translog/flush behavior, disk I/O and cache state.

01

Build an evidence timeline from application SLOs and index/node storage metrics before changing settings.

02

Distinguish excessive refresh/small-segment pressure from merge debt, translog/flush pressure and cold-cache I/O.

03

Use segment, docs.deleted, merge, refresh, flush, translog and filesystem statistics together.

04

Run one-variable experiments and reject repairs that only improve averages while p99 or recovery risk worsens.

05

Produce a reversible AtlasMart storage runbook with acceptance, rollback and cleanup criteria.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual runtime with GET _nodes/jvm instead of hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly says otherwise. No moving latest tags, no manual editing of Lucene files, and no production force merge are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Start from the user-visible regression, not from a favorite metric

Collect the time window in which AtlasMart search/indexing degraded. Record p50/p95/p99 latency, indexing rate, errors/rejections, freshness lag and deployment/configuration changes. Only then align server metrics. Segment count alone is not a diagnosis; a small busy index can have many harmless young segments, while a few huge segments can still be I/O-bound.

The goal is a falsifiable hypothesis such as “a refresh-per-document change created small-segment/merge pressure that saturated storage, raising write and search tail latency.”

First-pass evidence bundle
GET _cluster/health
GET _cat/indices/atlasmart-products*?v&h=health,status,index,pri,rep,docs.count,docs.deleted,store.size,pri.store.size
GET _cat/segments/atlasmart-products*?v
GET atlasmart-products*/_stats/docs,store,indexing,search,refresh,flush,merge,segments,translog?level=shards
GET _nodes/stats/fs,indices,jvm,process,thread_pool
GET _nodes/hot_threads

2. Read symptom patterns, then verify competing explanations

Observed pattern Hypothesis to test Disconfirming evidence
Refresh count/rate jumps + many small segments Forced/too-frequent refresh creates merge work Refresh unchanged; segment sizes/count stable.
docs.deleted/store grow with update rate Update/delete churn awaits merge reclamation No churn; deleted docs flat.
merge current/bytes rise with disk await and p99 Merge/storage contention Disk idle and CPU/query profile dominates instead.
Translog grows; recovery/flush behavior changes Flush threshold/durability/write pattern changed Translog/flush counters normal for same write rate.
First queries slow after restart, then improve Cold filesystem cache Repeated warm runs remain equally slow.
Search slow but merge/storage normal Query/relevance/heap/thread-pool cause outside this chapter Profile, GC/rejections/hot threads point elsewhere.

3. Build a controlled AtlasMart failure fixture

In a disposable lab only, compare two write modes against the same dataset: A uses bulk writes and a normal/explicit refresh cadence; B deliberately uses refresh=true per write for a bounded number of documents. Capture segment/refresh/merge metrics and latency. This safely demonstrates mechanism without filling disks or killing host processes.

Do not expect exact numeric outcomes across machines. The pass condition is directional evidence on your environment plus a clear causal explanation.

Bounded experiment outline
Index A: atlasmart-storage-good
  1 primary, 0 replicas
  bulk 1,000 deterministic documents
  refresh only at test boundary

Index B: atlasmart-storage-bad
  same mapping + documents
  index each document with refresh=true
  keep document count bounded

For both capture:
  _stats/refresh,merge,segments,docs,store
  _segments
  wall-clock ingest time
  fixed-query p50/p95/p99 after indexing
  host disk/CPU evidence

Cleanup both disposable indices after comparison.
Safety boundary

Do not intentionally fill the filesystem, disable safety mechanisms, or force merge a production/hot index to create a “better demo.” Controlled request patterns are enough to expose segment/refresh behavior.

4. One-variable repair: remove forced refresh and retest

If the evidence supports the refresh hypothesis, change only that variable first: remove per-write forced refresh, batch writes, and use the product requirement to choose normal periodic refresh or wait_for where a caller truly needs search visibility. Rerun the identical fixture and query set.

Verify that the repair preserves correctness: expected documents are searchable within the agreed freshness bound, no writes are lost, and application behavior that depended on immediate search is either redesigned or explicitly waits for visibility.

Acceptance record
Before -> after comparison:
  document count/checksum: must match
  search relevance fixture: must pass
  freshness SLO: must pass
  indexing docs/s: record
  p50/p95/p99 search + indexing latency: record
  refresh total/rate: record
  segment count/size distribution: record
  merge bytes/time/current: record
  disk latency/utilization: record
  translog/flush stats: record

Reject the change if correctness or tail-latency/recovery objectives regress.

5. Force merge is not the repair for a bad hot-index refresh policy

A force merge might temporarily reduce segment count, but if the application continues forcing refreshes, small segments begin accumulating again. Worse, the force merge itself adds heavy I/O and temporary disk demand. Fix the cause first.

If a rolled-over read-only index has a lifecycle reason for force merge, schedule and observe that operation separately, as demonstrated in Lesson 3. Keep the hot write index on normal background merge policy unless benchmark evidence and vendor guidance justify a change.

6. Flush/translog evidence prevents a second category mistake

If operators see a large translog and manually flush repeatedly, ask why. Automatic flush normally manages translog growth. A manual flush might shorten replay state but can add I/O; it does not make new documents searchable and does not remove the need to diagnose storage pressure.

Likewise, a refresh does not guarantee a Lucene commit. Keep the five distinct states from Lesson 2 in the incident runbook: acknowledged, GET-visible, search-visible, translog-persisted per policy, Lucene-committed.

7. Cold-cache and storage checks

If the incident followed node restart or shard relocation, compare cold and warmed query behavior before blaming merges. If remote storage or a degraded device shows elevated tail latency, correlate it with node filesystem and host/storage telemetry. A merge can expose storage weakness, but the root cause may be device/network latency rather than “too many segments.”

Managed-service users may not have host-level iostat access. Use provider metrics and the product’s node/index statistics, and document observability gaps rather than inventing values.

8. AtlasMart storage runbook and rollback

  • Freeze the evidence window: application p95/p99, throughput, freshness, errors, deployment/config changes.
  • Capture index/node stats before changing anything.
  • State one causal hypothesis and one disconfirming observation.
  • Reproduce safely on a disposable bounded fixture when possible.
  • Change one variable; preserve mapping/query/data fixture.
  • Validate document counts, search fixtures, freshness and tail latency.
  • Keep rollback command/config and owner.
  • Delete only disposable lab indices; never manually remove files under path.data.

Production judgment

Storage incidents cross layers: application refresh behavior, Lucene segment lifecycle, translog/flush, OS cache, device latency, merge/recovery and query workload. Good diagnosis narrows those layers with evidence. The next chapter deliberately continues upward into JVM heap, caches, circuit breakers, thread pools and backpressure so storage and runtime saturation are not conflated.

Check your understanding

  1. Why is segment count alone insufficient for diagnosis?
  2. What is a strong refresh-pressure hypothesis?
  3. Why not force merge the hot index as first repair?
  4. How do you distinguish cold cache from persistent storage/query problems?
  5. What must every tuning experiment preserve?
Review the answers

1. It lacks workload, segment-size, refresh, merge, disk and latency context; many segments can be normal or pathological depending on evidence.

2. A documented increase in forced refreshes aligns with more/smaller segments, more merge work/storage pressure and degraded foreground latency under the same workload.

3. It treats a symptom, adds heavy I/O/temporary disk demand, and new writes can immediately recreate small segments.

4. Run controlled repeated query sets and correlate latency with device reads/page-cache warmup and server metrics.

5. Data/query correctness, freshness requirements, tail-latency/error objectives, recovery safety and a reversible change record.

Summary and next step

You can now connect application-visible latency to segment creation, merge debt, deleted docs, translog/flush behavior, filesystem cache and storage I/O without mixing the mechanisms. Chapter 15 adds JVM heap, caches, circuit breakers, thread pools, queues and backpressure to the same evidence-first diagnostic method.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.