Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals

Segment Merging, Deleted Docs, Merge Policy, Force Merge Risks, and Write Amplification

Explain why updates and deletes create reclaimable work, how background merges rewrite immutable segments, and when force merge is safe only after writes have stopped.

Intermediate115–150 minutesLucene storage internals & evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart updates prices and inventory frequently. Store bytes rise even though logical document count barely changes, and segment counts fluctuate. The reason is immutable indexing: updates create replacement documents, deletes mark old Lucene documents as deleted, and background merges later rewrite segments to consolidate live data and reclaim space.

01

Explain deletion tombstones and why delete/update does not immediately shrink segment files.

02

Relate background merge policy to segment count, deleted-doc reclamation and write amplification.

03

Read merge/deleted-document statistics without interpreting every merge as a problem.

04

Explain why force merge is dangerous on actively written indices and why it can temporarily require large extra disk space.

05

Run a force-merge experiment only on a disposable read-only copy and verify before/after evidence.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual runtime with GET _nodes/jvm instead of hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly says otherwise. No moving latest tags, no manual editing of Lucene files, and no production force merge are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Immutable segments turn updates into future merge work

When AtlasMart updates P-1430, Lucene does not overwrite bytes in the segment that contains the old version. The old Lucene document becomes deleted (logically invisible), and a new document is indexed. Deletes behave similarly: a deletion marker makes a document invisible, while physical bytes are reclaimed when a merge rewrites affected segments without the deleted documents.

This is why docs.deleted can be non-zero and why store size can lag behind logical deletion. It is also why update-heavy workloads create more write amplification than append-only workloads.

Create controlled update/delete churn
DELETE atlasmart-merge-lab
PUT atlasmart-merge-lab
{
  "settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"1s"},
  "mappings":{"properties":{"sku":{"type":"keyword"},"name":{"type":"text"},"price":{"type":"scaled_float","scaling_factor":100}}}
}

PUT atlasmart-merge-lab/_doc/P-1430
{"sku":"P-1430","name":"AtlasMart Dock","price":79.00}
POST atlasmart-merge-lab/_refresh

PUT atlasmart-merge-lab/_doc/P-1430
{"sku":"P-1430","name":"AtlasMart Dock","price":74.00}
POST atlasmart-merge-lab/_refresh

GET atlasmart-merge-lab/_stats/docs,segments,merge,store
GET atlasmart-merge-lab/_segments

2. Background merge is normal housekeeping with real I/O cost

Lucene’s merge policy selects sets of segments and writes their live documents into larger replacement segments. Once readers no longer need old files, those files can be deleted. Merging reduces segment fragmentation and reclaims deleted documents, but it reads and writes substantial bytes and consumes CPU, disk bandwidth and filesystem-cache capacity.

A high merge time counter does not automatically mean a fault—the total is cumulative. Diagnose with rates and concurrency: current merges, bytes/docs being merged, indexing/search latency, disk utilization and whether merge backlog grows under steady load.

Observe merge and deleted-document evidence
GET atlasmart-merge-lab/_stats/merge,docs,segments,store
GET _nodes/stats/indices,fs,process,jvm
GET _cat/segments/atlasmart-merge-lab?v
GET atlasmart-merge-lab/_segments

3. Write amplification is workload + mapping + merge behavior

If the application writes 1 GB of new source data, the storage device may write more than 1 GB because analysis builds index structures and merges rewrite existing segments. Updates can amplify further because old versions remain until merged. There is no universal amplification factor: measure it using host/device metrics under a fixed dataset and workload.

Refresh frequency interacts with this system. Very frequent refreshes can produce more small searchable segments that must later be merged. Chapter 04’s “refresh every write” anti-pattern therefore reappears here as storage/merge debt.

Signal Possible interpretation Evidence needed before acting
Many small segments High refresh cadence or young index Refresh rate, indexing rate, segment age/size.
High docs.deleted Update/delete churn not yet reclaimed Update/delete rate, merge progress, store trend.
Merge current/time increasing Background consolidation active Disk latency/utilization and foreground p95/p99.
Store temporarily grows during merge Old + new segment files coexist during rewrite Free-space headroom and active merge bytes.
Search slower after restart Cold filesystem cache rather than “bad merge” Cold/warm repeated query comparison, disk reads.

4. Force merge is not a maintenance ritual for hot indices

The force merge API requests aggressive merging, for example toward one segment. Both products warn that it is resource intensive and can temporarily require significant additional disk space. OpenSearch documents that a one-segment force merge can temporarily double shard storage; Elastic similarly recommends force merge only after an index has stopped receiving writes in lifecycle scenarios.

Force-merging a hot write index can create very large segments, compete with indexing/search, and future writes immediately create new segments again. It does not make writes durable, does not replace flush, and does not replace snapshots.

Wrong approach: “nightly force merge everything”

This treats segment count as the objective and ignores workload/lifecycle state. Let background merge policy operate for write-active indices. Consider force merge only after rollover/read-only transition, with measured disk and latency headroom.

5. Safe lab: clone/reindex to a disposable read-only copy, then force merge

Create disposable cold copy
PUT atlasmart-merge-coldcopy
{
  "settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"-1"},
  "mappings":{"properties":{"sku":{"type":"keyword"},"name":{"type":"text"},"price":{"type":"scaled_float","scaling_factor":100}}}
}

POST _reindex
{
  "source":{"index":"atlasmart-merge-lab"},
  "dest":{"index":"atlasmart-merge-coldcopy"}
}
POST atlasmart-merge-coldcopy/_refresh
PUT atlasmart-merge-coldcopy/_settings
{"index.blocks.write":true}

GET atlasmart-merge-coldcopy/_stats/docs,store,segments,merge
GET atlasmart-merge-coldcopy/_segments
POST atlasmart-merge-coldcopy/_forcemerge?max_num_segments=1
GET atlasmart-merge-coldcopy/_stats/docs,store,segments,merge
GET atlasmart-merge-coldcopy/_segments

Record free disk before starting. The fixture is intentionally tiny; it teaches state transitions, not production performance. On real large shards, force merge can run for a long time and consume large temporary disk capacity.

6. Merge policy is not a first-line tuning surface

Changing merge-policy settings without a workload trace can trade one bottleneck for another. Start with mapping correctness, sensible refresh/freshness requirements, bulk indexing, adequate fast storage and lifecycle design. If merge debt remains causal, benchmark one change at a time with the exact workload.

Managed services may restrict low-level settings or perform storage/lifecycle operations differently. Document the deployment flavor before copying a self-managed setting.

Production judgment

Treat merge activity as background debt service. The objective is not zero merges; it is sustainable merge throughput while foreground indexing/search meets SLOs and disk headroom remains safe. Deleted-doc percentage alone is not an incident—connect it to store growth, merge backlog and user-visible behavior.

Check your understanding

  1. Why does deleting a document not immediately free all disk bytes?
  2. Why can updates increase docs.deleted?
  3. What resource cost does merging impose?
  4. When is force merge most defensible?
  5. Does force merge improve durability?
Review the answers

1. The document is logically marked deleted in immutable segment state; physical bytes are reclaimed when merges rewrite segments without it.

2. An update creates a replacement document and marks the prior Lucene document deleted.

3. It reads and rewrites segment data, consuming disk bandwidth, CPU, temporary disk space and cache capacity.

4. After writes have stopped, such as a rolled-over/read-only lifecycle stage, with disk/latency headroom measured.

5. No. Durability is not its purpose; flush/translog/replication/snapshot mechanisms address different guarantees.

Summary and next step

You can now interpret deleted-doc and merge evidence as consequences of immutable segments rather than as mysteries. Lesson 4 moves from logical storage structures to the operating system: compression, mmap, filesystem cache, cold/warm behavior and storage latency.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.