Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals

Refresh Creates Searchable Segments; Flush Commits; Translog Supports Recovery—Separate the Concepts

Separate search visibility, request acknowledgement, crash recovery, Lucene commit, and translog lifecycle so refresh and flush are never used as interchangeable controls.

Intermediate115–150 minutesLucene storage internals & evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart writes a new product and receives an HTTP success, but an immediate search may not find it. An operator suggests _flush; another suggests refresh=true. Both commands touch index lifecycle, but they solve different problems. This lesson separates acknowledgement, real-time GET, search visibility, translog durability, and Lucene commit.

01

Explain why a successful write can be retrievable by GET before it appears in search.

02

Define refresh as search visibility through new/opened Lucene segments rather than durability.

03

Define flush as a Lucene commit plus a new translog generation rather than search refresh.

04

Explain the translog role in acknowledged-write recovery and inspect its statistics.

05

Choose refresh/wait-for behavior from freshness requirements without forcing a refresh per write.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual runtime with GET _nodes/jvm instead of hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly says otherwise. No moving latest tags, no manual editing of Lucene files, and no production force merge are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Follow one write through five different milestones

Milestone What it means What it does NOT prove
Write acknowledged The request met the product’s configured acknowledgement/durability requirements Not necessarily visible to ordinary search yet.
Real-time GET succeeds The primary can retrieve the latest document using real-time mechanisms Does not mean a search reader includes the document.
Refresh occurs A search reader can see newly refreshed segments Not a Lucene commit/durability substitute.
Translog fsync/commit per configured durability Recent operations are persisted according to translog durability semantics Not equivalent to segment merge or force merge.
Flush occurs Lucene commit is performed and a new translog generation begins Not required after every write and not primarily a freshness control.

2. Reproduce GET-versus-search visibility

Dev Tools · controlled visibility experiment
DELETE atlasmart-refresh-lab
PUT atlasmart-refresh-lab
{
  "settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"-1"},
  "mappings":{"properties":{"sku":{"type":"keyword"},"name":{"type":"text"}}}
}

PUT atlasmart-refresh-lab/_doc/P-1420
{"sku":"P-1420","name":"AtlasMart Portable SSD"}

GET atlasmart-refresh-lab/_doc/P-1420
GET atlasmart-refresh-lab/_search
{"query":{"term":{"sku":"P-1420"}}}

POST atlasmart-refresh-lab/_refresh
GET atlasmart-refresh-lab/_search
{"query":{"term":{"sku":"P-1420"}}}

With automatic refresh disabled, the expected invariant is that real-time GET can retrieve the acknowledged document while ordinary search does not return it until refresh. Exact response metadata varies by product/version; verify the document identity and hit count rather than memorizing incidental fields.

This demonstrates search visibility—not crash recovery. Do not conclude that “refresh persisted the document safely.”

3. Refresh creates/open search-visible segment state

A refresh publishes recent indexing work to a new searcher so queries can see it. Frequent refreshes can create many small segments, which increases later merge work and can reduce indexing efficiency. The refresh interval is therefore a freshness-versus-resource decision.

For request-level read-after-write search semantics, refresh=wait_for is often safer than refresh=true: it waits for a refresh rather than forcing one immediately. Whether it is appropriate depends on latency and throughput objectives.

Observe refresh counters and segments
GET atlasmart-refresh-lab/_stats/refresh,segments,docs
GET atlasmart-refresh-lab/_segments

PUT atlasmart-refresh-lab/_doc/P-1421?refresh=wait_for
{"sku":"P-1421","name":"AtlasMart NVMe Enclosure"}
Wrong approach: refresh after every write

A forced refresh per document can create inefficient small-segment patterns. Repair the application contract first: decide whether the caller needs real-time GET, eventual search visibility, or explicit wait-for-search visibility.

4. The translog bridges acknowledgement and Lucene commit

Index operations are recorded in a transaction log (translog) so operations newer than the latest Lucene commit can be replayed during shard recovery. Elasticsearch currently documents index.translog.durability=request as the default, where success is reported after required translog persistence on the primary and allocated replicas. OpenSearch provides corresponding translog durability settings; inspect the effective value rather than assuming an environment override did not change it.

A larger translog can reduce flush frequency but lengthen recovery replay; a smaller one can increase flush activity. This is an operational tradeoff, not a “make it huge for speed” knob.

Inspect effective durability and translog evidence
GET atlasmart-refresh-lab/_settings?include_defaults=true&flat_settings=true&filter_path=*.settings.index.translog.*,*.defaults.index.translog.*
GET atlasmart-refresh-lab/_stats/translog,flush
GET _nodes/jvm

5. Flush = Lucene commit + new translog generation

A flush performs a Lucene commit and starts a new translog generation. Both platforms normally flush automatically. Manual flush is a specialized operational action, not part of an application write loop.

Flush may coincide with other activity, so do not use “search found my document after flush” as proof that flush caused visibility—the document may already have refreshed. To test concepts, disable automatic refresh as above and invoke refresh/flush separately.

Safe disposable flush evidence
GET atlasmart-refresh-lab/_stats/flush,translog,segments
POST atlasmart-refresh-lab/_flush
GET atlasmart-refresh-lab/_stats/flush,translog,segments
GET atlasmart-refresh-lab/_segments
Flush is not force merge

Flush commits the current Lucene index state and rolls the translog generation. It does not mean “merge down to one segment,” and it should not be used as a substitute for snapshot/backup policy.

6. AtlasMart acceptance test: visibility and durability are separate assertions

Pseudo-test contract
1. Create index with refresh_interval=-1.
2. Index stable ID P-1422 and record acknowledgement metadata.
3. GET P-1422 -> must exist.
4. Search sku:P-1422 -> must have 0 hits before explicit refresh.
5. POST _refresh.
6. Search -> must have 1 hit.
7. Capture _stats/translog,flush and _segments.
8. POST _flush; capture stats again.
9. Do not claim crash durability from search visibility alone.
10. Delete the disposable lab index.

A production test for actual crash recovery requires controlled process/node failure and replica/snapshot planning; this lesson does not fake a crash result. The deterministic local lab proves the separation of APIs and observable states.

Production judgment

Set freshness from product requirements and then measure indexing throughput, p95/p99 search freshness and merge pressure. Let automatic flush manage normal operation unless a documented maintenance/recovery case requires manual action. Treat translog durability changes as data-loss-risk decisions requiring explicit approval, not performance folklore.

Check your understanding

  1. Does refresh make an acknowledged write durable?
  2. What does flush do?
  3. Why can GET see a document before search?
  4. Why is refresh=true risky at high write rates?
  5. What should be measured before changing translog settings?
Review the answers

1. No. Refresh makes recent indexing changes visible to search readers; durability is governed by the write/translog/replication path.

2. It performs a Lucene commit and starts a new translog generation.

3. Real-time GET can retrieve the latest document without waiting for the search reader to refresh.

4. It forces immediate refreshes, potentially creating many small segments and extra merge/search overhead.

5. Flush/translog size, recovery objectives, indexing latency/throughput, failure tolerance and the effective durability configuration.

Summary and next step

Search visibility, Lucene commit and recovery logging are now separate mechanisms. Lesson 3 explains what happens after many immutable segments, updates and deletes accumulate: background merging, reclaimable deleted docs, write amplification and the narrow safe role of force merge.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.