Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals
Refresh Creates Searchable Segments; Flush Commits; Translog Supports Recovery—Separate the Concepts
Separate search visibility, request acknowledgement, crash recovery, Lucene commit, and translog lifecycle so refresh and flush are never used as interchangeable controls.
Learning outcomes
AtlasMart writes a new product and receives an HTTP success, but
an immediate search may not find it. An operator suggests
_flush; another suggests refresh=true.
Both commands touch index lifecycle, but they solve different
problems. This lesson separates
acknowledgement,
real-time GET,
search visibility,
translog durability, and
Lucene commit.
Explain why a successful write can be retrievable by GET before it appears in search.
Define refresh as search visibility through new/opened Lucene segments rather than durability.
Define flush as a Lucene commit plus a new translog generation rather than search refresh.
Explain the translog role in acknowledged-write recovery and inspect its statistics.
Choose refresh/wait-for behavior from freshness requirements without forcing a refresh per write.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual runtime with
GET _nodes/jvm instead of hard-coding a JDK
patch. Labs use one primary and zero replicas unless a step
explicitly says otherwise. No moving latest tags,
no manual editing of Lucene files, and no production force
merge are used.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Follow one write through five different milestones
| Milestone | What it means | What it does NOT prove |
|---|---|---|
| Write acknowledged | The request met the product’s configured acknowledgement/durability requirements | Not necessarily visible to ordinary search yet. |
| Real-time GET succeeds | The primary can retrieve the latest document using real-time mechanisms | Does not mean a search reader includes the document. |
| Refresh occurs | A search reader can see newly refreshed segments | Not a Lucene commit/durability substitute. |
| Translog fsync/commit per configured durability | Recent operations are persisted according to translog durability semantics | Not equivalent to segment merge or force merge. |
| Flush occurs | Lucene commit is performed and a new translog generation begins | Not required after every write and not primarily a freshness control. |
2. Reproduce GET-versus-search visibility
DELETE atlasmart-refresh-lab
PUT atlasmart-refresh-lab
{
"settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"-1"},
"mappings":{"properties":{"sku":{"type":"keyword"},"name":{"type":"text"}}}
}
PUT atlasmart-refresh-lab/_doc/P-1420
{"sku":"P-1420","name":"AtlasMart Portable SSD"}
GET atlasmart-refresh-lab/_doc/P-1420
GET atlasmart-refresh-lab/_search
{"query":{"term":{"sku":"P-1420"}}}
POST atlasmart-refresh-lab/_refresh
GET atlasmart-refresh-lab/_search
{"query":{"term":{"sku":"P-1420"}}}
With automatic refresh disabled, the expected invariant is that real-time GET can retrieve the acknowledged document while ordinary search does not return it until refresh. Exact response metadata varies by product/version; verify the document identity and hit count rather than memorizing incidental fields.
This demonstrates search visibility—not crash recovery. Do not conclude that “refresh persisted the document safely.”
3. Refresh creates/open search-visible segment state
A refresh publishes recent indexing work to a new searcher so queries can see it. Frequent refreshes can create many small segments, which increases later merge work and can reduce indexing efficiency. The refresh interval is therefore a freshness-versus-resource decision.
For request-level read-after-write search semantics,
refresh=wait_for is often safer than
refresh=true: it waits for a refresh rather than
forcing one immediately. Whether it is appropriate depends on
latency and throughput objectives.
GET atlasmart-refresh-lab/_stats/refresh,segments,docs
GET atlasmart-refresh-lab/_segments
PUT atlasmart-refresh-lab/_doc/P-1421?refresh=wait_for
{"sku":"P-1421","name":"AtlasMart NVMe Enclosure"}
A forced refresh per document can create inefficient small-segment patterns. Repair the application contract first: decide whether the caller needs real-time GET, eventual search visibility, or explicit wait-for-search visibility.
4. The translog bridges acknowledgement and Lucene commit
Index operations are recorded in a transaction log
(translog) so operations newer than the latest
Lucene commit can be replayed during shard recovery.
Elasticsearch currently documents
index.translog.durability=request as the default,
where success is reported after required translog persistence on
the primary and allocated replicas. OpenSearch provides
corresponding translog durability settings; inspect the
effective value rather than assuming an environment override did
not change it.
A larger translog can reduce flush frequency but lengthen recovery replay; a smaller one can increase flush activity. This is an operational tradeoff, not a “make it huge for speed” knob.
GET atlasmart-refresh-lab/_settings?include_defaults=true&flat_settings=true&filter_path=*.settings.index.translog.*,*.defaults.index.translog.*
GET atlasmart-refresh-lab/_stats/translog,flush
GET _nodes/jvm
5. Flush = Lucene commit + new translog generation
A flush performs a Lucene commit and starts a new translog generation. Both platforms normally flush automatically. Manual flush is a specialized operational action, not part of an application write loop.
Flush may coincide with other activity, so do not use “search found my document after flush” as proof that flush caused visibility—the document may already have refreshed. To test concepts, disable automatic refresh as above and invoke refresh/flush separately.
GET atlasmart-refresh-lab/_stats/flush,translog,segments
POST atlasmart-refresh-lab/_flush
GET atlasmart-refresh-lab/_stats/flush,translog,segments
GET atlasmart-refresh-lab/_segments
Flush commits the current Lucene index state and rolls the translog generation. It does not mean “merge down to one segment,” and it should not be used as a substitute for snapshot/backup policy.
6. AtlasMart acceptance test: visibility and durability are separate assertions
1. Create index with refresh_interval=-1.
2. Index stable ID P-1422 and record acknowledgement metadata.
3. GET P-1422 -> must exist.
4. Search sku:P-1422 -> must have 0 hits before explicit refresh.
5. POST _refresh.
6. Search -> must have 1 hit.
7. Capture _stats/translog,flush and _segments.
8. POST _flush; capture stats again.
9. Do not claim crash durability from search visibility alone.
10. Delete the disposable lab index.
A production test for actual crash recovery requires controlled process/node failure and replica/snapshot planning; this lesson does not fake a crash result. The deterministic local lab proves the separation of APIs and observable states.
Production judgment
Set freshness from product requirements and then measure indexing throughput, p95/p99 search freshness and merge pressure. Let automatic flush manage normal operation unless a documented maintenance/recovery case requires manual action. Treat translog durability changes as data-loss-risk decisions requiring explicit approval, not performance folklore.
Check your understanding
- Does refresh make an acknowledged write durable?
- What does flush do?
- Why can GET see a document before search?
- Why is refresh=true risky at high write rates?
- What should be measured before changing translog settings?
Review the answers
1. No. Refresh makes recent indexing changes visible to search readers; durability is governed by the write/translog/replication path.
2. It performs a Lucene commit and starts a new translog generation.
3. Real-time GET can retrieve the latest document without waiting for the search reader to refresh.
4. It forces immediate refreshes, potentially creating many small segments and extra merge/search overhead.
5. Flush/translog size, recovery objectives, indexing latency/throughput, failure tolerance and the effective durability configuration.
Summary and next step
Search visibility, Lucene commit and recovery logging are now separate mechanisms. Lesson 3 explains what happens after many immutable segments, updates and deletes accumulate: background merging, reclaimable deleted docs, write amplification and the narrow safe role of force merge.
Authoritative references
- Elastic index segments API — Low-level Lucene segment metadata for index shards.
- Elastic index stats API — Refresh, flush, merge, segment, translog, docs and store statistics.
- Elastic translog settings — Flush as Lucene commit plus new translog generation and request durability semantics.
- Elastic tune for indexing speed — Filesystem-cache and indexing guidance.
- Elastic tune for search speed — Filesystem cache, storage latency and local-versus-remote storage guidance.
- OpenSearch Index Segments API — Segment committed/search flags and Lucene writer-version evidence.
- OpenSearch Index Stats API — Refresh, flush, merge, segments, translog and deleted-document statistics.
- OpenSearch Refresh API — Search-visibility semantics and refresh cost guidance.
- OpenSearch Flush API — Flush and transaction-log lifecycle semantics.
- OpenSearch Force Merge API — Merge/deleted-document behavior and temporary disk-space risk.
- Elastic refresh API — Explicit search visibility refresh.
- Elastic flush API — Manual flush semantics and parameters.
- OpenSearch transaction log settings — Current index/translog settings.