Chapter 04 · Indexing and CRUD: Bulk APIs, Refresh, Concurrency Control, and Idempotent Ingestion
Refresh vs Flush vs Force Merge: Search Visibility, Durability, Segments, and Performance
Distinguish search visibility from acknowledged writes, translog-backed recovery, Lucene commits, and segment merging so refresh, flush, and force merge are never used as interchangeable durability knobs.
Learning outcomes
AtlasMart operators see three tempting commands—refresh, flush, and force merge—and each seems to “make indexing more complete.” They solve different problems. Refresh changes what search can see. Flush establishes a Lucene commit and allows older translog generations to be discarded according to the engine’s recovery rules. Segment merge rewrites Lucene segments and can reclaim deletes. Confusing them can produce tiny segments, heavy I/O, long pauses, or false claims about durability.
Distinguish acknowledged write, real-time GET visibility, near-real-time search visibility, translog recovery, Lucene commit, and segment merge.
Use request refresh=false,
refresh=wait_for, and
refresh=true intentionally and explain their
cost.
Explain what a flush changes and why manual flush is rarely a per-write application requirement.
Explain why force merge is normally a maintenance action for indices that have stopped receiving writes, not a durability mechanism.
Observe document visibility, refresh/flush counters, translog stats, and segment counts without inventing performance numbers.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. The portable core uses document
and index APIs verified on both products; product-specific
behavior is labeled rather than normalized away. Elasticsearch
examples assume the default self-managed distribution with its
bundled JVM. OpenSearch examples assume the upstream 3.8.0
distribution with the Security plugin present. Keep the
earlier course endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, and OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD.
The generation environment does not run these two search
servers. Requests were checked against current official
documentation but were not executed here, so example responses
describe expected fields, status classes, and invariants
rather than fabricated captured output or benchmarks. Use only
disposable atlasmart-* indices, preserve the
CA/certificate paths established in Chapter 01, and never
point cleanup, force-merge, or failure-injection commands at
unrelated or production data.
1. Acknowledged does not mean searchable yet
Elasticsearch and OpenSearch are near-real-time search systems. A document mutation can be acknowledged before a periodic refresh opens a new searcher over recently written Lucene segment data. A direct GET by ID is designed for current document retrieval and can observe state before a search query sees the same mutation. That difference is exactly what the lab should prove; it is not evidence that the write was “lost.”
PUT atlasmart-products-v4-write-lab/_doc/P-720?refresh=false{"product_id":"P-720","name":"Laptop Stand","price":44.00,"stock":7,"updated_at":"2026-09-11T05:40:00Z","event_id":"evt-720"}GET atlasmart-products-v4-write-lab/_doc/P-720GET atlasmart-products-v4-write-lab/_search{"query":{"term":{"product_id":"P-720"}}}POST atlasmart-products-v4-write-lab/_refreshGET atlasmart-products-v4-write-lab/_search{"query":{"term":{"product_id":"P-720"}}}
The first search may already see the document if an automatic
refresh happened before you query. That race is expected. To
demonstrate the distinction reliably, use a long disposable
refresh_interval and record timing rather than
asserting that search must miss the document every time.
2. Request-level refresh choices
| Value | Behavior | Production consequence |
|---|---|---|
false (normal default) |
Do not force a refresh because of this request; periodic/background policy determines visibility. | Highest opportunity for efficient batching of segment creation. |
wait_for |
Wait until a refresh makes the change visible; do not itself force an immediate refresh. | Adds response latency but normally avoids the tiny-segment pattern of refresh=true. |
true |
Refresh affected primary/replica shards immediately after the operation. | Can create many small segments and increase later merge/search cost when used frequently. |
For Bulk requests, only shards that received operations participate in request-level refresh behavior. That detail matters when measuring latency and when a test incorrectly assumes every shard in the index has been refreshed by one small batch.
refresh=true after every
write.
This converts an application freshness requirement into
continuous searcher/segment churn. Prefer the default refresh
cadence or wait_for for operations that truly
must return only after search visibility. If the business
requires sub-second read-your-write behavior, first ask
whether direct GET/state retrieval or a different architecture
is the correct contract.
3. Flush is about recovery boundaries, not search visibility
Each engine uses a transaction log/translog so recently acknowledged operations can be recovered even when they are not yet represented in a durable Lucene commit. A flush creates a Lucene commit and starts a new translog generation, allowing older translog data to become unnecessary for recovery according to engine rules. Both products perform flushes automatically. Manual flush exists for operational cases; it is not something an application should issue after every mutation.
GET atlasmart-products-v4-write-lab/_stats?filter_path=indices.*.total.translog,indices.*.total.refresh,indices.*.total.flush,indices.*.primaries.segmentsGET atlasmart-products-v4-write-lab/_segmentsPOST atlasmart-products-v4-write-lab/_flushGET atlasmart-products-v4-write-lab/_stats?filter_path=indices.*.total.translog,indices.*.total.refresh,indices.*.total.flush,indices.*.primaries.segments
A manual flush is not “make my latest write durable.” Durability of an acknowledged write depends on translog durability configuration, primary/replica acknowledgement conditions, storage behavior, and the exact operation. Nor is a flush a backup: it does not create an independent recovery copy outside the cluster.
4. Force merge rewrites segments; it does not certify safety
Lucene segments are immutable. Background merge policy combines smaller segments and eventually reclaims space from deleted documents. Force merge asks the engine to perform additional merge work. Official guidance warns against routine force merge on actively written indices because it can create very large segments, consume disk and I/O, and interfere with normal merge behavior. Later lifecycle chapters can apply force merge after rollover when a backing index becomes read-only.
# First stop writes to this disposable index and verify that assumption.POST atlasmart-products-v4-write-lab/_forcemerge?max_num_segments=1GET atlasmart-products-v4-write-lab/_segments
It is a costly segment-maintenance operation. If the goal is search visibility, refresh is the relevant mechanism. If the goal is recovery/commit housekeeping, flush is the relevant mechanism. If the goal is independent disaster recovery, snapshots are the relevant family and are taught later.
5. Evidence lab: build a state-transition timeline
Create ten small documents with refresh=false,
capture the write acknowledgement timestamps, GET/search
visibility, stats, and segment inventory. Then perform one
refresh, one flush, and—only after stopping writes—an optional
force merge. Record what changed after each step. Do not claim
that any operation made p99 latency “better”; the lab is about
causal state evidence, not invented performance.
| Step | Evidence to capture | What it can prove |
|---|---|---|
| Write acknowledged | Response metadata + timestamp | The mutation was accepted under the request’s shard/durability conditions. |
| Direct GET | Found/source + sequence metadata | Current ID-level document state visible to GET. |
| Search | Hit count/source | Search reader currently exposes the mutation. |
| Refresh | Refresh counters + later search | A new searcher/visible segment state was opened for affected shards. |
| Flush | Flush/translog stats | Recovery boundary/translog generation housekeeping changed. |
| Force merge | Segment inventory/disk activity | Segments were merged; not that backup/durability requirements are satisfied. |
Check your understanding
- What is the primary purpose of refresh?
-
Why is
refresh=wait_fordifferent fromrefresh=true? - Does flush make a document searchable?
- When is force merge most defensible?
- Is replication plus flush a backup?
Review the answers
1. Make recent index changes visible to search by opening a new searcher/segment view.
2. wait_for waits for a refresh event; true forces an immediate refresh of affected shards.
3. Search visibility is a refresh concern. Flush is a recovery/Lucene-commit/translog concern.
4. On an index that has stopped receiving writes, commonly as lifecycle maintenance, with disk/I/O headroom and a measured reason.
5. No. Replicas and local commits are cluster availability/recovery mechanisms; independent snapshot/restore is a separate disaster-recovery contract.
Production judgment
Freshness is a product SLO and must be measured separately from write acknowledgement and disaster recovery. Tight refresh intervals improve visibility but increase segment churn; long intervals improve write efficiency but increase search lag. Manual flush and force merge are operational controls, not application acknowledgements. Capacity plans must leave disk and I/O headroom for translogs and merges, and benchmark reports must state refresh policy and segment state because they materially affect indexing/search results.
Summary and next step
You can now name which mechanism controls visibility, recovery boundaries, and segment rewriting. The next lesson uses sequence numbers and primary terms to prevent a stale client from overwriting a newer document, then connects that OCC contract to external versioning and idempotency keys.
Authoritative references
- Elastic Bulk API — Current bulk action, per-item result, refresh, versioning, routing, and OCC reference.
- Elastic refresh parameter — Current visibility semantics for index/update/delete/bulk requests.
- Elastic optimistic concurrency control — Sequence-number and primary-term concurrency semantics.
- OpenSearch Bulk API — Current NDJSON, per-item failure, OCC, versioning, and refresh behavior.
- OpenSearch Document APIs — Current document operation and sequence-number/primary-term overview.
- OpenSearch Refresh API — Near-real-time visibility and refresh guidance.
- OpenSearch Flush API — Flush and transaction-log recovery semantics.
- OpenSearch Force Merge API — Segment merge behavior and write-active warnings.
- Elastic Flush API — Current Elasticsearch flush API.
- Elastic Force Merge API — Current Elasticsearch force-merge API.