Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals
Immutable Segments, Inverted Index/Postings, Stored Fields, Doc Values, Norms, and Term Dictionaries
Connect AtlasMart search behavior to the immutable Lucene structures inside each shard and learn which structures serve matching, retrieval, sorting/aggregation, and scoring.
Learning outcomes
AtlasMart search works because every Elasticsearch/OpenSearch
shard is backed by a Lucene index. A Lucene index is not one
mutable table file: it is a collection of
immutable segments. Each segment contains
purpose-built structures for matching, scoring, sorting,
aggregating and retrieving fields. Understanding those
structures prevents incorrect claims such as “_source
is the inverted index” or “doc values are just another cache.”
Explain immutable Lucene segments and why updates create new document versions rather than editing a row in place.
Distinguish term dictionaries and postings from stored fields, doc values and norms.
Connect text analysis from Chapter 05 to the tokens that appear in postings.
Inspect segment metadata and term vectors without treating internal file layouts as stable application APIs.
Choose mapping features from query/retrieval requirements and predict their storage/search consequences.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual runtime with
GET _nodes/jvm instead of hard-coding a JDK
patch. Labs use one primary and zero replicas unless a step
explicitly says otherwise. No moving latest tags,
no manual editing of Lucene files, and no production force
merge are used.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. One shard is one Lucene index composed of segments
A segment is an immutable unit written by Lucene. New indexing work accumulates in memory and, after refresh, becomes visible through newly opened segment readers. Existing segment bytes are not edited to change a document. An update is logically a delete of the old Lucene document plus indexing of a new one; physical reclamation waits for merge work.
Immutability simplifies concurrent search because readers can continue using old segment files while new readers open newer segments. The cost is write amplification: updates, deletes and merges eventually rewrite data.
GET _cat/shards/atlasmart-products*?v&h=index,shard,prirep,state,docs,store,node
GET atlasmart-products*/_segments
GET atlasmart-products*/_stats/docs,store,segments?include_segment_file_sizes=true
GET _nodes/jvm
Segment names, file extensions and exact Lucene codec details
are implementation evidence, not an application contract. Use
supported Elasticsearch/OpenSearch APIs for operations; never
modify files under path.data.
2. Inverted index = term dictionary + postings, not the original JSON
For an analyzed text field, Chapter 05’s analyzer
produces tokens. Lucene’s
term dictionary organizes distinct indexed
terms so a query can locate a term efficiently. A
postings list records which Lucene document IDs
contain that term and may also retain frequencies, positions and
offsets depending on mapping/index options.
A match query analyzes user text and looks up
resulting terms in these structures. A phrase query needs
positional information. This is why changing an analyzer changes
the indexed token contract and normally requires reindexing: the
old postings do not magically change when a new analyzer
definition exists.
PUT atlasmart-storage-lab
{
"settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"-1"},
"mappings":{"properties":{
"sku":{"type":"keyword"},
"name":{"type":"text","fields":{"raw":{"type":"keyword"}}},
"price":{"type":"scaled_float","scaling_factor":100},
"popularity":{"type":"integer"}
}}
}
PUT atlasmart-storage-lab/_doc/P-1401
{"sku":"P-1401","name":"Café USB-C Dock","price":89.00,"popularity":42}
POST atlasmart-storage-lab/_refresh
GET atlasmart-storage-lab/_termvectors/P-1401?fields=name&positions=true&offsets=true&term_statistics=true
Term vectors are optional evidence derived from indexed terms
and settings; they are not the global term dictionary itself.
Use them to inspect token/position behavior for a document, then
use _analyze when debugging the analyzer pipeline.
3. Stored fields, _source and doc values answer different questions
| Structure | Primary job | Common AtlasMart use | Key boundary |
|---|---|---|---|
| _source | Retain the original/reconstructed JSON representation exposed by Elasticsearch/OpenSearch | Return product JSON; reindex/update source | Not the inverted index and not automatically optimized for sorting. |
| Stored fields | Store explicitly selected field values in Lucene for retrieval | Special retrieval cases |
Separate from _source; rarely a reason to
duplicate everything.
|
| Doc values | Column-oriented on-disk values optimized for per-document access by sort/aggregation/scripts | Sort by price; aggregate category/price | Mapping/type dependent; not “all fields in heap.” |
| Postings | Term → document occurrence data | Full-text and exact term matching | Optimized for term-driven retrieval, not arbitrary source reconstruction. |
| Norms | Compact per-document scoring normalization information | BM25 field-length effects | Relevant to scoring fields; disabling them changes scoring capability/space tradeoffs. |
4. Multi-fields deliberately create multiple index structures
The AtlasMart name field is indexed as analyzed
text and also as name.raw keyword.
That is not one representation with two query modes; it is two
field mappings with different indexed structures. The text field
supports lexical matching through analyzed terms, while the
keyword subfield supports exact value semantics and typically
doc values for sorting/aggregation.
This is a storage-for-capability tradeoff. Multi-fields, positions, norms and doc values should exist because a query or aggregation requires them, not because “more indexing is safer.”
GET atlasmart-storage-lab/_mapping
GET atlasmart-storage-lab/_field_caps?fields=name,name.raw,price,popularity
GET atlasmart-storage-lab/_search
{
"query":{"match":{"name":"cafe dock"}},
"sort":[{"price":"asc"}]
}
5. Boundary case: do not infer physical bytes from logical field names
Compression, codec selection, compound files and Lucene versions
can change the physical representation. A field that looks small
in JSON can create more index bytes because of
terms/positions/doc values; conversely, repeated values may
compress efficiently. Measure store, segment file
sizes and actual workload behavior.
Elasticsearch 9.5.3 and OpenSearch 3.8.0 both expose segment version metadata. Use it as diagnostic evidence, not as a promise that every segment in a rolling-upgraded cluster was written by one identical Lucene patch.
6. AtlasMart lab and production judgment
Index a deterministic 1,000-document fixture twice: once with only fields required by the query contract and once with an intentionally redundant multi-field or unnecessary scoring feature. Refresh both, capture mapping, segment, store and query behavior, and explain the size/capability difference. Do not fabricate a “% faster” result: record your machine, warm/cold state, corpus and measured values.
Production decisions should link every mapping/storage feature to a user-visible capability. Segment internals explain cost, but application code should depend on supported mapping/query semantics rather than codec/file layouts.
-
Capture
_segmentsand_stats/store,segmentsafter the same document count. - Verify matching, sorting and aggregation behavior before comparing size.
- Keep the fixture and commands in source control so mapping changes are regression-testable.
-
Cleanup: delete only
atlasmart-storage-lab*disposable indices.
Check your understanding
- Why are Lucene segments immutable?
- What do postings answer?
- Why are doc values not the same as _source?
- What do norms contribute?
- Should application code depend on Lucene segment file names?
Review the answers
1. Readers can safely search stable files while new segments are created; updates/deletes are represented through new data and deletion markers until merges reclaim old bytes.
2. They map indexed terms to Lucene document occurrences, optionally including frequencies, positions and offsets.
3. Doc values are column-oriented field data for operations such as sorting/aggregation, while _source represents the document JSON used for retrieval and reindex/update workflows.
4. Compact per-document scoring normalization data, such as field-length effects used by relevance scoring.
5. No. Treat them as implementation details and use supported Elasticsearch/OpenSearch APIs.
Summary and next step
You can now name the structures inside a shard without conflating them. Lesson 2 follows one acknowledged write through in-memory indexing, real-time GET, refresh, translog durability and flush/Lucene commit so visibility and durability are separated precisely.
Authoritative references
- Elastic index segments API — Low-level Lucene segment metadata for index shards.
- Elastic index stats API — Refresh, flush, merge, segment, translog, docs and store statistics.
- Elastic translog settings — Flush as Lucene commit plus new translog generation and request durability semantics.
- Elastic tune for indexing speed — Filesystem-cache and indexing guidance.
- Elastic tune for search speed — Filesystem cache, storage latency and local-versus-remote storage guidance.
- OpenSearch Index Segments API — Segment committed/search flags and Lucene writer-version evidence.
- OpenSearch Index Stats API — Refresh, flush, merge, segments, translog and deleted-document statistics.
- OpenSearch Refresh API — Search-visibility semantics and refresh cost guidance.
- OpenSearch Flush API — Flush and transaction-log lifecycle semantics.
- OpenSearch Force Merge API — Merge/deleted-document behavior and temporary disk-space risk.
- Elastic term vectors API — Token/term statistics for document-level inspection.
- OpenSearch term vectors API — Document term/position/offset inspection.
- Apache Lucene core — Lower-level implementation reference; not an Elasticsearch/OpenSearch application contract.