Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals

Immutable Segments, Inverted Index/Postings, Stored Fields, Doc Values, Norms, and Term Dictionaries

Connect AtlasMart search behavior to the immutable Lucene structures inside each shard and learn which structures serve matching, retrieval, sorting/aggregation, and scoring.

Intermediate115–150 minutesLucene storage internals & evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart search works because every Elasticsearch/OpenSearch shard is backed by a Lucene index. A Lucene index is not one mutable table file: it is a collection of immutable segments. Each segment contains purpose-built structures for matching, scoring, sorting, aggregating and retrieving fields. Understanding those structures prevents incorrect claims such as “_source is the inverted index” or “doc values are just another cache.”

01

Explain immutable Lucene segments and why updates create new document versions rather than editing a row in place.

02

Distinguish term dictionaries and postings from stored fields, doc values and norms.

03

Connect text analysis from Chapter 05 to the tokens that appear in postings.

04

Inspect segment metadata and term vectors without treating internal file layouts as stable application APIs.

05

Choose mapping features from query/retrieval requirements and predict their storage/search consequences.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual runtime with GET _nodes/jvm instead of hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly says otherwise. No moving latest tags, no manual editing of Lucene files, and no production force merge are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. One shard is one Lucene index composed of segments

A segment is an immutable unit written by Lucene. New indexing work accumulates in memory and, after refresh, becomes visible through newly opened segment readers. Existing segment bytes are not edited to change a document. An update is logically a delete of the old Lucene document plus indexing of a new one; physical reclamation waits for merge work.

Immutability simplifies concurrent search because readers can continue using old segment files while new readers open newer segments. The cost is write amplification: updates, deletes and merges eventually rewrite data.

Dev Tools · inspect AtlasMart shard/segment evidence
GET _cat/shards/atlasmart-products*?v&h=index,shard,prirep,state,docs,store,node
GET atlasmart-products*/_segments
GET atlasmart-products*/_stats/docs,store,segments?include_segment_file_sizes=true
GET _nodes/jvm
Implementation boundary

Segment names, file extensions and exact Lucene codec details are implementation evidence, not an application contract. Use supported Elasticsearch/OpenSearch APIs for operations; never modify files under path.data.

2. Inverted index = term dictionary + postings, not the original JSON

For an analyzed text field, Chapter 05’s analyzer produces tokens. Lucene’s term dictionary organizes distinct indexed terms so a query can locate a term efficiently. A postings list records which Lucene document IDs contain that term and may also retain frequencies, positions and offsets depending on mapping/index options.

A match query analyzes user text and looks up resulting terms in these structures. A phrase query needs positional information. This is why changing an analyzer changes the indexed token contract and normally requires reindexing: the old postings do not magically change when a new analyzer definition exists.

Smallest observable fixture
PUT atlasmart-storage-lab
{
  "settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"-1"},
  "mappings":{"properties":{
    "sku":{"type":"keyword"},
    "name":{"type":"text","fields":{"raw":{"type":"keyword"}}},
    "price":{"type":"scaled_float","scaling_factor":100},
    "popularity":{"type":"integer"}
  }}
}

PUT atlasmart-storage-lab/_doc/P-1401
{"sku":"P-1401","name":"Café USB-C Dock","price":89.00,"popularity":42}

POST atlasmart-storage-lab/_refresh
GET atlasmart-storage-lab/_termvectors/P-1401?fields=name&positions=true&offsets=true&term_statistics=true

Term vectors are optional evidence derived from indexed terms and settings; they are not the global term dictionary itself. Use them to inspect token/position behavior for a document, then use _analyze when debugging the analyzer pipeline.

3. Stored fields, _source and doc values answer different questions

Structure Primary job Common AtlasMart use Key boundary
_source Retain the original/reconstructed JSON representation exposed by Elasticsearch/OpenSearch Return product JSON; reindex/update source Not the inverted index and not automatically optimized for sorting.
Stored fields Store explicitly selected field values in Lucene for retrieval Special retrieval cases Separate from _source; rarely a reason to duplicate everything.
Doc values Column-oriented on-disk values optimized for per-document access by sort/aggregation/scripts Sort by price; aggregate category/price Mapping/type dependent; not “all fields in heap.”
Postings Term → document occurrence data Full-text and exact term matching Optimized for term-driven retrieval, not arbitrary source reconstruction.
Norms Compact per-document scoring normalization information BM25 field-length effects Relevant to scoring fields; disabling them changes scoring capability/space tradeoffs.

4. Multi-fields deliberately create multiple index structures

The AtlasMart name field is indexed as analyzed text and also as name.raw keyword. That is not one representation with two query modes; it is two field mappings with different indexed structures. The text field supports lexical matching through analyzed terms, while the keyword subfield supports exact value semantics and typically doc values for sorting/aggregation.

This is a storage-for-capability tradeoff. Multi-fields, positions, norms and doc values should exist because a query or aggregation requires them, not because “more indexing is safer.”

Inspect mapping and capabilities
GET atlasmart-storage-lab/_mapping
GET atlasmart-storage-lab/_field_caps?fields=name,name.raw,price,popularity
GET atlasmart-storage-lab/_search
{
  "query":{"match":{"name":"cafe dock"}},
  "sort":[{"price":"asc"}]
}

5. Boundary case: do not infer physical bytes from logical field names

Compression, codec selection, compound files and Lucene versions can change the physical representation. A field that looks small in JSON can create more index bytes because of terms/positions/doc values; conversely, repeated values may compress efficiently. Measure store, segment file sizes and actual workload behavior.

Elasticsearch 9.5.3 and OpenSearch 3.8.0 both expose segment version metadata. Use it as diagnostic evidence, not as a promise that every segment in a rolling-upgraded cluster was written by one identical Lucene patch.

6. AtlasMart lab and production judgment

Index a deterministic 1,000-document fixture twice: once with only fields required by the query contract and once with an intentionally redundant multi-field or unnecessary scoring feature. Refresh both, capture mapping, segment, store and query behavior, and explain the size/capability difference. Do not fabricate a “% faster” result: record your machine, warm/cold state, corpus and measured values.

Production decisions should link every mapping/storage feature to a user-visible capability. Segment internals explain cost, but application code should depend on supported mapping/query semantics rather than codec/file layouts.

  • Capture _segments and _stats/store,segments after the same document count.
  • Verify matching, sorting and aggregation behavior before comparing size.
  • Keep the fixture and commands in source control so mapping changes are regression-testable.
  • Cleanup: delete only atlasmart-storage-lab* disposable indices.

Check your understanding

  1. Why are Lucene segments immutable?
  2. What do postings answer?
  3. Why are doc values not the same as _source?
  4. What do norms contribute?
  5. Should application code depend on Lucene segment file names?
Review the answers

1. Readers can safely search stable files while new segments are created; updates/deletes are represented through new data and deletion markers until merges reclaim old bytes.

2. They map indexed terms to Lucene document occurrences, optionally including frequencies, positions and offsets.

3. Doc values are column-oriented field data for operations such as sorting/aggregation, while _source represents the document JSON used for retrieval and reindex/update workflows.

4. Compact per-document scoring normalization data, such as field-length effects used by relevance scoring.

5. No. Treat them as implementation details and use supported Elasticsearch/OpenSearch APIs.

Summary and next step

You can now name the structures inside a shard without conflating them. Lesson 2 follows one acknowledged write through in-memory indexing, real-time GET, refresh, translog durability and flush/Lucene commit so visibility and durability are separated precisely.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.