Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals

Compression, Filesystem Cache, mmap, Storage Latency, and Why Local SSD Often Matters

Trace search and merge I/O through compressed Lucene files, memory mapping and the operating-system filesystem cache, then reason about storage latency without pretending heap is an index cache.

Intermediate115–150 minutesLucene storage internals & evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

An AtlasMart query is fast after repeated runs but slow immediately after a node restart. Increasing JVM heap seems tempting, yet the index files live mostly outside the Java heap and Elasticsearch/OpenSearch rely heavily on the operating system’s filesystem page cache. This lesson connects Lucene compression, memory mapping, page faults and storage latency to that cold-versus-warm behavior.

01

Separate Java heap from filesystem cache and explain why the entire index should not be copied into heap.

02

Explain mmap/hybrid file access as an OS virtual-memory mechanism rather than a second Elasticsearch cache.

03

Relate Lucene compression and doc-values/postings access patterns to disk and cache behavior.

04

Design cold-cache versus warm-cache benchmarks without fabricating hardware-independent latency claims.

05

Evaluate local SSD versus remote storage from measured I/O latency, recovery and operational requirements.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. The containers use their bundled JVMs; record the actual runtime with GET _nodes/jvm instead of hard-coding a JDK patch. Labs use one primary and zero replicas unless a step explicitly says otherwise. No moving latest tags, no manual editing of Lucene files, and no production force merge are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Heap manages JVM objects; filesystem cache holds hot file pages

Elasticsearch/OpenSearch need heap for cluster/index metadata, request execution, buffers, caches and many runtime objects. Lucene index files, however, are normally read through filesystem-backed I/O and benefit from the host operating system’s page cache. Allocating every byte of RAM to the JVM can starve that cache and make searches more I/O-bound.

Elastic’s current performance guidance explicitly emphasizes leaving substantial memory to filesystem cache; OpenSearch similarly warns against swapping and documents heap/system-memory tradeoffs. Treat those as starting guidance, then verify with workload measurements and container/host memory limits.

Inspect JVM and filesystem-facing evidence
GET _nodes/stats/jvm,fs,process,indices
GET _nodes/jvm
GET atlasmart-products*/_stats/store,segments,search

# On a Linux host (outside the container where appropriate):
# free -h
# vmstat 1
# iostat -x 1
# lsblk -o NAME,RA,MOUNTPOINT,TYPE,SIZE

2. mmap maps files into virtual address space; it does not preload everything into RAM

Memory mapping (mmap) lets a process address file regions through virtual memory. The operating system loads pages as they are accessed and can evict them under memory pressure. OpenSearch exposes node.store.allow_mmap; Elasticsearch/Lucene store implementations likewise use memory mapping where supported. The exact store implementation is platform/configuration sensitive.

A huge virtual address mapping is not equal to the same amount of resident physical RAM. Use OS metrics to distinguish virtual mappings, resident memory and actual page-cache pressure.

Do not “fix” mmap by copying index files into heap

Heap and filesystem cache solve different problems. Oversizing heap can worsen garbage collection and reduce file-cache memory. Chapter 15 covers heap/GC/circuit-breaker diagnosis in depth.

3. Compression trades CPU for fewer storage bytes and I/O

Lucene codecs compress stored structures and may use specialized encodings for postings/doc values. Compression can reduce disk footprint and I/O, but the CPU/latency outcome depends on data shape and workload. Elasticsearch/OpenSearch expose mapping/index modes and codec-related settings selectively; do not assume a low-level Lucene codec knob is safe or supported just because Lucene implements it.

Measure end-to-end: store bytes, CPU, disk read/write, p95/p99 query/indexing latency and merge throughput. Smaller store size is not automatically better if CPU becomes the bottleneck—or vice versa.

Condition Likely bottleneck clue What to verify
High disk await + low CPU Storage latency / cold cache Page-cache warmness, random I/O, device/network latency.
High CPU + low disk wait Query/aggregation/decompression/scoring work Profile/hot threads; query shape; compression tradeoff.
Fast second run, slow first run Filesystem-cache warmup likely contributes Repeat controlled cold/warm runs; do not claim causality from one query.
Merge causes search p99 spike Shared disk/cache contention Merge bytes/current merges, device utilization, cache churn.
Remote volume jitter Network/storage service tail latency Block/storage telemetry plus node/search latency correlation.

4. Local SSD often wins on latency, but topology and durability still matter

Elastic’s current guidance notes that directly attached local storage generally performs better than remote storage because it avoids communications overhead and complex remote I/O behavior; SSDs tend to beat spinning disks for mixed random/sequential search access. OpenSearch performance guidance likewise emphasizes fast storage and avoiding swap.

“Use local SSD” is still not a universal architecture answer. Ephemeral local disks require replica/snapshot/recovery design. Managed services may provide remote-backed or decoupled storage with different cache layers and guarantees. Benchmark the actual deployment flavor rather than transplanting self-managed assumptions.

Storage question Evidence
Can foreground search meet p99 under merge/recovery? Mixed workload benchmark plus device latency/utilization.
Can a node be rebuilt within RTO? Shard bytes, recovery bandwidth, remote/local source behavior.
What happens after restart? Cold-cache p95/p99 and warmup duration.
What protects data if local media fails? Replica failure domains + snapshots; not RAID folklore alone.
Can the platform tune readahead/mmap? Self-managed host access vs managed-service restrictions.

5. Benchmark cold and warm states honestly

A benchmark that reports only the fastest warmed run hides an important operational state. Conversely, dropping host caches on a shared production system is unsafe. Use a disposable lab host/container, restart the node or use a controlled dataset larger than available cache, and disclose exactly how “cold” was approximated.

Record at least p50/p95/p99, CPU, device read latency, bytes read, page faults where available, segment/store bytes and merge activity. Keep query corpus/order fixed between candidates.

Benchmark record template
Environment:
  server/version: Elasticsearch 9.5.3 OR OpenSearch 3.8.0
  host RAM / JVM heap: <record>
  storage: <device or managed profile>
  corpus/index bytes: <record>
  shard/segment count: <record>

Cold approximation: <node restart / controlled method>
Warm protocol: <N identical query-set passes>
Queries: <fixed file + seed>
Concurrency: <fixed>

Capture:
  p50/p95/p99 latency
  CPU / GC
  disk read/write + await
  node fs stats
  segment/store stats
  errors/timeouts

Never report undocumented "cache cleared" claims.

6. Preloading is an expert tradeoff, not default tuning

Both ecosystems expose mechanisms/settings related to preloading selected index file types into filesystem cache. Preloading too much can evict more useful pages and slow the system. Use it only after a repeatable cold-start problem is demonstrated and the hot working set fits memory.

The safer default is to preserve memory for OS cache, use appropriate storage, and let measured access patterns warm the working set naturally.

Production judgment

Storage design is a latency distribution and recovery design. Keep heap, filesystem cache, device latency, merge/recovery traffic and workload concurrency in one model. Do not compare local SSD and remote storage using only average query latency or a warm single-thread benchmark.

Check your understanding

  1. Why can a larger JVM heap make search worse?
  2. Does mmap mean the whole mapped file is resident in RAM?
  3. Why disclose cache warmness in benchmarks?
  4. Is local SSD always the correct architecture?
  5. What should be correlated with merge-related latency?
Review the answers

1. It can leave less RAM for the filesystem cache and may increase GC costs, while Lucene index files still depend heavily on OS-backed I/O.

2. No. It maps file regions into virtual address space; the OS brings pages into physical memory on access and can evict them.

3. Cold and warm states can have materially different storage/page-fault behavior and latency.

4. No. It often offers low latency, but durability, failure domains, recovery, managed-service behavior and workload measurements determine the design.

5. Merge bytes/current merges plus disk/device latency, CPU, cache behavior and application p95/p99.

Summary and next step

You can now reason about Lucene storage through the OS rather than treating heap as the index cache. Lesson 5 combines all chapter evidence into a diagnosis workflow for an AtlasMart indexing/storage regression.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.