Chapter 14 · Lucene Segments, Refresh, Merge, Translog, Flush, and Storage Internals
Compression, Filesystem Cache, mmap, Storage Latency, and Why Local SSD Often Matters
Trace search and merge I/O through compressed Lucene files, memory mapping and the operating-system filesystem cache, then reason about storage latency without pretending heap is an index cache.
Learning outcomes
An AtlasMart query is fast after repeated runs but slow immediately after a node restart. Increasing JVM heap seems tempting, yet the index files live mostly outside the Java heap and Elasticsearch/OpenSearch rely heavily on the operating system’s filesystem page cache. This lesson connects Lucene compression, memory mapping, page faults and storage latency to that cold-versus-warm behavior.
Separate Java heap from filesystem cache and explain why the entire index should not be copied into heap.
Explain mmap/hybrid file access as an OS virtual-memory mechanism rather than a second Elasticsearch cache.
Relate Lucene compression and doc-values/postings access patterns to disk and cache behavior.
Design cold-cache versus warm-cache benchmarks without fabricating hardware-independent latency claims.
Evaluate local SSD versus remote storage from measured I/O latency, recovery and operational requirements.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. The containers use
their bundled JVMs; record the actual runtime with
GET _nodes/jvm instead of hard-coding a JDK
patch. Labs use one primary and zero replicas unless a step
explicitly says otherwise. No moving latest tags,
no manual editing of Lucene files, and no production force
merge are used.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Heap manages JVM objects; filesystem cache holds hot file pages
Elasticsearch/OpenSearch need heap for cluster/index metadata, request execution, buffers, caches and many runtime objects. Lucene index files, however, are normally read through filesystem-backed I/O and benefit from the host operating system’s page cache. Allocating every byte of RAM to the JVM can starve that cache and make searches more I/O-bound.
Elastic’s current performance guidance explicitly emphasizes leaving substantial memory to filesystem cache; OpenSearch similarly warns against swapping and documents heap/system-memory tradeoffs. Treat those as starting guidance, then verify with workload measurements and container/host memory limits.
GET _nodes/stats/jvm,fs,process,indices
GET _nodes/jvm
GET atlasmart-products*/_stats/store,segments,search
# On a Linux host (outside the container where appropriate):
# free -h
# vmstat 1
# iostat -x 1
# lsblk -o NAME,RA,MOUNTPOINT,TYPE,SIZE
2. mmap maps files into virtual address space; it does not preload everything into RAM
Memory mapping (mmap) lets a process address
file regions through virtual memory. The operating system loads
pages as they are accessed and can evict them under memory
pressure. OpenSearch exposes node.store.allow_mmap;
Elasticsearch/Lucene store implementations likewise use memory
mapping where supported. The exact store implementation is
platform/configuration sensitive.
A huge virtual address mapping is not equal to the same amount of resident physical RAM. Use OS metrics to distinguish virtual mappings, resident memory and actual page-cache pressure.
Heap and filesystem cache solve different problems. Oversizing heap can worsen garbage collection and reduce file-cache memory. Chapter 15 covers heap/GC/circuit-breaker diagnosis in depth.
3. Compression trades CPU for fewer storage bytes and I/O
Lucene codecs compress stored structures and may use specialized encodings for postings/doc values. Compression can reduce disk footprint and I/O, but the CPU/latency outcome depends on data shape and workload. Elasticsearch/OpenSearch expose mapping/index modes and codec-related settings selectively; do not assume a low-level Lucene codec knob is safe or supported just because Lucene implements it.
Measure end-to-end: store bytes, CPU, disk read/write, p95/p99 query/indexing latency and merge throughput. Smaller store size is not automatically better if CPU becomes the bottleneck—or vice versa.
| Condition | Likely bottleneck clue | What to verify |
|---|---|---|
| High disk await + low CPU | Storage latency / cold cache | Page-cache warmness, random I/O, device/network latency. |
| High CPU + low disk wait | Query/aggregation/decompression/scoring work | Profile/hot threads; query shape; compression tradeoff. |
| Fast second run, slow first run | Filesystem-cache warmup likely contributes | Repeat controlled cold/warm runs; do not claim causality from one query. |
| Merge causes search p99 spike | Shared disk/cache contention | Merge bytes/current merges, device utilization, cache churn. |
| Remote volume jitter | Network/storage service tail latency | Block/storage telemetry plus node/search latency correlation. |
4. Local SSD often wins on latency, but topology and durability still matter
Elastic’s current guidance notes that directly attached local storage generally performs better than remote storage because it avoids communications overhead and complex remote I/O behavior; SSDs tend to beat spinning disks for mixed random/sequential search access. OpenSearch performance guidance likewise emphasizes fast storage and avoiding swap.
“Use local SSD” is still not a universal architecture answer. Ephemeral local disks require replica/snapshot/recovery design. Managed services may provide remote-backed or decoupled storage with different cache layers and guarantees. Benchmark the actual deployment flavor rather than transplanting self-managed assumptions.
| Storage question | Evidence |
|---|---|
| Can foreground search meet p99 under merge/recovery? | Mixed workload benchmark plus device latency/utilization. |
| Can a node be rebuilt within RTO? | Shard bytes, recovery bandwidth, remote/local source behavior. |
| What happens after restart? | Cold-cache p95/p99 and warmup duration. |
| What protects data if local media fails? | Replica failure domains + snapshots; not RAID folklore alone. |
| Can the platform tune readahead/mmap? | Self-managed host access vs managed-service restrictions. |
5. Benchmark cold and warm states honestly
A benchmark that reports only the fastest warmed run hides an important operational state. Conversely, dropping host caches on a shared production system is unsafe. Use a disposable lab host/container, restart the node or use a controlled dataset larger than available cache, and disclose exactly how “cold” was approximated.
Record at least p50/p95/p99, CPU, device read latency, bytes read, page faults where available, segment/store bytes and merge activity. Keep query corpus/order fixed between candidates.
Environment:
server/version: Elasticsearch 9.5.3 OR OpenSearch 3.8.0
host RAM / JVM heap: <record>
storage: <device or managed profile>
corpus/index bytes: <record>
shard/segment count: <record>
Cold approximation: <node restart / controlled method>
Warm protocol: <N identical query-set passes>
Queries: <fixed file + seed>
Concurrency: <fixed>
Capture:
p50/p95/p99 latency
CPU / GC
disk read/write + await
node fs stats
segment/store stats
errors/timeouts
Never report undocumented "cache cleared" claims.
6. Preloading is an expert tradeoff, not default tuning
Both ecosystems expose mechanisms/settings related to preloading selected index file types into filesystem cache. Preloading too much can evict more useful pages and slow the system. Use it only after a repeatable cold-start problem is demonstrated and the hot working set fits memory.
The safer default is to preserve memory for OS cache, use appropriate storage, and let measured access patterns warm the working set naturally.
Production judgment
Storage design is a latency distribution and recovery design. Keep heap, filesystem cache, device latency, merge/recovery traffic and workload concurrency in one model. Do not compare local SSD and remote storage using only average query latency or a warm single-thread benchmark.
Check your understanding
- Why can a larger JVM heap make search worse?
- Does mmap mean the whole mapped file is resident in RAM?
- Why disclose cache warmness in benchmarks?
- Is local SSD always the correct architecture?
- What should be correlated with merge-related latency?
Review the answers
1. It can leave less RAM for the filesystem cache and may increase GC costs, while Lucene index files still depend heavily on OS-backed I/O.
2. No. It maps file regions into virtual address space; the OS brings pages into physical memory on access and can evict them.
3. Cold and warm states can have materially different storage/page-fault behavior and latency.
4. No. It often offers low latency, but durability, failure domains, recovery, managed-service behavior and workload measurements determine the design.
5. Merge bytes/current merges plus disk/device latency, CPU, cache behavior and application p95/p99.
Summary and next step
You can now reason about Lucene storage through the OS rather than treating heap as the index cache. Lesson 5 combines all chapter evidence into a diagnosis workflow for an AtlasMart indexing/storage regression.
Authoritative references
- Elastic index segments API — Low-level Lucene segment metadata for index shards.
- Elastic index stats API — Refresh, flush, merge, segment, translog, docs and store statistics.
- Elastic translog settings — Flush as Lucene commit plus new translog generation and request durability semantics.
- Elastic tune for indexing speed — Filesystem-cache and indexing guidance.
- Elastic tune for search speed — Filesystem cache, storage latency and local-versus-remote storage guidance.
- OpenSearch Index Segments API — Segment committed/search flags and Lucene writer-version evidence.
- OpenSearch Index Stats API — Refresh, flush, merge, segments, translog and deleted-document statistics.
- OpenSearch Refresh API — Search-visibility semantics and refresh cost guidance.
- OpenSearch Flush API — Flush and transaction-log lifecycle semantics.
- OpenSearch Force Merge API — Merge/deleted-document behavior and temporary disk-space risk.
- Elastic filesystem-cache preload — Expert setting for selected file preloading.
- OpenSearch system configuration — Memory-lock and mmap-related node settings.
- OpenSearch index settings — Index store preload and related index settings.