Chapter 04 · Indexing and CRUD: Bulk APIs, Refresh, Concurrency Control, and Idempotent Ingestion
if_seq_no/if_primary_term Concurrency Control, External Versioning Concepts, and Idempotency Keys
Protect read-modify-write paths with sequence-number/primary-term optimistic concurrency, understand external versioning limits, and design stable idempotency keys for ambiguous outcomes.
Learning outcomes
Two AtlasMart services can read the same product, compute different changes, and then write in the opposite order. Without a precondition, the later network arrival can overwrite newer business state even though both requests were individually valid. Optimistic concurrency control (OCC) makes that race observable: the client writes only if the document still has the sequence number and primary term it previously read.
Explain _seq_no as operation ordering metadata
and _primary_term as primary-generation
metadata rather than business versions.
Use if_seq_no and
if_primary_term on index/update/delete
operations and deliberately produce a stale-write conflict.
Distinguish internal OCC from external/external_gte versioning and identify the upstream ordering assumption external versioning requires.
Design stable event/idempotency keys so a network timeout does not automatically produce duplicate logical effects.
Reconcile conflicts instead of treating
retry_on_conflict or generic retries as
substitutes for business merge logic.
Examples are written against
Elasticsearch 9.5.3 and
OpenSearch 3.8.0. The portable core uses document
and index APIs verified on both products; product-specific
behavior is labeled rather than normalized away. Elasticsearch
examples assume the default self-managed distribution with its
bundled JVM. OpenSearch examples assume the upstream 3.8.0
distribution with the Security plugin present. Keep the
earlier course endpoints: Elasticsearch at
https://localhost:9200 with
ELASTIC_PASSWORD, and OpenSearch at
https://localhost:9201 with
OPENSEARCH_INITIAL_ADMIN_PASSWORD.
The generation environment does not run these two search
servers. Requests were checked against current official
documentation but were not executed here, so example responses
describe expected fields, status classes, and invariants
rather than fabricated captured output or benchmarks. Use only
disposable atlasmart-* indices, preserve the
CA/certificate paths established in Chapter 01, and never
point cleanup, force-merge, or failure-injection commands at
unrelated or production data.
Current Elasticsearch has an experimental
index.disable_sequence_numbers setting for
specific scenarios. When sequence numbers are disabled, normal
if_seq_no/if_primary_term OCC is
unavailable and update operations are not supported. This is
not a portable OpenSearch assumption. The mandatory AtlasMart
lab uses ordinary sequence-number-enabled indices on both
engines.
1. Read metadata, then write with a precondition
Every successful mutation receives a sequence number assigned by the primary shard. The primary term identifies the current primary generation. A client that reads a document can carry both values into a later write. If any mutation changed the document—or primary history changed incompatibly—the stale precondition fails instead of silently overwriting newer state.
GET atlasmart-products-v4-write-lab/_doc/P-701# Record the returned _seq_no = S and _primary_term = T
The pair (S,T) is an engine concurrency token, not
a domain revision such as “catalog version 2026-09-11.” Never
expose it as the business’s durable cross-cluster version
contract.
2. Force a lost-update race, then prevent it
Imagine Service A and Service B both read P-701 at
(S,T). Service A successfully changes price.
Service B then tries to write a stock change while still
asserting the old pair. The second write must fail with a
conflict rather than overwrite Service A’s newer source.
# Service A uses the pair read earlierPOST atlasmart-products-v4-write-lab/_update/P-701?if_seq_no=S&if_primary_term=T{"doc":{"price":34.90,"updated_at":"2026-09-11T06:00:00Z"}}# Reusing the same stale pair is now wrongPOST atlasmart-products-v4-write-lab/_update/P-701?if_seq_no=S&if_primary_term=T{"doc":{"stock":9,"updated_at":"2026-09-11T06:00:01Z"}}
Replace S and T with values captured
from the real GET. The stale request should return a conflict
class such as version_conflict_engine_exception.
The repair is not “retry the same stale update.” Re-read the
current document, re-evaluate the business intent, and issue a
new conditional write if it is still valid.
3. External versioning is a different contract
Both current products support external version concepts for
index/delete workflows. With version_type=external,
the caller supplies a monotonically increasing version and the
engine accepts only a strictly newer one.
external_gte allows equal versions as well. This is
useful only when the upstream system already owns a reliable
ordering number. It does not create such an ordering for you.
| Mechanism | Version owner | Good fit | Failure if assumption is false |
|---|---|---|---|
| Sequence number + primary term | Search engine | Protect a read-modify-write against concurrent changes in this index | Cannot be reused as a portable business revision across migrations/clusters. |
external |
Upstream system | Events/snapshots with strictly increasing authoritative revision | Out-of-order or reused upstream numbers reject/skip valid business changes. |
external_gte |
Upstream system | Replay where equal-version same-state writes are explicitly acceptable | Equal version with different payload can mask producer defects; use only with a deterministic contract. |
| Stable ID + create | Application identity/event key | Accept one logical event once | Conflicts require application interpretation; create alone does not merge business state. |
4. Idempotency keys turn ambiguous outcomes into evidence
A timeout does not prove the server rejected the write. For an
event whose logical identity is
inventory:SKU-42:warehouse-7:offset-18823, derive a
stable event ID and use a create-only record or a deterministic
product-state update that can be safely replayed. Keep the
original event ID in the source for audit. If the retry receives
a create conflict, fetch/reconcile the existing document and
compare the stored event identity before declaring success.
PUT atlasmart-product-events/_create/inventory-SKU42-W7-18823{ "event_id":"inventory-SKU42-W7-18823", "product_id":"SKU42", "warehouse_id":"W7", "source_offset":18823, "absolute_stock":17, "occurred_at":"2026-09-11T06:10:00Z"}
If the same event both creates a dedupe marker and separately decrements another document, the two writes are not automatically atomic. A crash between them can still create inconsistency. Prefer absolute-state events where possible, or use an application-level transactional/outbox/reconciliation design that explicitly handles cross-document effects.
5. Conflict/retry decision table
| Situation | Safe immediate retry? | Correct next step |
|---|---|---|
| 429 before item executes / rejected item | Often, if operation is idempotent | Backoff with jitter and retry only the rejected item. |
409 stale if_seq_no/if_primary_term
|
No | Re-read current state, merge/re-evaluate, issue a new conditional write. |
| Duplicate create for known event ID | Usually no retry needed | Verify existing event identity/state; classify as deduplicated success if it matches. |
| Network timeout after sending request | Unknown | Reconcile by stable ID/version/event key; safe replay only when semantics permit. |
| Scripted increment/decrement timeout | No blind retry | Read/reconcile using event identity or redesign to absolute state. |
Check your understanding
-
What problem does
if_seq_no/if_primary_termsolve? - Are sequence numbers business versions?
- What must be true for external versioning to be meaningful?
- Why can a timeout not be classified as a failed write?
-
Why is
retry_on_conflictnot a universal fix?
Review the answers
1. It prevents a client from applying a write based on stale document state after another mutation has occurred.
2. No. They are engine operation-ordering metadata tied to a shard/index history.
3. The upstream producer must supply a trustworthy monotonic ordering/version contract.
4. The server may have committed before the response was lost, so the result is ambiguous.
5. Retrying the server-side update does not decide whether the business operation is still valid or idempotent after concurrent changes.
Production judgment
Concurrency correctness should be designed before throughput tuning. Identify which fields have one authoritative writer, which require compare-and-set semantics, which events have stable identities, and which operations are safe to replay. Preserve conflict metrics and reconciliation outcomes: a falling error rate caused by blind overwrites is worse than visible 409s. Treat external versioning as an integration contract with the source system and test its reset/migration behavior before relying on it.
Summary and next step
You can now turn a stale write into an explicit conflict and give ambiguous events stable identity. The final lesson combines these rules with Bulk item classification, bounded backpressure, retry queues, and a reconciliation ledger.
Authoritative references
- Elastic Bulk API — Current bulk action, per-item result, refresh, versioning, routing, and OCC reference.
- Elastic refresh parameter — Current visibility semantics for index/update/delete/bulk requests.
- Elastic optimistic concurrency control — Sequence-number and primary-term concurrency semantics.
- OpenSearch Bulk API — Current NDJSON, per-item failure, OCC, versioning, and refresh behavior.
- OpenSearch Document APIs — Current document operation and sequence-number/primary-term overview.
- Elastic index settings — sequence number caveat — Current experimental sequence-number disabling behavior and its limitations.