Chapter 04 · Indexing and CRUD: Bulk APIs, Refresh, Concurrency Control, and Idempotent Ingestion

Index/Create/Update/Delete Semantics, IDs, Upserts, Scripts, and Partial Document Updates

Separate replace, create-only, partial update, scripted mutation, upsert, and delete semantics so AtlasMart can choose stable document identities and retry-safe write behavior deliberately.

Intermediate → Advanced115–135 minutesDual-platform CRUD semantics labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart receives authoritative catalog snapshots, price corrections, inventory deltas, and supplier events. All of them can reach the same product document, but they do not have the same write semantics. Replacing a document, requiring nonexistence, applying a partial patch, incrementing a value, and creating a fallback document are different contracts. If the ingestion service treats them as interchangeable, a retry can overwrite fields, duplicate a counter, or silently accept a duplicate event.

01

Distinguish index, create, update, and delete by precondition, source behavior, and retry implications.

02

Use stable AtlasMart document IDs as application identity rather than relying on auto-generated IDs where deduplication matters.

03

Explain partial updates, scripts, doc_as_upsert/upsert patterns, and why a server-side update is still a read/compute/write operation internally.

04

Inspect _version, _seq_no, _primary_term, result, shard acknowledgement, and GET/search evidence after each mutation.

05

Choose an operation from business intent and prove a deliberately wrong choice before repairing it.

Chapter baseline reviewed 11 September 2026

Examples are written against Elasticsearch 9.5.3 and OpenSearch 3.8.0. The portable core uses document and index APIs verified on both products; product-specific behavior is labeled rather than normalized away. Elasticsearch examples assume the default self-managed distribution with its bundled JVM. OpenSearch examples assume the upstream 3.8.0 distribution with the Security plugin present. Keep the earlier course endpoints: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD, and OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD.

Execution and safety note

The generation environment does not run these two search servers. Requests were checked against current official documentation but were not executed here, so example responses describe expected fields, status classes, and invariants rather than fabricated captured output or benchmarks. Use only disposable atlasmart-* indices, preserve the CA/certificate paths established in Chapter 01, and never point cleanup, force-merge, or failure-injection commands at unrelated or production data.

1. Four verbs, four different contracts

The document endpoint looks uniform, but the operation type changes the application guarantee. An index operation writes the supplied document under an ID, creating it when absent and replacing the existing source when present. A create operation adds the document only when the ID does not already exist; a duplicate normally produces a conflict. An update operation targets an existing document and can merge a partial doc, execute a script, or use an upsert policy. A delete removes the current logical document, while Lucene-level reclamation happens later through segment merging.

Intent Typical API shape Critical precondition Retry question
Replace desired state PUT index/_doc/id No create-only precondition Is sending the same full state again harmless?
Create once PUT index/_create/id or op_type=create ID must be absent Should duplicate delivery become a visible conflict?
Patch/compute POST index/_update/id Document usually exists unless upsert is supplied Is the patch or script idempotent?
Delete desired identity DELETE index/_doc/id Absence may be acceptable depending on business intent Can not-found be treated as converged state?

A stable product_id such as P-701 is also a useful document _id when AtlasMart wants retries to converge on one logical document. An auto-generated ID is appropriate when each request intentionally creates a distinct document; it is a poor deduplication key for at-least-once delivery.

2. Observe replace, create-only, patch, script, and upsert

Create a disposable governed index using the mapping contract established in Chapter 03. This lab keeps one primary and zero replicas because it is demonstrating document semantics, not high availability.

portable REST · disposable write-semantics index
DELETE atlasmart-products-v4-write-labPUT atlasmart-products-v4-write-lab{  "settings":{"number_of_shards":1,"number_of_replicas":0,"refresh_interval":"30s"},  "mappings":{"dynamic":"strict","properties":{    "product_id":{"type":"keyword"},    "name":{"type":"text","fields":{"keyword":{"type":"keyword"}}},    "price":{"type":"double"},    "stock":{"type":"integer"},    "updated_at":{"type":"date"},    "event_id":{"type":"keyword"}  }}}

Now write and inspect one product. The response should identify the document and expose mutation metadata such as _version, _seq_no, _primary_term, result, and shard acknowledgement. Exact numbers are not universal; their monotonic/change relationships are the evidence.

portable REST · create, replace, update, script, upsert
PUT atlasmart-products-v4-write-lab/_create/P-701{"product_id":"P-701","name":"USB-C Charger 65W","price":39.90,"stock":12,"updated_at":"2026-09-11T05:00:00Z","event_id":"evt-701-create"}# Repeat the same create: expect a 409 conflict, not a duplicate documentPUT atlasmart-products-v4-write-lab/_create/P-701{"product_id":"P-701","name":"USB-C Charger 65W","price":39.90,"stock":12,"updated_at":"2026-09-11T05:00:00Z","event_id":"evt-701-create"}# index replaces the source for the same IDPUT atlasmart-products-v4-write-lab/_doc/P-701{"product_id":"P-701","name":"USB-C Charger 65W","price":37.90,"stock":12,"updated_at":"2026-09-11T05:03:00Z","event_id":"evt-701-snapshot-2"}# partial update preserves unspecified source fieldsPOST atlasmart-products-v4-write-lab/_update/P-701{"doc":{"stock":11,"updated_at":"2026-09-11T05:04:00Z"}}# scripted mutation: useful, but repeated delivery changes state againPOST atlasmart-products-v4-write-lab/_update/P-701{"script":{"source":"ctx._source.stock -= params.n","params":{"n":1}}}# upsert: patch if present, create fallback if absentPOST atlasmart-products-v4-write-lab/_update/P-799{"doc":{"stock":5},"upsert":{"product_id":"P-799","name":"Travel Adapter","price":24.50,"stock":5,"updated_at":"2026-09-11T05:05:00Z","event_id":"evt-799-create"}}GET atlasmart-products-v4-write-lab/_doc/P-701GET atlasmart-products-v4-write-lab/_doc/P-799
Scripts change the retry analysis.

Repeating an absolute assignment such as “stock is 11” can converge; repeating “decrement stock by one” does not. A timeout after the server committed the first decrement leaves the client uncertain. Retrying blindly can decrement twice. The right repair is an idempotent event/state design, or a concurrency/reconciliation workflow—not simply a larger retry count.

3. Wrong approach: replace a partial object with the Index API

Suppose a service reads only price and then sends {"price":36.90} with an index operation under P-701, assuming it means “patch this field.” The index operation treats the supplied source as the new document. Under a strict mapping it may still be accepted, but the previously stored product_id, name, stock, timestamp, and event ID disappear from _source. The bug is semantic, not syntactic.

wrong operation · replacement is not a patch
PUT atlasmart-products-v4-write-lab/_doc/P-701{"price":36.90}GET atlasmart-products-v4-write-lab/_doc/P-701

Repair by sending the complete authoritative state with index, or use _update with a bounded partial doc when patch semantics are intended. If concurrent writers can modify the same product, Lesson 4 adds an explicit sequence-number/primary-term precondition.

4. Delete is logical state change, not immediate byte erasure

A successful delete means the current document is no longer logically available through normal reads/search after visibility catches up. At the Lucene layer, deleted documents are typically marked and later reclaimed as segments merge. This is why delete count, disk usage, refresh visibility, and segment merging must not be conflated. Recreating the same ID is a new mutation with new concurrency metadata.

portable REST · delete and verify logical state
DELETE atlasmart-products-v4-write-lab/_doc/P-799GET atlasmart-products-v4-write-lab/_doc/P-799POST atlasmart-products-v4-write-lab/_refreshGET atlasmart-products-v4-write-lab/_search{"query":{"term":{"product_id":"P-799"}}}

5. Lab acceptance criteria

Repeat the sequence on both local engines and preserve the responses. The numbers may differ, but the following invariants should hold: create-only duplicate delivery conflicts; index replacement can remove omitted source fields; partial update preserves fields not mentioned by the patch; a scripted decrement changes state on every successful execution; an upsert creates the missing target; delete removes the logical document; and each accepted mutation advances concurrency metadata.

Check your understanding

  1. Why is index not the same as a patch?
  2. When is create useful for idempotency?
  3. Why is a decrement script dangerous under ambiguous retry?
  4. What do _seq_no and _primary_term prepare you to do?
  5. Does a delete immediately remove all bytes from Lucene segments?
Review the answers

1. It writes the supplied document as replacement desired state for the ID; omitted source fields are not implicitly preserved.

2. When a stable business/event ID should be accepted once and duplicate delivery should surface as a conflict rather than create another logical record.

3. Because the operation is non-idempotent: a committed first attempt plus a retry applies the decrement twice.

4. Attach a compare-and-set style precondition to a later write so stale read-modify-write attempts fail instead of overwriting newer state.

5. No. It removes the logical document; physical reclamation is associated with later segment merge behavior.

Production judgment

Choose the narrowest operation that matches business intent. Stable IDs and create-only writes are powerful for event deduplication, full index replacement is appropriate for authoritative snapshots, partial updates reduce payload size but add server-side read/compute/write work, and scripts create a larger correctness and security surface. Record whether retries are safe, which failures require reconciliation, and whether the write is allowed to overwrite concurrent changes. Throughput optimizations never justify an operation whose retry semantics are undefined.

Summary and next step

You can now explain what a single AtlasMart mutation means. The next lesson batches many such actions into one NDJSON Bulk request and shows why an HTTP-level success can still contain document-level failures that must be classified individually.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.