Chapter 01 · Search Engine Foundations, Elasticsearch vs OpenSearch, Deployment Models, and Lab Setup
Build a Reproducible Lab with Sample Data, TLS/Auth, Persistent Storage, Metrics, and Safe Reset Automation
Turn the dual-platform setup into a persistent, deterministic, measurable and safely resettable AtlasMart course lab.
Learning outcomes
A useful course lab must survive a browser refresh, produce the same sample data on demand, reveal its resource/security state, and be safely disposable. This lesson turns the two Chapter 01 containers into a repeatable AtlasMart baseline that later chapters can extend without inventing new names or silently changing assumptions.
Define stable Chapter 01 names, ports, versions, credentials handling, volumes, index names and dataset IDs.
Load a deterministic AtlasMart catalog through the bulk API and verify item-level success plus searchable document counts.
Collect cluster/node/index/JVM/filesystem evidence without presenting synthetic performance numbers as measurements.
Distinguish TLS encryption, certificate verification and authentication for Elasticsearch versus the OpenSearch demo-security path.
Perform soft reset and full destructive reset with an explicit blast radius and verification checklist.
This chapter pins Elasticsearch 9.5.3 (released 3
September 2026) and OpenSearch 3.8.0 (released 4
August 2026) for reproducible examples. OpenSearch 3.9.0 is
scheduled for 29 September 2026 and is therefore not treated
as current. Re-check both projects before reusing these
commands later. Elasticsearch and OpenSearch are independent
products: shared Lucene ancestry does not make their APIs,
plugins, security, lifecycle, vector features, clients, or
managed offerings interchangeable.
The environment used to generate this lesson does not provide
Docker, Elasticsearch, OpenSearch, Kibana, or OpenSearch
Dashboards. The commands and API shapes were reviewed against
the current official documentation but were not executed here.
Expected output is described by invariant and field shape
rather than presented as captured benchmark evidence. Every
destructive action is scoped to
atlasmart-* course containers, volumes, indices,
and local loopback ports.
1. Freeze the lab contract before adding features
| Contract item | Elasticsearch baseline | OpenSearch baseline |
|---|---|---|
| Server version | 9.5.3 | 3.8.0 |
| Container | atlasmart-es |
atlasmart-os |
| Host endpoint | https://localhost:9200 |
https://localhost:9201 |
| Persistence | atlasmart-es-data |
atlasmart-os-data |
| Security | Auto-configured Elasticsearch TLS/auth; trust copied HTTP CA |
Security-plugin demo TLS/auth; local -k only
until proper lab certs are introduced
|
| Index | atlasmart-products-v1 |
same logical fixture name |
| Topology | 1 node, 1 primary shard, 0 replicas for fixture | 1 node, 1 primary shard, 0 replicas for fixture |
| Source data | Five synthetic products with stable IDs | identical logical dataset |
The zero-replica setting is intentionally not a production recommendation. It prevents a single-node teaching cluster from implying replica availability it cannot provide. Later topology chapters change this assumption explicitly.
2. Load deterministic data and inspect every bulk item
The bulk endpoint uses newline-delimited JSON. An HTTP-level
success does not imply every item succeeded, so the lab verifies
the top-level errors flag and item statuses. That
invariant becomes important later when retries and idempotency
are introduced.
{"index":{"_index":"atlasmart-products-v1","_id":"P-1001"}}{"product_id":"P-1001","name":"Quiet Wireless Keyboard","category":"keyboards","brand":"AtlasKey","description":"compact wireless keyboard with quiet keys","price":49.90}{"index":{"_index":"atlasmart-products-v1","_id":"P-1002"}}{"product_id":"P-1002","name":"Mechanical Gaming Keyboard","category":"keyboards","brand":"AtlasKey","description":"wired mechanical keyboard with tactile switches","price":89.00}{"index":{"_index":"atlasmart-products-v1","_id":"P-1003"}}{"product_id":"P-1003","name":"Wireless Travel Mouse","category":"mice","brand":"RoutePoint","description":"compact wireless mouse for travel","price":29.50}{"index":{"_index":"atlasmart-products-v1","_id":"P-1004"}}{"product_id":"P-1004","name":"USB-C Travel Hub","category":"adapters","brand":"RoutePoint","description":"portable usb c hub with hdmi and ethernet","price":54.75}{"index":{"_index":"atlasmart-products-v1","_id":"P-1005"}}{"product_id":"P-1005","name":"Ergonomic Keyboard Wrist Rest","category":"accessories","brand":"DeskCare","description":"memory foam wrist rest for full size keyboards","price":18.25}
Ensure the file ends with a newline. Submit it to each endpoint
with Content-Type: application/x-ndjson and
?refresh=true only for this deterministic lab.
# Elasticsearchcurl --cacert atlasmart-es-http-ca.crt -u "elastic:$ELASTIC_PASSWORD" -H "Content-Type: application/x-ndjson" --data-binary @atlasmart-products.bulk.ndjson "https://localhost:9200/_bulk?refresh=true"# OpenSearch demo securitycurl -k -u "admin:$OPENSEARCH_INITIAL_ADMIN_PASSWORD" -H "Content-Type: application/x-ndjson" --data-binary @atlasmart-products.bulk.ndjson "https://localhost:9201/_bulk?refresh=true"
Acceptance: top-level errors is false,
five item operations report successful statuses, and
GET /atlasmart-products-v1/_count returns five. If
errors is true, inspect each failed item before
retrying. Never retry the whole bulk request blindly when IDs or
operations are non-idempotent.
3. Query evidence is separate from operational evidence
GET /atlasmart-products-v1/_search{ "query": { "bool": { "must": [{"match": {"description": "wireless keyboard"}}], "filter": [{"term": {"category": "keyboards"}}] } }, "sort": ["_score", {"product_id":"asc"}], "_source": ["product_id","name","category","brand","price"]}
The expected product ID is P-1001. Record the
returned IDs and scores, but assert semantic expectations
primarily on IDs/order, not frozen score constants.
GET /_cluster/healthGET /_nodes?filter_path=cluster_name,nodes.*.name,nodes.*.version,nodes.*.roles,nodes.*.jvm.versionGET /_nodes/stats/jvm,process,fs,http,indices?filter_path=nodes.*.jvm.mem.heap_used_percent,nodes.*.process.cpu.percent,nodes.*.fs.total.available_in_bytes,nodes.*.http.current_open,nodes.*.indices.docs,nodes.*.indices.storeGET /_cat/indices/atlasmart-products-v1?v=true
These metrics are snapshots of a tiny lab, not performance conclusions. Do not publish “Elasticsearch used X% less memory” from one idle sample. Later benchmarks disclose resources, warmup, segment/cache state, query mix, concurrency and tail-latency distributions.
4. Security evidence: encrypted is not the same as verified
Elasticsearch’s local Docker path gives this chapter a CA
certificate that curl can trust explicitly. OpenSearch’s demo
security path gives encrypted HTTPS but the documented
quickstart uses -k because the demo certificates
are self-signed and not suitable for normal hostname
verification. Therefore the two labs intentionally have
different verification strength.
Authentication is separate again: elastic and
OpenSearch demo admin credentials prove only that
those test identities were accepted. Later security chapters
replace administrator usage with least-privilege application
identities and product-specific role models. Never place these
credentials, CA private keys or future cloud secrets in the
academy repository.
Put passwords directly in docker-compose.yml,
commit them, use -k for every product, and expose
0.0.0.0:9200. That collapses secret management,
server identity and network exposure into unsafe convenience.
Repair by using shell/.env secrets excluded from source
control, CA verification where available, proper certificates
for production, loopback/private networking for the local lab,
and least privilege.
5. Reset automation with two blast-radius levels
A soft reset deletes only the course index while retaining containers, credentials and volumes. A full reset destroys the named course containers and volumes. The scripts must refuse to use wildcards or unrelated names.
DELETE /atlasmart-products-v1# verify absenceGET /_cat/indices/atlasmart-products-v1?v=true
docker rm -f atlasmart-es atlasmart-osdocker volume rm atlasmart-es-data atlasmart-os-datadocker network rm atlasmart-search# local certificate copy created by this course only# remove ./atlasmart-es-http-ca.crt manually after verifying the path
Blast radius: the full reset permanently
removes only data stored in atlasmart-es-data and
atlasmart-os-data. Before execution, verify the
names with docker ps -a and
docker volume ls. Do not replace them with patterns
such as docker system prune in course instructions.
Verification checklist
- Root response matches the intended product and pinned version.
- TLS/auth behavior is documented next to each endpoint.
- Both fixture indices have one primary, zero replicas and five documents after load.
- The bulk response has no failed items.
-
The search fixture returns
P-1001for the required query/filter. - Cluster/node/JVM/filesystem evidence is captured without inventing performance conclusions.
- Soft reset removes only the fixture index; full reset names only course containers/volumes/network.
Production judgment
A production platform needs stronger identity, least privilege, replica/failure-domain design, snapshot/restore validation, monitoring, capacity headroom, upgrade policy and managed-service-aware controls. This Chapter 01 lab intentionally proves only the baseline mechanisms needed for later lessons. Chapter 02 will expand the document/index/shard model without changing these names silently.
Check your understanding
- Why must bulk responses be checked item by item?
-
Why is
refresh=trueacceptable here but not a default production ingestion pattern? -
What does
-kfail to verify in the OpenSearch demo request? - Why are the fixture indices configured with zero replicas?
- What is the difference between soft reset and full reset?
Review the answers
1. The bulk request can succeed at the HTTP level while individual operations fail; retries and reconciliation must be based on item results.
2. It makes the tiny lab immediately searchable and deterministic, but forcing refresh frequently can reduce indexing efficiency and increase segment work.
3. It skips normal server-certificate/hostname verification, so it does not authenticate the server identity in the way a trusted CA/hostname path should.
4. A one-node lab cannot place a replica on a second node; zero replicas avoids teaching an impossible HA guarantee and keeps health interpretation deterministic.
5. Soft reset removes only the course index and keeps the environment; full reset destroys the named course containers, their persistent volumes and course network.
Summary and next step
Chapter 01 now has a stable dual-platform contract: pinned versions, isolated endpoints, explicit security differences, persistent storage, deterministic data, observable health/metrics, and safe reset rules. Chapter 02 can build on this foundation to explain documents, indices, data streams, shards, replicas, nodes and distributed request flow.
Authoritative references
- Elasticsearch Docker installation — Pinned Docker, secure startup and certificate workflow.
- Elasticsearch bulk API — Bulk request format and item-level result semantics.
- Elasticsearch nodes stats API — Structured node/JVM/filesystem/index statistics.
- OpenSearch Docker installation — Pinned local Docker and admin-password requirements.
- OpenSearch bulk API — OpenSearch bulk format and behavior.
- OpenSearch security demo configuration — Demo TLS/auth configuration and limitations.