Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering
match, multi_match, match_phrase, query_string / simple_query_string, and Phrase/Proximity Semantics
Build lexical queries from analyzer behavior rather than syntax memorization: compare match, multi_match, phrase/proximity, strict query_string, and fault-tolerant simple_query_string against deterministic AtlasMart fixtures.
Learning outcomes
AtlasMart's search box accepts ordinary phrases such as
wireless headphones, but its support team also wants
fielded expert syntax for internal troubleshooting. Those are
different contracts. A useful search API must decide when text
is analyzed, which fields are searched, whether term order
matters, how proximity is interpreted, and whether the caller is
allowed to supply operators. This lesson makes those choices
observable instead of treating every text query as a string
passed to _search.
Explain why full-text queries analyze input while term-level queries do not, and predict the resulting indexed/query terms from Chapter 05 analyzers.
Compare match, multi_match and match_phrase semantics, including field selection, field boosts, phrase positions and slop.
Distinguish strict query_string parsing from fault-tolerant simple_query_string and explain why neither should receive unconstrained user input by default.
Use named queries, hits, _score, matched_queries and targeted Explain/Profile evidence without treating scores as absolute probabilities.
Build adversarial tests that separate ordinary free-text intent from explicit expert-query syntax and expensive wildcard/regex constructs.
Examples target Elasticsearch 9.5.3 and OpenSearch 3.8.0, both
with a plain lexical text/keyword
mapping. Elasticsearch can give match special
semantic behavior on semantic field types; that is outside
this lexical fixture and must not be generalized to
OpenSearch.
This generation environment does not run Elasticsearch/OpenSearch containers. Requests below are checked against current official documentation and use deterministic fixtures, but any response fragments are described as expected invariants rather than captured benchmark output. Run them against the pinned Chapter 01 lab before recording timings or scores.
1. Build one deterministic lexical fixture
Use a dedicated disposable index so query behavior is not
contaminated by earlier mappings, synonyms or hidden data. The
name and description fields are
analyzed text; brand, category and
sku are exact-value keyword fields. The six
documents deliberately overlap on words such as
headphones, Bluetooth, and
noise cancelling.
PUT atlasmart-products-query-v1
{
"settings": {
"number_of_shards": 1,
"number_of_replicas": 0
},
"mappings": {
"properties": {
"sku": { "type": "keyword" },
"name": { "type": "text", "fields": { "raw": { "type": "keyword" } } },
"description": { "type": "text" },
"brand": { "type": "keyword" },
"category": { "type": "keyword" },
"price": { "type": "double" },
"available":{ "type": "boolean" },
"rating": { "type": "float" },
"warehouse":{ "type": "geo_point" },
"attributes": {
"type": "nested",
"properties": {
"name": { "type": "keyword" },
"value": { "type": "keyword" }
}
}
}
}
}
POST atlasmart-products-query-v1/_bulk?refresh=true
{ "index": { "_id": "p1" } }
{ "sku":"AM-AU-100","name":"Wireless Noise Cancelling Headphones","description":"Over-ear Bluetooth headphones with active noise cancellation","brand":"Auralux","category":"electronics/audio","price":199.99,"available":true,"rating":4.7,"warehouse":{"lat":40.7128,"lon":-74.0060},"attributes":[{"name":"color","value":"black"},{"name":"connectivity","value":"bluetooth"}] }
{ "index": { "_id": "p2" } }
{ "sku":"AM-AU-200","name":"Wired Studio Headphones","description":"Closed-back monitoring headphones for studio recording","brand":"Auralux","category":"electronics/audio","price":89.99,"available":true,"rating":4.4,"warehouse":{"lat":40.7306,"lon":-73.9352},"attributes":[{"name":"color","value":"black"},{"name":"connectivity","value":"wired"}] }
{ "index": { "_id": "p3" } }
{ "sku":"AM-WB-300","name":"Portable Bluetooth Speaker","description":"Compact waterproof speaker for travel","brand":"WaveBox","category":"electronics/audio","price":79.99,"available":false,"rating":4.2,"warehouse":{"lat":34.0522,"lon":-118.2437},"attributes":[{"name":"color","value":"blue"},{"name":"connectivity","value":"bluetooth"}] }
{ "index": { "_id": "p4" } }
{ "sku":"AM-AU-400","name":"USB-C Noise Cancelling Earbuds","description":"In-ear USB-C earbuds with active noise cancellation","brand":"Auralux","category":"electronics/audio","price":129.99,"available":true,"rating":4.5,"warehouse":{"lat":40.6500,"lon":-73.9496},"attributes":[{"name":"color","value":"white"},{"name":"connectivity","value":"usb-c"}] }
{ "index": { "_id": "p5" } }
{ "sku":"AM-WN-500","name":"Aluminum Laptop Stand","description":"Adjustable desktop stand for laptops","brand":"WorkNest","category":"office/accessories","price":49.00,"available":true,"rating":4.6,"warehouse":{"lat":41.8781,"lon":-87.6298},"attributes":[{"name":"color","value":"silver"},{"name":"material","value":"aluminum"}] }
{ "index": { "_id": "p6" } }
{ "sku":"AM-ST-600","name":"Lightweight Running Shoes","description":"Breathable road running shoes","brand":"Stride","category":"sports/running","price":69.00,"available":true,"rating":4.1,"warehouse":{"lat":42.3601,"lon":-71.0589},"attributes":[{"name":"color","value":"black"},{"name":"size","value":"42"}] }
Because bulk indexing uses refresh=true only for
this small teaching fixture, the first search can run
immediately. Chapter 04 explained why production ingestion
should not force refresh on every batch.
2. match turns user text into analyzed query terms
A match query is appropriate when the caller
supplies natural language for an analyzed field. The search
analyzer transforms Wireless HEADPHONES before the
query is built. On the standard analyzer, case is normalized and
punctuation/word boundaries affect terms. The important evidence
is not the original string; it is the query token stream
established in Chapter 05.
GET atlasmart-products-query-v1/_search
{
"query": {
"match": {
"name": {
"query": "wireless headphones",
"operator": "and",
"_name": "name_terms"
}
}
}
}
With this fixture, p1 is the deterministic target
because its name contains both analyzed terms. A different
analyzer, synonym graph or stemming rule can change the built
query; therefore Chapter 05 token tests are a dependency of
Query DSL tests.
3. multi_match is not “search everything”
multi_match applies match-like logic across an
explicit field set. Field selection is part of the contract.
AtlasMart can search name and
description while boosting name,
instead of querying every mapped field and hoping the score is
useful. Modes such as best_fields,
most_fields, cross_fields,
phrase, phrase_prefix and
bool_prefix create different internal query shapes;
pick one because its semantics fit the data model, not because
one happens to rank a demo better.
GET atlasmart-products-query-v1/_search
{
"query": {
"multi_match": {
"query": "noise cancelling",
"fields": ["name^3", "description"],
"type": "best_fields"
}
}
}
| Choice | Useful when | Failure to avoid |
|---|---|---|
| best_fields | One field should dominate relevance | Assuming scores from different fields have identical distributions |
| most_fields | Several similarly analyzed fields contribute evidence | Duplicating the same content across fields and accidentally over-weighting it |
| cross_fields | Words may be split across fields with compatible analyzers | Using incompatible analyzers and expecting one logical bag of terms |
| phrase/phrase_prefix | Ordered positional matching is intentional | Treating phrase-prefix expansion as free autocomplete at arbitrary scale |
4. Phrases are about token positions, not substring matching
match_phrase analyzes the input and then requires
the resulting terms to satisfy phrase-position constraints.
slop relaxes the allowed positional movement; it is
not a character distance. This is why stop filters, multiword
synonyms and token graphs can alter phrase behavior even when
the visible text looks similar.
GET atlasmart-products-query-v1/_search
{
"query": {
"match_phrase": {
"name": {
"query": "noise cancelling",
"slop": 0
}
}
}
}
The deterministic target set is p1 and
p4. The lesson's acceptance criterion is not their
numeric _score; it is that the expected IDs match,
no unrelated ID appears, and the analyzer/phrase contract is
recorded.
5. query_string and
simple_query_string are parsers, not escaping
shortcuts
query_string exposes a strict query syntax with
Boolean operators, field syntax, phrases, ranges, wildcards and
more. Invalid syntax can fail the request. That power makes it
appropriate for an explicitly documented expert-query surface,
not for blindly interpolating a public search-box value.
simple_query_string uses a smaller, fault-tolerant
syntax and ignores invalid syntax rather than failing, but it
still supports operators and can construct expensive queries
depending on enabled flags and input.
# BAD: raw q is pasted into a query_string query
q = user_input # e.g. headphones OR *:*
body = {"query":{"query_string":{"query": q}}}
GET atlasmart-products-query-v1/_search
{
"query": {
"multi_match": {
"query": "headphones OR *:*",
"fields": ["name^3", "description"],
"type": "best_fields"
}
}
}
In the safer request, OR and punctuation are
analyzed as text according to the field analyzer; they are not
granted query-language authority. If AtlasMart intentionally
offers expert syntax, place it behind authentication, field
allowlists, length/complexity limits and expensive-query policy
rather than pretending that HTML escaping solves query
semantics.
6. Observe matching evidence without over-reading scores
Add unique _name values to clauses when you need to
learn which branch matched. Returned hits can report
matched_queries. Use Explain for a specific
document when you need score reasoning and Profile for
controlled diagnosis of execution structure; both add overhead
and neither replaces a production-like latency benchmark. Always
inspect _shards.failed and errors before accepting
a search response as complete.
Check your understanding
- Why is match appropriate for AtlasMart free-text product names?
- What does match_phrase constrain that match does not?
- Why is query_string a risky default for a public search box?
- Does simple_query_string make arbitrary input automatically safe?
- What should a deterministic lexical regression test assert before comparing scores?
Review the answers
1. It analyzes the user text using the field/search analyzer and builds lexical clauses from the resulting terms instead of requiring one exact stored term.
2. It preserves token-position/order semantics and optionally allows controlled positional movement with slop.
3. The user controls a strict query language with operators and potentially expensive constructs; malformed syntax can also fail the request.
4. No. It is more fault tolerant and has limited syntax, but operators/expansions still exist; validation, field allowlists and resource controls remain necessary.
5. The expected document IDs, clause behavior/token contract, shard-success state and relevant mappings/versions; raw score values are query-relative implementation output.
Production judgment
Expose the smallest query language your product actually needs. Ordinary shoppers normally need analyzed text plus explicit filters; expert operators are a separate product feature with separate abuse controls. Keep field lists explicit, measure phrase/prefix expansion on representative data, and regression-test IDs/ranking judgments after analyzer or synonym changes. Chapter 07 will address ranking quality; Chapter 06 first establishes whether the query means what the API claims it means.
Summary and next step
You can now distinguish analyzed free-text search from phrase/proximity semantics and from user-supplied query languages. Next, use term-level queries for exact values and reason about the cost of prefix, wildcard and regular-expression expansion.
Authoritative references
- Elastic: Query DSL overview — Current leaf/compound query, query/filter context, and expensive-query reference.
- Elastic: full-text queries — Current full-text query families and analyzed-input semantics.
- Elastic: term-level queries — Exact-value and structured-query semantics.
- Elastic: bool query — must/should/filter/must_not, minimum_should_match, and named-query behavior.
- Elastic: query and filter context — Scoring versus binary filtering semantics.
- OpenSearch: Query DSL — Current OpenSearch query categories and expensive-query controls.
- OpenSearch: bool query — Boolean clause semantics and minimum_should_match behavior.
- OpenSearch: wildcard query — Wildcard cost, rewrite behavior, and search.allow_expensive_queries interaction.
- OpenSearch: script query — Painless script-query behavior and cost warning.
- Elastic: match query — Analyzed match semantics and parameters.
- Elastic: multi_match query — Multi-field modes and field boosts.
- Elastic: match_phrase query — Phrase analysis and slop semantics.
- Elastic: query_string query — Strict query-string parser and search-box warning.
- Elastic: simple_query_string query — Fault-tolerant parser and supported operators.