Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering

match, multi_match, match_phrase, query_string / simple_query_string, and Phrase/Proximity Semantics

Build lexical queries from analyzer behavior rather than syntax memorization: compare match, multi_match, phrase/proximity, strict query_string, and fault-tolerant simple_query_string against deterministic AtlasMart fixtures.

Intermediate100–120 minutesDeterministic Query DSL evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart's search box accepts ordinary phrases such as wireless headphones, but its support team also wants fielded expert syntax for internal troubleshooting. Those are different contracts. A useful search API must decide when text is analyzed, which fields are searched, whether term order matters, how proximity is interpreted, and whether the caller is allowed to supply operators. This lesson makes those choices observable instead of treating every text query as a string passed to _search.

01

Explain why full-text queries analyze input while term-level queries do not, and predict the resulting indexed/query terms from Chapter 05 analyzers.

02

Compare match, multi_match and match_phrase semantics, including field selection, field boosts, phrase positions and slop.

03

Distinguish strict query_string parsing from fault-tolerant simple_query_string and explain why neither should receive unconstrained user input by default.

04

Use named queries, hits, _score, matched_queries and targeted Explain/Profile evidence without treating scores as absolute probabilities.

05

Build adversarial tests that separate ordinary free-text intent from explicit expert-query syntax and expensive wildcard/regex constructs.

Chapter baseline reviewed 11 September 2026

Examples target Elasticsearch 9.5.3 and OpenSearch 3.8.0, both with a plain lexical text/keyword mapping. Elasticsearch can give match special semantic behavior on semantic field types; that is outside this lexical fixture and must not be generalized to OpenSearch.

Execution note

This generation environment does not run Elasticsearch/OpenSearch containers. Requests below are checked against current official documentation and use deterministic fixtures, but any response fragments are described as expected invariants rather than captured benchmark output. Run them against the pinned Chapter 01 lab before recording timings or scores.

1. Build one deterministic lexical fixture

Use a dedicated disposable index so query behavior is not contaminated by earlier mappings, synonyms or hidden data. The name and description fields are analyzed text; brand, category and sku are exact-value keyword fields. The six documents deliberately overlap on words such as headphones, Bluetooth, and noise cancelling.

setup · mapping shared by Chapter 06
PUT atlasmart-products-query-v1
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0
  },
  "mappings": {
    "properties": {
      "sku":      { "type": "keyword" },
      "name":     { "type": "text", "fields": { "raw": { "type": "keyword" } } },
      "description": { "type": "text" },
      "brand":    { "type": "keyword" },
      "category": { "type": "keyword" },
      "price":    { "type": "double" },
      "available":{ "type": "boolean" },
      "rating":   { "type": "float" },
      "warehouse":{ "type": "geo_point" },
      "attributes": {
        "type": "nested",
        "properties": {
          "name":  { "type": "keyword" },
          "value": { "type": "keyword" }
        }
      }
    }
  }
}
setup · deterministic AtlasMart documents
POST atlasmart-products-query-v1/_bulk?refresh=true
{ "index": { "_id": "p1" } }
{ "sku":"AM-AU-100","name":"Wireless Noise Cancelling Headphones","description":"Over-ear Bluetooth headphones with active noise cancellation","brand":"Auralux","category":"electronics/audio","price":199.99,"available":true,"rating":4.7,"warehouse":{"lat":40.7128,"lon":-74.0060},"attributes":[{"name":"color","value":"black"},{"name":"connectivity","value":"bluetooth"}] }
{ "index": { "_id": "p2" } }
{ "sku":"AM-AU-200","name":"Wired Studio Headphones","description":"Closed-back monitoring headphones for studio recording","brand":"Auralux","category":"electronics/audio","price":89.99,"available":true,"rating":4.4,"warehouse":{"lat":40.7306,"lon":-73.9352},"attributes":[{"name":"color","value":"black"},{"name":"connectivity","value":"wired"}] }
{ "index": { "_id": "p3" } }
{ "sku":"AM-WB-300","name":"Portable Bluetooth Speaker","description":"Compact waterproof speaker for travel","brand":"WaveBox","category":"electronics/audio","price":79.99,"available":false,"rating":4.2,"warehouse":{"lat":34.0522,"lon":-118.2437},"attributes":[{"name":"color","value":"blue"},{"name":"connectivity","value":"bluetooth"}] }
{ "index": { "_id": "p4" } }
{ "sku":"AM-AU-400","name":"USB-C Noise Cancelling Earbuds","description":"In-ear USB-C earbuds with active noise cancellation","brand":"Auralux","category":"electronics/audio","price":129.99,"available":true,"rating":4.5,"warehouse":{"lat":40.6500,"lon":-73.9496},"attributes":[{"name":"color","value":"white"},{"name":"connectivity","value":"usb-c"}] }
{ "index": { "_id": "p5" } }
{ "sku":"AM-WN-500","name":"Aluminum Laptop Stand","description":"Adjustable desktop stand for laptops","brand":"WorkNest","category":"office/accessories","price":49.00,"available":true,"rating":4.6,"warehouse":{"lat":41.8781,"lon":-87.6298},"attributes":[{"name":"color","value":"silver"},{"name":"material","value":"aluminum"}] }
{ "index": { "_id": "p6" } }
{ "sku":"AM-ST-600","name":"Lightweight Running Shoes","description":"Breathable road running shoes","brand":"Stride","category":"sports/running","price":69.00,"available":true,"rating":4.1,"warehouse":{"lat":42.3601,"lon":-71.0589},"attributes":[{"name":"color","value":"black"},{"name":"size","value":"42"}] }

Because bulk indexing uses refresh=true only for this small teaching fixture, the first search can run immediately. Chapter 04 explained why production ingestion should not force refresh on every batch.

2. match turns user text into analyzed query terms

A match query is appropriate when the caller supplies natural language for an analyzed field. The search analyzer transforms Wireless HEADPHONES before the query is built. On the standard analyzer, case is normalized and punctuation/word boundaries affect terms. The important evidence is not the original string; it is the query token stream established in Chapter 05.

query · analyzed free text
GET atlasmart-products-query-v1/_search
{
  "query": {
    "match": {
      "name": {
        "query": "wireless headphones",
        "operator": "and",
        "_name": "name_terms"
      }
    }
  }
}

With this fixture, p1 is the deterministic target because its name contains both analyzed terms. A different analyzer, synonym graph or stemming rule can change the built query; therefore Chapter 05 token tests are a dependency of Query DSL tests.

3. multi_match is not “search everything”

multi_match applies match-like logic across an explicit field set. Field selection is part of the contract. AtlasMart can search name and description while boosting name, instead of querying every mapped field and hoping the score is useful. Modes such as best_fields, most_fields, cross_fields, phrase, phrase_prefix and bool_prefix create different internal query shapes; pick one because its semantics fit the data model, not because one happens to rank a demo better.

query · explicit fields and relative boost
GET atlasmart-products-query-v1/_search
{
  "query": {
    "multi_match": {
      "query": "noise cancelling",
      "fields": ["name^3", "description"],
      "type": "best_fields"
    }
  }
}
Choice Useful when Failure to avoid
best_fields One field should dominate relevance Assuming scores from different fields have identical distributions
most_fields Several similarly analyzed fields contribute evidence Duplicating the same content across fields and accidentally over-weighting it
cross_fields Words may be split across fields with compatible analyzers Using incompatible analyzers and expecting one logical bag of terms
phrase/phrase_prefix Ordered positional matching is intentional Treating phrase-prefix expansion as free autocomplete at arbitrary scale

4. Phrases are about token positions, not substring matching

match_phrase analyzes the input and then requires the resulting terms to satisfy phrase-position constraints. slop relaxes the allowed positional movement; it is not a character distance. This is why stop filters, multiword synonyms and token graphs can alter phrase behavior even when the visible text looks similar.

query · exact analyzed phrase
GET atlasmart-products-query-v1/_search
{
  "query": {
    "match_phrase": {
      "name": {
        "query": "noise cancelling",
        "slop": 0
      }
    }
  }
}

The deterministic target set is p1 and p4. The lesson's acceptance criterion is not their numeric _score; it is that the expected IDs match, no unrelated ID appears, and the analyzer/phrase contract is recorded.

5. query_string and simple_query_string are parsers, not escaping shortcuts

query_string exposes a strict query syntax with Boolean operators, field syntax, phrases, ranges, wildcards and more. Invalid syntax can fail the request. That power makes it appropriate for an explicitly documented expert-query surface, not for blindly interpolating a public search-box value. simple_query_string uses a smaller, fault-tolerant syntax and ignores invalid syntax rather than failing, but it still supports operators and can construct expensive queries depending on enabled flags and input.

wrong · user becomes part of the query language
# BAD: raw q is pasted into a query_string query
q = user_input   # e.g. headphones OR *:*
body = {"query":{"query_string":{"query": q}}}
safer default · user is data, structure is fixed
GET atlasmart-products-query-v1/_search
{
  "query": {
    "multi_match": {
      "query": "headphones OR *:*",
      "fields": ["name^3", "description"],
      "type": "best_fields"
    }
  }
}

In the safer request, OR and punctuation are analyzed as text according to the field analyzer; they are not granted query-language authority. If AtlasMart intentionally offers expert syntax, place it behind authentication, field allowlists, length/complexity limits and expensive-query policy rather than pretending that HTML escaping solves query semantics.

6. Observe matching evidence without over-reading scores

Add unique _name values to clauses when you need to learn which branch matched. Returned hits can report matched_queries. Use Explain for a specific document when you need score reasoning and Profile for controlled diagnosis of execution structure; both add overhead and neither replaces a production-like latency benchmark. Always inspect _shards.failed and errors before accepting a search response as complete.

Check your understanding

  1. Why is match appropriate for AtlasMart free-text product names?
  2. What does match_phrase constrain that match does not?
  3. Why is query_string a risky default for a public search box?
  4. Does simple_query_string make arbitrary input automatically safe?
  5. What should a deterministic lexical regression test assert before comparing scores?
Review the answers

1. It analyzes the user text using the field/search analyzer and builds lexical clauses from the resulting terms instead of requiring one exact stored term.

2. It preserves token-position/order semantics and optionally allows controlled positional movement with slop.

3. The user controls a strict query language with operators and potentially expensive constructs; malformed syntax can also fail the request.

4. No. It is more fault tolerant and has limited syntax, but operators/expansions still exist; validation, field allowlists and resource controls remain necessary.

5. The expected document IDs, clause behavior/token contract, shard-success state and relevant mappings/versions; raw score values are query-relative implementation output.

Production judgment

Expose the smallest query language your product actually needs. Ordinary shoppers normally need analyzed text plus explicit filters; expert operators are a separate product feature with separate abuse controls. Keep field lists explicit, measure phrase/prefix expansion on representative data, and regression-test IDs/ranking judgments after analyzer or synonym changes. Chapter 07 will address ranking quality; Chapter 06 first establishes whether the query means what the API claims it means.

Summary and next step

You can now distinguish analyzed free-text search from phrase/proximity semantics and from user-supplied query languages. Next, use term-level queries for exact values and reason about the cost of prefix, wildcard and regular-expression expansion.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.