Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering

term/terms, range, exists, prefix/wildcard/regexp, IDs, and Cost Awareness

Use exact-term and structured predicates deliberately: distinguish term/terms/range/exists/IDs from analyzed full-text search, constrain prefix/wildcard/regexp expansion, and prove expensive-query behavior instead of guessing.

Intermediate100–120 minutesDeterministic Query DSL evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

After AtlasMart can interpret natural-language text, it still needs exact SKU/brand/category filters, numeric price windows, existence checks and controlled prefix behavior. Those predicates should not be expressed as full-text guesses. This lesson separates exact indexed terms from analyzed text and makes multi-term expansion cost visible before it reaches production traffic.

01

Use term and terms for exact indexed values and explain why a term query on analyzed text often surprises users.

02

Use range and exists queries with field-type and missing-value semantics appropriate to numeric/date/keyword data.

03

Use IDs queries for document identity without pretending document IDs are a business secondary index.

04

Compare prefix, wildcard and regexp as term-enumeration tools and relate their cost to mappings, patterns and search.allow_expensive_queries.

05

Design validation and query-shape alternatives that avoid unbounded expansion while preserving user intent.

Portability rule

The examples use Query DSL constructs shared by Elasticsearch 9.5.3 and OpenSearch 3.8.0. Rewrite options, optimized field types/settings and expensive-query exceptions differ over time. Treat search.allow_expensive_queries as a guardrail to verify, not as proof that every allowed query is cheap.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. term and terms operate on indexed terms

A term-level query does not run the field's full-text analyzer. On keyword fields this is exactly what AtlasMart needs for SKU, brand and category. On text, the original human string may have been lowercased or split into multiple indexed terms, so copying the original visible value into a term query can miss unexpectedly.

exact filters · keyword fields
GET atlasmart-products-query-v1/_search
{
  "query": {
    "bool": {
      "filter": [
        { "term":  { "brand": "Auralux" } },
        { "terms": { "category": ["electronics/audio", "office/accessories"] } }
      ]
    }
  }
}
wrong · exact term lookup on analyzed visible phrase
GET atlasmart-products-query-v1/_search
{
  "query": {
    "term": { "name": "Wireless Noise Cancelling Headphones" }
  }
}

The wrong query asks whether that entire exact term exists in the analyzed name field. Use match for analyzed user text, or query name.raw if exact whole-value matching is truly required.

2. Range semantics belong to the mapped type

A numeric price range is a structured predicate: gte/gt and lte/lt define the boundary. Dates add parsing/time-zone concerns. Ranges on text/keyword can also be classified as expensive depending on product/settings; do not use lexical range tricks as a substitute for the right field type.

price window · binary filter
GET atlasmart-products-query-v1/_search
{
  "query": {
    "range": {
      "price": { "gte": 80, "lt": 150 }
    }
  },
  "sort": [ { "price": "asc" } ]
}

In the six-row fixture, p2 and p4 are the deterministic IDs inside that half-open interval. The test should assert those IDs and price boundaries, not a wall-clock duration from one laptop run.

3. exists asks whether an indexed value exists

exists is not the same as “the JSON key visually appears in _source.” Null values, empty arrays, mapping options such as ignore_above or malformed-value handling, and explicit null_value mappings can change whether a field contributes an indexed value. Empty strings can still count as values. Therefore define missing-data semantics at ingestion and mapping time rather than using exists as a generic data-quality oracle.

exists · indexed-value presence
GET atlasmart-products-query-v1/_search
{
  "query": {
    "exists": { "field": "rating" }
  }
}

4. IDs are direct identity lookups; business keys still need mappings

The ids query matches document _id values. It is useful when the application already knows document identities, but it does not replace a mapped sku, tenant ID or external business key. If AtlasMart must search, aggregate or sort on a business identifier, map that identifier explicitly.

ids · known document identities
GET atlasmart-products-query-v1/_search
{
  "query": {
    "ids": { "values": ["p1", "p4"] }
  }
}

5. Prefix, wildcard and regexp enumerate possible terms

Multi-term queries can expand one pattern into many indexed terms. Expansion depends on vocabulary size and pattern shape, not merely on the number of matching documents. A bounded SKU prefix such as AM-AU-* is qualitatively different from a leading wildcard such as *phones*. Leading wildcards and broad regular expressions can force expensive term enumeration. Elasticsearch and OpenSearch both expose search.allow_expensive_queries, but optimized mappings/settings and exact behavior differ.

bounded prefix · exact keyword vocabulary
GET atlasmart-products-query-v1/_search
{
  "query": {
    "prefix": {
      "sku": { "value": "AM-AU-" }
    }
  }
}
failure drill · disable expensive query classes
PUT _cluster/settings
{
  "transient": {
    "search.allow_expensive_queries": false
  }
}

GET atlasmart-products-query-v1/_search
{
  "query": {
    "wildcard": { "name.raw": "*head*" }
  }
}

PUT _cluster/settings
{
  "transient": {
    "search.allow_expensive_queries": null
  }
}

Run this only on the disposable course cluster and restore the setting. The exact error/allowance can depend on query type and optimized mapping settings, so record the product/version and the actual response rather than hard-coding one universal error string.

6. Cost-aware alternatives

Need Prefer Avoid
Exact category/brand/SKU keyword + term/terms match on identifiers
Known prefix/autocomplete purpose-built prefix/autocomplete mapping or bounded prefix query leading wildcard over a large vocabulary
Contains-style machine identifiers consider wildcard-oriented field/mapping where supported and benchmark general regexp on analyzed prose
Price/date window typed numeric/date range lexicographic text range
Known document IDs ids query scanning _source for IDs

Validate query length, wildcard/regex operator count, target fields and permitted range widths before building DSL. A timeout is not a resource-governance strategy: a costly query can consume resources before the client gives up.

Check your understanding

  1. Why can a term query miss the visible product name?
  2. What does exists actually test?
  3. Why are leading wildcards dangerous?
  4. Does search.allow_expensive_queries=false prove every remaining query is cheap?
  5. When should sku be queried with term rather than match?
Review the answers

1. The text field was analyzed into indexed terms; term does not run the full-text analyzer over the provided phrase.

2. Whether the field has an indexed value according to mapping/indexing semantics, not simply whether a JSON key visually appears in _source.

3. They can require enumeration across a large portion of the term dictionary and create high CPU/memory/latency cost.

4. No. It blocks selected query classes/paths; allowed queries can still be costly under large data, high concurrency, poor mappings or broad result sets.

5. When sku is mapped as keyword/exact data and the caller intends exact indexed-value equality.

Production judgment

Make exactness visible in the API contract. Structured filters should target typed/keyword fields and have bounded input domains. Pattern search deserves dedicated mappings and quotas, not an unrestricted regex field in a public request. Observe rejected queries, slow logs, top-query telemetry and per-tenant budgets before relaxing expensive-query controls.

Summary and next step

Term-level queries translate typed/exact intent into precise predicates; prefix/wildcard/regexp add vocabulary-expansion cost that must be bounded. Next, combine lexical scoring and binary filters correctly with Boolean Query DSL.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.