Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering
term/terms, range, exists, prefix/wildcard/regexp, IDs, and Cost Awareness
Use exact-term and structured predicates deliberately: distinguish term/terms/range/exists/IDs from analyzed full-text search, constrain prefix/wildcard/regexp expansion, and prove expensive-query behavior instead of guessing.
Learning outcomes
After AtlasMart can interpret natural-language text, it still needs exact SKU/brand/category filters, numeric price windows, existence checks and controlled prefix behavior. Those predicates should not be expressed as full-text guesses. This lesson separates exact indexed terms from analyzed text and makes multi-term expansion cost visible before it reaches production traffic.
Use term and terms for exact indexed values and explain why a term query on analyzed text often surprises users.
Use range and exists queries with field-type and missing-value semantics appropriate to numeric/date/keyword data.
Use IDs queries for document identity without pretending document IDs are a business secondary index.
Compare prefix, wildcard and regexp as term-enumeration tools and relate their cost to mappings, patterns and search.allow_expensive_queries.
Design validation and query-shape alternatives that avoid unbounded expansion while preserving user intent.
The examples use Query DSL constructs shared by Elasticsearch
9.5.3 and OpenSearch 3.8.0. Rewrite options, optimized field
types/settings and expensive-query exceptions differ over
time. Treat search.allow_expensive_queries as a
guardrail to verify, not as proof that every allowed query is
cheap.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. term and terms operate on indexed
terms
A term-level query does not run the field's full-text analyzer.
On keyword fields this is exactly what AtlasMart
needs for SKU, brand and category. On text, the
original human string may have been lowercased or split into
multiple indexed terms, so copying the original visible value
into a term query can miss unexpectedly.
GET atlasmart-products-query-v1/_search
{
"query": {
"bool": {
"filter": [
{ "term": { "brand": "Auralux" } },
{ "terms": { "category": ["electronics/audio", "office/accessories"] } }
]
}
}
}
GET atlasmart-products-query-v1/_search
{
"query": {
"term": { "name": "Wireless Noise Cancelling Headphones" }
}
}
The wrong query asks whether that entire exact term exists in
the analyzed name field. Use match for
analyzed user text, or query name.raw if exact
whole-value matching is truly required.
2. Range semantics belong to the mapped type
A numeric price range is a structured predicate:
gte/gt and lte/lt
define the boundary. Dates add parsing/time-zone concerns.
Ranges on text/keyword can also be
classified as expensive depending on product/settings; do not
use lexical range tricks as a substitute for the right field
type.
GET atlasmart-products-query-v1/_search
{
"query": {
"range": {
"price": { "gte": 80, "lt": 150 }
}
},
"sort": [ { "price": "asc" } ]
}
In the six-row fixture, p2 and p4 are
the deterministic IDs inside that half-open interval. The test
should assert those IDs and price boundaries, not a wall-clock
duration from one laptop run.
3. exists asks whether an indexed value exists
exists is not the same as “the JSON key visually
appears in _source.” Null values, empty arrays,
mapping options such as ignore_above or
malformed-value handling, and explicit
null_value mappings can change whether a field
contributes an indexed value. Empty strings can still count as
values. Therefore define missing-data semantics at ingestion and
mapping time rather than using exists as a generic
data-quality oracle.
GET atlasmart-products-query-v1/_search
{
"query": {
"exists": { "field": "rating" }
}
}
4. IDs are direct identity lookups; business keys still need mappings
The ids query matches document
_id values. It is useful when the application
already knows document identities, but it does not replace a
mapped sku, tenant ID or external business key. If
AtlasMart must search, aggregate or sort on a business
identifier, map that identifier explicitly.
GET atlasmart-products-query-v1/_search
{
"query": {
"ids": { "values": ["p1", "p4"] }
}
}
5. Prefix, wildcard and regexp enumerate possible terms
Multi-term queries can expand one pattern into many indexed
terms. Expansion depends on vocabulary size and pattern shape,
not merely on the number of matching documents. A bounded SKU
prefix such as AM-AU-* is qualitatively different
from a leading wildcard such as *phones*. Leading
wildcards and broad regular expressions can force expensive term
enumeration. Elasticsearch and OpenSearch both expose
search.allow_expensive_queries, but optimized
mappings/settings and exact behavior differ.
GET atlasmart-products-query-v1/_search
{
"query": {
"prefix": {
"sku": { "value": "AM-AU-" }
}
}
}
PUT _cluster/settings
{
"transient": {
"search.allow_expensive_queries": false
}
}
GET atlasmart-products-query-v1/_search
{
"query": {
"wildcard": { "name.raw": "*head*" }
}
}
PUT _cluster/settings
{
"transient": {
"search.allow_expensive_queries": null
}
}
Run this only on the disposable course cluster and restore the setting. The exact error/allowance can depend on query type and optimized mapping settings, so record the product/version and the actual response rather than hard-coding one universal error string.
6. Cost-aware alternatives
| Need | Prefer | Avoid |
|---|---|---|
| Exact category/brand/SKU | keyword + term/terms | match on identifiers |
| Known prefix/autocomplete | purpose-built prefix/autocomplete mapping or bounded prefix query | leading wildcard over a large vocabulary |
| Contains-style machine identifiers | consider wildcard-oriented field/mapping where supported and benchmark | general regexp on analyzed prose |
| Price/date window | typed numeric/date range | lexicographic text range |
| Known document IDs | ids query | scanning _source for IDs |
Validate query length, wildcard/regex operator count, target fields and permitted range widths before building DSL. A timeout is not a resource-governance strategy: a costly query can consume resources before the client gives up.
Check your understanding
- Why can a term query miss the visible product name?
- What does exists actually test?
- Why are leading wildcards dangerous?
- Does search.allow_expensive_queries=false prove every remaining query is cheap?
- When should sku be queried with term rather than match?
Review the answers
1. The text field was analyzed into indexed terms; term does not run the full-text analyzer over the provided phrase.
2. Whether the field has an indexed value according to mapping/indexing semantics, not simply whether a JSON key visually appears in _source.
3. They can require enumeration across a large portion of the term dictionary and create high CPU/memory/latency cost.
4. No. It blocks selected query classes/paths; allowed queries can still be costly under large data, high concurrency, poor mappings or broad result sets.
5. When sku is mapped as keyword/exact data and the caller intends exact indexed-value equality.
Production judgment
Make exactness visible in the API contract. Structured filters should target typed/keyword fields and have bounded input domains. Pattern search deserves dedicated mappings and quotas, not an unrestricted regex field in a public request. Observe rejected queries, slow logs, top-query telemetry and per-tenant budgets before relaxing expensive-query controls.
Summary and next step
Term-level queries translate typed/exact intent into precise predicates; prefix/wildcard/regexp add vocabulary-expansion cost that must be bounded. Next, combine lexical scoring and binary filters correctly with Boolean Query DSL.
Authoritative references
- Elastic: Query DSL overview — Current leaf/compound query, query/filter context, and expensive-query reference.
- Elastic: full-text queries — Current full-text query families and analyzed-input semantics.
- Elastic: term-level queries — Exact-value and structured-query semantics.
- Elastic: bool query — must/should/filter/must_not, minimum_should_match, and named-query behavior.
- Elastic: query and filter context — Scoring versus binary filtering semantics.
- OpenSearch: Query DSL — Current OpenSearch query categories and expensive-query controls.
- OpenSearch: bool query — Boolean clause semantics and minimum_should_match behavior.
- OpenSearch: wildcard query — Wildcard cost, rewrite behavior, and search.allow_expensive_queries interaction.
- OpenSearch: script query — Painless script-query behavior and cost warning.
- Elastic: wildcard query — Wildcard pattern and expensive-query behavior.
- OpenSearch: regexp query — Regexp automaton limits and expensive-query warning.
- OpenSearch: prefix query — Prefix semantics, rewrites and optimized index_prefixes path.