Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering
Construct a Search API from User Intent with Safe Query Templates, Explicit Filters, and Test Cases
Turn product-search requirements into a validated request builder that emits fixed Query DSL structures, explicit filters and nested clauses, rejects unsafe inputs, and regression-tests expected hits and failure cases.
Learning outcomes
The safest Query DSL is often the Query DSL the user never writes. AtlasMart's public endpoint needs a small, documented request model: free text, typed price bounds, availability, approved categories, and one optional nested attribute. Application code owns the DSL structure; the caller only supplies validated values. This lesson turns Chapters 03–06 into an executable API contract.
Design a typed public search contract that is smaller than the underlying Query DSL capability surface.
Build fixed DSL structures from validated values rather than concatenate raw JSON or query_string syntax.
Keep full-text clauses in query context and exact/range/nested eligibility clauses in filter context.
Write deterministic positive, boundary and adversarial tests for expected IDs and rejected inputs.
Capture shard/error/version evidence and separate semantic regression tests from relevance and performance benchmarks.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Define the public contract before the JSON body
AtlasMart exposes only inputs the product actually supports:
q, min_price, max_price,
available, a bounded category list, and an optional
attribute name/value pair. It does not expose arbitrary field
names, raw Query DSL, scripts, wildcard syntax, regex, sort
scripts or arbitrary result windows. Internal/admin search can
have a different authenticated contract.
| Public input | Validation | Generated DSL |
|---|---|---|
| q | 1..120 non-space chars | multi_match over name^3 + description |
| price bounds | finite numbers; min <= max | range filter |
| available | boolean only | term filter |
| categories | allowlist; max 10 unique | terms filter |
| attribute pair | both present; allowed name/value domain as product requires | nested bool/filter |
| size | server-owned fixed/default cap | top-level size |
2. Build objects, never concatenate query syntax
from dataclasses import dataclass
@dataclass(frozen=True)
class SearchInput:
q: str
min_price: float | None = None
max_price: float | None = None
available: bool | None = None
categories: tuple[str, ...] = ()
attr_name: str | None = None
attr_value: str | None = None
ALLOWED_CATEGORIES = {
"electronics/audio", "office/accessories", "sports/running"
}
def build_query(x: SearchInput) -> dict:
q = x.q.strip()
if not q or len(q) > 120:
raise ValueError("q must contain 1..120 non-space characters")
filters: list[dict] = []
if x.available is not None:
filters.append({"term": {"available": x.available}})
if x.categories:
cats = tuple(dict.fromkeys(x.categories))
if len(cats) > 10 or any(c not in ALLOWED_CATEGORIES for c in cats):
raise ValueError("invalid category selection")
filters.append({"terms": {"category": list(cats)}})
if x.min_price is not None or x.max_price is not None:
r = {}
if x.min_price is not None: r["gte"] = x.min_price
if x.max_price is not None: r["lte"] = x.max_price
if "gte" in r and "lte" in r and r["gte"] > r["lte"]:
raise ValueError("min_price must be <= max_price")
filters.append({"range": {"price": r}})
if (x.attr_name is None) != (x.attr_value is None):
raise ValueError("attribute name and value must be supplied together")
if x.attr_name is not None:
filters.append({
"nested": {
"path": "attributes",
"score_mode": "none",
"query": {"bool": {"filter": [
{"term": {"attributes.name": x.attr_name}},
{"term": {"attributes.value": x.attr_value}},
]}}
}
})
return {
"size": 20,
"track_total_hits": True,
"query": {
"bool": {
"must": [{
"multi_match": {
"query": q,
"fields": ["name^3", "description"],
"type": "best_fields"
}
}],
"filter": filters
}
}
}
The function returns a Python object that a maintained
Elasticsearch/OpenSearch client or ordinary HTTPS layer can
serialize. The key safety property is structural: user input
occupies value positions only. If product requirements later add
prefix search, geo radius or an expert syntax mode, add a new
typed input with its own validation and tests rather than
smuggling operators through q.
3. Positive regression cases use deterministic IDs
case A: q="noise cancelling", available=true, max_price=150
expect IDs: {p4}
case B: q="headphones", categories=("electronics/audio",)
expect IDs: {p1,p2}
case C: q="bluetooth", attr_name="connectivity", attr_value="bluetooth"
expect IDs: {p1,p3}
case D: q="laptop", min_price=40, max_price=60, available=true
expect IDs: {p5}
For every response also assert:
timed_out == false
_shards.failed == 0
returned IDs equal the expected set
product/server version recorded with the test run
Do not freeze numeric _score values as the primary
contract. Ranking judgments belong in Chapter 07, and exact
score values can shift with corpus/statistics/version changes
even when the intended ordering remains acceptable.
4. Adversarial cases must fail before reaching the cluster
reject: q=""
reject: q of 121+ characters
reject: categories=("../../system-index",)
reject: 11 categories
reject: min_price=200, max_price=100
reject: attr_name="color" with no attr_value
accept as ordinary text, NOT operators:
q="headphones OR *:*"
q="name:(headphones)"
q="/.*?/"
The fixed multi_match structure prevents these strings from becoming Query DSL syntax.
Length limits alone do not solve abusive search; also rate-limit requests, cap concurrency/tenant budgets, bound result size/pagination, restrict sort/facet choices, and monitor slow/top queries. But application validation removes a large class of accidental query-language injection before cluster controls are needed.
5. Optional expert syntax is a different endpoint/permission
If AtlasMart later needs an internal support endpoint with
simple_query_string, document its operators and
restrict fields, flags, maximum input
length and expensive features. Do not silently switch the public
endpoint to query_string because an internal user
asks for wildcards. Product semantics, abuse risk and support
burden all change.
GET atlasmart-products-query-v1/_search
{
"query": {
"simple_query_string": {
"query": ""noise cancelling" +Auralux",
"fields": ["name^3", "description"],
"default_operator": "and",
"flags": "AND|OR|NOT|PHRASE|PRECEDENCE"
}
}
}
6. Release gate: semantics, relevance, performance and failure are separate
- Semantic gate: expected IDs and filters/nested correlation are correct.
- Relevance gate: judged query set meets ranking metrics/thresholds established in Chapter 07.
- Performance gate: measure p50/p95/p99, throughput, shard fan-out, CPU/heap/cache state and concurrency on representative data; never infer from the six-document fixture.
- Failure gate: timeouts, rejected/expensive queries, partial shard failures and dependency errors are surfaced rather than silently converted into an empty result set.
- Compatibility gate: rerun the semantic/adversarial suite on each supported Elasticsearch/OpenSearch version before rollout.
Check your understanding
- Why not expose raw Query DSL to AtlasMart shoppers?
- How does the builder prevent query_string-style operator injection?
- Why are categories allowlisted?
- Why assert _shards.failed == 0 in regression tests?
- Why separate semantic, relevance and performance gates?
Review the answers
1. It gives untrusted callers a powerful cluster resource language, including expensive/complex query shapes and fields the product never intended to expose.
2. It constructs a fixed multi_match/bool structure and places q only in the query value, so operator-looking characters are analyzed text rather than query-language syntax.
3. To keep the public contract bounded, prevent arbitrary field/index-like values from becoming behavior, and make authorization/product semantics explicit.
4. A hit set from a partial search can look plausible while some shards failed; semantic acceptance requires a complete successful search for this fixture.
5. Correct matching does not prove good ranking or acceptable latency/resource use, and benchmark success does not prove the query means the right thing.
Production judgment
Search APIs are compilers from product intent to Query DSL. Keep the public language intentionally small, typed and versioned. Add new capabilities only with a mapping plan, bounded inputs, deterministic tests, resource budgets and observability. Do not let frontend convenience become an accidental promise that arbitrary Query DSL will remain supported forever.
Summary and next step
Chapter 06 now connects analysis, mappings and Query DSL into a safe request-construction discipline. Chapter 07 can focus on BM25, field weighting, business signals and relevance evaluation because the underlying match/filter semantics are already explicit and regression-tested.
Authoritative references
- Elastic: Query DSL overview — Current leaf/compound query, query/filter context, and expensive-query reference.
- Elastic: full-text queries — Current full-text query families and analyzed-input semantics.
- Elastic: term-level queries — Exact-value and structured-query semantics.
- Elastic: bool query — must/should/filter/must_not, minimum_should_match, and named-query behavior.
- Elastic: query and filter context — Scoring versus binary filtering semantics.
- OpenSearch: Query DSL — Current OpenSearch query categories and expensive-query controls.
- OpenSearch: bool query — Boolean clause semantics and minimum_should_match behavior.
- OpenSearch: wildcard query — Wildcard cost, rewrite behavior, and search.allow_expensive_queries interaction.
- OpenSearch: script query — Painless script-query behavior and cost warning.
- Elastic: simple_query_string query — Restricted fault-tolerant query syntax for intentionally exposed expert search.
- Elastic: multi_match query — Explicit multi-field lexical request construction.
- OpenSearch: query_string query — OpenSearch expert query syntax and wildcard/regex cost considerations.