Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering

Construct a Search API from User Intent with Safe Query Templates, Explicit Filters, and Test Cases

Turn product-search requirements into a validated request builder that emits fixed Query DSL structures, explicit filters and nested clauses, rejects unsafe inputs, and regression-tests expected hits and failure cases.

Intermediate100–120 minutesDeterministic Query DSL evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

The safest Query DSL is often the Query DSL the user never writes. AtlasMart's public endpoint needs a small, documented request model: free text, typed price bounds, availability, approved categories, and one optional nested attribute. Application code owns the DSL structure; the caller only supplies validated values. This lesson turns Chapters 03–06 into an executable API contract.

01

Design a typed public search contract that is smaller than the underlying Query DSL capability surface.

02

Build fixed DSL structures from validated values rather than concatenate raw JSON or query_string syntax.

03

Keep full-text clauses in query context and exact/range/nested eligibility clauses in filter context.

04

Write deterministic positive, boundary and adversarial tests for expected IDs and rejected inputs.

05

Capture shard/error/version evidence and separate semantic regression tests from relevance and performance benchmarks.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Define the public contract before the JSON body

AtlasMart exposes only inputs the product actually supports: q, min_price, max_price, available, a bounded category list, and an optional attribute name/value pair. It does not expose arbitrary field names, raw Query DSL, scripts, wildcard syntax, regex, sort scripts or arbitrary result windows. Internal/admin search can have a different authenticated contract.

Public input Validation Generated DSL
q 1..120 non-space chars multi_match over name^3 + description
price bounds finite numbers; min <= max range filter
available boolean only term filter
categories allowlist; max 10 unique terms filter
attribute pair both present; allowed name/value domain as product requires nested bool/filter
size server-owned fixed/default cap top-level size

2. Build objects, never concatenate query syntax

Python 3.11+ · pure request builder, no client dependency
from dataclasses import dataclass

@dataclass(frozen=True)
class SearchInput:
    q: str
    min_price: float | None = None
    max_price: float | None = None
    available: bool | None = None
    categories: tuple[str, ...] = ()
    attr_name: str | None = None
    attr_value: str | None = None

ALLOWED_CATEGORIES = {
    "electronics/audio", "office/accessories", "sports/running"
}

def build_query(x: SearchInput) -> dict:
    q = x.q.strip()
    if not q or len(q) > 120:
        raise ValueError("q must contain 1..120 non-space characters")

    filters: list[dict] = []
    if x.available is not None:
        filters.append({"term": {"available": x.available}})

    if x.categories:
        cats = tuple(dict.fromkeys(x.categories))
        if len(cats) > 10 or any(c not in ALLOWED_CATEGORIES for c in cats):
            raise ValueError("invalid category selection")
        filters.append({"terms": {"category": list(cats)}})

    if x.min_price is not None or x.max_price is not None:
        r = {}
        if x.min_price is not None: r["gte"] = x.min_price
        if x.max_price is not None: r["lte"] = x.max_price
        if "gte" in r and "lte" in r and r["gte"] > r["lte"]:
            raise ValueError("min_price must be <= max_price")
        filters.append({"range": {"price": r}})

    if (x.attr_name is None) != (x.attr_value is None):
        raise ValueError("attribute name and value must be supplied together")
    if x.attr_name is not None:
        filters.append({
            "nested": {
                "path": "attributes",
                "score_mode": "none",
                "query": {"bool": {"filter": [
                    {"term": {"attributes.name": x.attr_name}},
                    {"term": {"attributes.value": x.attr_value}},
                ]}}
            }
        })

    return {
        "size": 20,
        "track_total_hits": True,
        "query": {
            "bool": {
                "must": [{
                    "multi_match": {
                        "query": q,
                        "fields": ["name^3", "description"],
                        "type": "best_fields"
                    }
                }],
                "filter": filters
            }
        }
    }

The function returns a Python object that a maintained Elasticsearch/OpenSearch client or ordinary HTTPS layer can serialize. The key safety property is structural: user input occupies value positions only. If product requirements later add prefix search, geo radius or an expert syntax mode, add a new typed input with its own validation and tests rather than smuggling operators through q.

3. Positive regression cases use deterministic IDs

test matrix · semantic acceptance, not benchmark output
case A: q="noise cancelling", available=true, max_price=150
expect IDs: {p4}

case B: q="headphones", categories=("electronics/audio",)
expect IDs: {p1,p2}

case C: q="bluetooth", attr_name="connectivity", attr_value="bluetooth"
expect IDs: {p1,p3}

case D: q="laptop", min_price=40, max_price=60, available=true
expect IDs: {p5}

For every response also assert:
  timed_out == false
  _shards.failed == 0
  returned IDs equal the expected set
  product/server version recorded with the test run

Do not freeze numeric _score values as the primary contract. Ranking judgments belong in Chapter 07, and exact score values can shift with corpus/statistics/version changes even when the intended ordering remains acceptable.

4. Adversarial cases must fail before reaching the cluster

adversarial input suite
reject: q=""
reject: q of 121+ characters
reject: categories=("../../system-index",)
reject: 11 categories
reject: min_price=200, max_price=100
reject: attr_name="color" with no attr_value

accept as ordinary text, NOT operators:
  q="headphones OR *:*"
  q="name:(headphones)"
  q="/.*?/"

The fixed multi_match structure prevents these strings from becoming Query DSL syntax.

Length limits alone do not solve abusive search; also rate-limit requests, cap concurrency/tenant budgets, bound result size/pagination, restrict sort/facet choices, and monitor slow/top queries. But application validation removes a large class of accidental query-language injection before cluster controls are needed.

5. Optional expert syntax is a different endpoint/permission

If AtlasMart later needs an internal support endpoint with simple_query_string, document its operators and restrict fields, flags, maximum input length and expensive features. Do not silently switch the public endpoint to query_string because an internal user asks for wildcards. Product semantics, abuse risk and support burden all change.

restricted expert example · still requires policy
GET atlasmart-products-query-v1/_search
{
  "query": {
    "simple_query_string": {
      "query": ""noise cancelling" +Auralux",
      "fields": ["name^3", "description"],
      "default_operator": "and",
      "flags": "AND|OR|NOT|PHRASE|PRECEDENCE"
    }
  }
}

6. Release gate: semantics, relevance, performance and failure are separate

  • Semantic gate: expected IDs and filters/nested correlation are correct.
  • Relevance gate: judged query set meets ranking metrics/thresholds established in Chapter 07.
  • Performance gate: measure p50/p95/p99, throughput, shard fan-out, CPU/heap/cache state and concurrency on representative data; never infer from the six-document fixture.
  • Failure gate: timeouts, rejected/expensive queries, partial shard failures and dependency errors are surfaced rather than silently converted into an empty result set.
  • Compatibility gate: rerun the semantic/adversarial suite on each supported Elasticsearch/OpenSearch version before rollout.

Check your understanding

  1. Why not expose raw Query DSL to AtlasMart shoppers?
  2. How does the builder prevent query_string-style operator injection?
  3. Why are categories allowlisted?
  4. Why assert _shards.failed == 0 in regression tests?
  5. Why separate semantic, relevance and performance gates?
Review the answers

1. It gives untrusted callers a powerful cluster resource language, including expensive/complex query shapes and fields the product never intended to expose.

2. It constructs a fixed multi_match/bool structure and places q only in the query value, so operator-looking characters are analyzed text rather than query-language syntax.

3. To keep the public contract bounded, prevent arbitrary field/index-like values from becoming behavior, and make authorization/product semantics explicit.

4. A hit set from a partial search can look plausible while some shards failed; semantic acceptance requires a complete successful search for this fixture.

5. Correct matching does not prove good ranking or acceptable latency/resource use, and benchmark success does not prove the query means the right thing.

Production judgment

Search APIs are compilers from product intent to Query DSL. Keep the public language intentionally small, typed and versioned. Add new capabilities only with a mapping plan, bounded inputs, deterministic tests, resource budgets and observability. Do not let frontend convenience become an accidental promise that arbitrary Query DSL will remain supported forever.

Summary and next step

Chapter 06 now connects analysis, mappings and Query DSL into a safe request-construction discipline. Chapter 07 can focus on BM25, field weighting, business signals and relevance evaluation because the underlying match/filter semantics are already explicit and regression-tested.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.