Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering
Nested Queries, Geo Queries, Script Queries, and Avoiding Expensive Query Anti-Patterns
Preserve nested-object correlation, use geo predicates from mapped coordinates, isolate scripts as an expensive last resort, and design guardrails for nested, geo, script, wildcard and regexp abuse.
Learning outcomes
AtlasMart's API now needs “Bluetooth connectivity in the same variant,” “warehouse within 30 km,” and a one-off derived condition that the standard DSL cannot express. These needs are where incorrect object modeling, unbounded geo work and user-controlled scripts often enter production. This lesson keeps each mechanism isolated and measurable.
Use nested queries only on fields mapped as nested and preserve same-object correlation across multiple predicates.
Use geo_distance against geo_point fields with explicit distance units and validate user-supplied coordinates/radii.
Explain script query execution cost and prohibit arbitrary user-controlled script source/parameters by default.
Use search.allow_expensive_queries and application-level budgets as complementary guardrails, not interchangeable protections.
Diagnose partial/error responses and choose a cheaper mapping/query alternative before simply raising limits.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Nested queries preserve tuple correlation
Chapter 03 mapped attributes as
nested. Each nested object is indexed as a separate
hidden Lucene document associated with its parent, allowing one
nested query to require name and
value from the same attribute object. A plain
object array would flatten fields and can produce cross-object
false positives.
GET atlasmart-products-query-v1/_search
{
"query": {
"nested": {
"path": "attributes",
"query": {
"bool": {
"filter": [
{ "term": { "attributes.name": "connectivity" } },
{ "term": { "attributes.value": "bluetooth" } }
]
}
},
"score_mode": "none"
}
}
}
The deterministic matching IDs are p1 and
p3. A query requiring
attributes.name=color and
attributes.value=bluetooth should match none,
proving same-object correlation. Do not rewrite nested fields as
ordinary object fields just to simplify query JSON; that changes
correctness.
2. Geo queries are typed spatial predicates
geo_distance works on
geo_point mappings. The API should validate
latitude/longitude ranges, distance units and a maximum radius
before building DSL. A radius that expands from “nearby
warehouse” to a continent-sized scan is a product-policy change,
not just another number.
GET atlasmart-products-query-v1/_search
{
"query": {
"geo_distance": {
"distance": "30km",
"warehouse": { "lat": 40.7128, "lon": -74.0060 }
}
}
}
Given the teaching coordinates, p1,
p2 and p4 are the intended New
York-area hits. Exact distance calculations and boundary
inclusion must be verified if a business rule depends on a
precise radius.
3. Script queries execute code per candidate document
A script query can express a condition that is awkward in normal DSL, but it introduces per-document computation and a security/operational surface. Both current products support Painless-style scripting for this class of query. Use stored/fixed script logic or server-owned source with validated parameters; never allow an external caller to submit arbitrary script source.
GET atlasmart-products-query-v1/_search
{
"query": {
"script": {
"script": {
"source": "doc['rating'].size()!=0 && doc['price'].size()!=0 && doc['rating'].value >= params.min_rating && doc['price'].value <= params.max_price",
"params": { "min_rating": 4.5, "max_price": 150.0 }
}
}
}
}
Before accepting this in production, ask whether the same rule
can be precomputed at ingest time or expressed as ordinary
range/bool filters. The cheaper
representation is usually easier to secure, cache and reason
about.
4. Expensive-query controls are circuit breakers for query classes, not quotas
search.allow_expensive_queries=false can block
selected query families such as scripts, broad
wildcard/regexp/prefix paths and some other expensive patterns
depending on product/mapping. It does not enforce per-user CPU
budgets, maximum nested clauses, request body size, concurrency
or result-window discipline. Application validation still has to
constrain inputs before requests reach the cluster.
| Risk | Application guardrail | Cluster/engine evidence |
|---|---|---|
| Nested fan-out | Limit allowed nested filters/values | Profile only in diagnosis; monitor latency and shard work |
| Geo abuse | Validate coordinate/radius; cap radius | Slow/query telemetry and shard timing |
| Script abuse | No user script source; parameter allowlist | Expensive-query policy, script/cache metrics where exposed |
| Wildcard/regexp | Reject leading/broad patterns; length/operator caps | search.allow_expensive_queries, slow/query telemetry |
| Large Boolean trees | Clause/value count caps | Rejected queries, CPU/heap, tail latency |
5. Failure drill: wrong nested path and blocked script
Two failures are educational because they are deterministic and
reversible. First, issue a nested query against a
field that is not mapped as nested and inspect the error.
Second, on the disposable cluster only, disable expensive
queries and issue the script query, recording the actual
product/version error. Restore the setting immediately.
PUT _cluster/settings
{ "transient": { "search.allow_expensive_queries": false } }
# Re-run the script query here and record the actual response.
PUT _cluster/settings
{ "transient": { "search.allow_expensive_queries": null } }
Do not treat the failure text as a portable API contract. Treat the rejected/allowed state, query type, mapping and product version as evidence.
Check your understanding
- Why does AtlasMart need a nested query for attributes?
- What should be validated before building a geo_distance query?
- Why is a script query a last resort?
- What does search.allow_expensive_queries not provide?
- What should happen after an expensive-query failure in the lab?
Review the answers
1. Because the mapping is nested and the business rule requires multiple predicates to match the same attribute object, preserving tuple correlation.
2. Coordinate ranges, distance unit, maximum radius, target field and the user/tenant authorization for the spatial search.
3. It can execute per-document code, is expensive, harder to cache/reason about, and creates an additional security/operational surface.
4. It is not per-user quota, rate limiting, body-size control, concurrency control or a guarantee that allowed queries are cheap.
5. Record the actual response and product/version, diagnose the blocked query class, choose a cheaper representation when possible, and restore the temporary cluster setting.
Production judgment
Prefer mappings that make common queries cheap and explicit. Nested data is justified by same-object semantics, not by JSON aesthetics. Geo radii are product inputs with quotas. Scripts belong behind server-owned logic and monitoring. If a public endpoint can generate arbitrary Query DSL, you have effectively exposed a cluster resource language to untrusted callers.
Summary and next step
You can now preserve nested correctness, add typed geo constraints and recognize scripts/pattern queries as privileged resource consumers. The final lesson turns these pieces into a fixed request builder with adversarial tests.
Authoritative references
- Elastic: Query DSL overview — Current leaf/compound query, query/filter context, and expensive-query reference.
- Elastic: full-text queries — Current full-text query families and analyzed-input semantics.
- Elastic: term-level queries — Exact-value and structured-query semantics.
- Elastic: bool query — must/should/filter/must_not, minimum_should_match, and named-query behavior.
- Elastic: query and filter context — Scoring versus binary filtering semantics.
- OpenSearch: Query DSL — Current OpenSearch query categories and expensive-query controls.
- OpenSearch: bool query — Boolean clause semantics and minimum_should_match behavior.
- OpenSearch: wildcard query — Wildcard cost, rewrite behavior, and search.allow_expensive_queries interaction.
- OpenSearch: script query — Painless script-query behavior and cost warning.
- Elastic: nested query — Nested-object query semantics and score modes.
- Elastic: geo-distance query — geo_point distance filtering.
- OpenSearch: nested query — OpenSearch nested-query behavior.
- OpenSearch: geo-distance query — OpenSearch geo-distance syntax and semantics.