Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering
bool must/should/filter/must_not, Minimum Should Match, Filter Context, and Cacheability
Compose Boolean intent correctly with must, should, filter, must_not and minimum_should_match; separate scoring from binary eligibility and treat cacheability as an observed optimization rather than a promise.
Learning outcomes
AtlasMart's product endpoint must combine “how well does this
match?” with “is this product eligible at all?”. A Boolean query
is the boundary between those concerns. Misplacing a clause can
change ranking, matching, cache eligibility and even the meaning
of should. This lesson makes that logic explicit
and testable.
Explain scoring query context versus binary filter context and place clauses according to product intent.
Compose must, should, filter and must_not without confusing Boolean eligibility with score contribution.
Predict the default minimum_should_match rule and set it explicitly when API behavior should not depend on surrounding clauses.
Use named clauses and deterministic hit sets to debug Boolean logic before tuning relevance.
Treat cacheability as an engine optimization measured under workload, not as a contractual guarantee that every filter is cached.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Translate requirements into two questions
For each AtlasMart criterion, ask:
must this affect relevance? If yes, it belongs
in query context such as must/should.
If it is a binary eligibility rule such as inventory
availability, category, tenant, price or policy status, place it
in filter context when semantics allow. Filters skip relevance
scoring and may benefit from caching/bitset reuse according to
engine heuristics.
| Requirement | Clause | Why |
|---|---|---|
| Text “noise cancelling” | must: multi_match | Lexical evidence should affect ranking |
| available=true | filter: term | Binary eligibility; no score needed |
| price <= 150 | filter: range | Binary structured constraint |
| prefer Auralux | should: term on brand | Optional preference can add score |
| exclude discontinued tag | must_not: term | Binary exclusion |
2. One Boolean request, with clause names
GET atlasmart-products-query-v1/_search
{
"query": {
"bool": {
"must": [
{ "multi_match": {
"query": "noise cancelling",
"fields": ["name^3", "description"],
"_name": "lexical"
}}
],
"filter": [
{ "term": { "available": { "value": true, "_name": "in_stock" } } },
{ "range": { "price": { "lte": 150, "_name": "budget" } } }
],
"should": [
{ "term": { "brand": { "value": "Auralux", "_name": "preferred_brand" } } }
]
}
}
}
For the fixture, p4 is the deterministic eligible
lexical hit: it matches “noise cancelling”, is available, and
costs at most 150. The brand preference can affect score but is
not required because a must/filter
exists and no explicit minimum_should_match was
set.
3. The should default changes with surrounding
clauses
If a Boolean query has one or more should clauses
and no must or filter, the default
minimum_should_match is 1. If a
must or filter clause exists, the
default becomes 0. This is convenient for relevance boosts but
dangerous when an API designer meant “at least one selected
preference is required.” Make the requirement explicit.
GET atlasmart-products-query-v1/_search
{
"query": {
"bool": {
"filter": [ { "term": { "available": true } } ],
"should": [
{ "term": { "brand": "Auralux" } },
{ "term": { "category": "office/accessories" } }
],
"minimum_should_match": 1
}
}
}
Conditional and percentage forms are available, but complexity should follow a tested business rule. “75% sounds relevant” is not a requirement.
4. must_not is filter context, not negative scoring
must_not excludes documents that match its clauses;
it does not merely decrease their score. If you need a soft
penalty, that is a relevance-design problem for Chapter 07.
Confusing exclusion with demotion can silently remove valid
products.
GET atlasmart-products-query-v1/_search
{
"query": {
"bool": {
"filter": { "term": { "available": true } },
"must_not": { "term": { "category": "sports/running" } }
}
}
}
5. Cacheability is observable, not promised per clause
Filter context avoids score calculation and can enable reusable cached structures, but engines decide what to cache based on query shape, segment state, frequency and other heuristics. Do not write a design document that says “all filters are cached.” Use node/index query-cache statistics and representative repeated workloads if cache behavior matters. Separate the shard-level request cache from query/filter caching concepts, and do not benchmark with one warmed query then claim general p99 improvement.
Moving a full-text match clause into
filter preserves yes/no matching but discards its
score contribution. Moving exact availability/category checks
into must can compute scores that the application
does not use. Correctness first; then measure performance.
6. Debug Boolean logic with names before score tuning
Name important clauses uniquely. For each expected hit, record
which names matched and which eligibility rules were applied.
This catches logic errors such as a should clause
accidentally becoming optional after a new filter was added. If
a query behaves unexpectedly, simplify it to one clause at a
time before opening Profile.
Check your understanding
- When is filter context the better fit?
- What is the default minimum_should_match for only should clauses?
- What happens to that default after adding a must or filter?
- Does must_not lower a score?
- Why should cache behavior be measured rather than assumed?
Review the answers
1. When the criterion is binary eligibility and should not contribute to relevance scoring, such as availability, tenant/category or a price window.
2. 1 when there is at least one should and no must/filter clause.
3. It becomes 0, making should clauses optional unless minimum_should_match is set explicitly.
4. No. It excludes documents matching the clause in filter context.
5. Caching is heuristic and depends on query/segment/workload state; filter context makes reuse possible but does not guarantee every clause is cached.
Production judgment
Encode eligibility and relevance separately in application code.
Make minimum_should_match explicit when business
meaning depends on it. Give each externally visible filter a
typed allowlist and bounds, and monitor clause growth because
deeply nested Boolean trees and large term lists can become
resource problems even if each leaf looks harmless.
Summary and next step
Boolean Query DSL is now a translation of product intent rather than a bag of clauses. Next, preserve object correlation with nested queries, add geographic constraints, and isolate scripts/expensive queries behind stronger guardrails.
Authoritative references
- Elastic: Query DSL overview — Current leaf/compound query, query/filter context, and expensive-query reference.
- Elastic: full-text queries — Current full-text query families and analyzed-input semantics.
- Elastic: term-level queries — Exact-value and structured-query semantics.
- Elastic: bool query — must/should/filter/must_not, minimum_should_match, and named-query behavior.
- Elastic: query and filter context — Scoring versus binary filtering semantics.
- OpenSearch: Query DSL — Current OpenSearch query categories and expensive-query controls.
- OpenSearch: bool query — Boolean clause semantics and minimum_should_match behavior.
- OpenSearch: wildcard query — Wildcard cost, rewrite behavior, and search.allow_expensive_queries interaction.
- OpenSearch: script query — Painless script-query behavior and cost warning.