Chapter 06 · Query DSL Fundamentals: Term-Level, Full-Text, Boolean, Range, Exists, and Filtering

bool must/should/filter/must_not, Minimum Should Match, Filter Context, and Cacheability

Compose Boolean intent correctly with must, should, filter, must_not and minimum_should_match; separate scoring from binary eligibility and treat cacheability as an observed optimization rather than a promise.

Intermediate100–120 minutesDeterministic Query DSL evidence labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart's product endpoint must combine “how well does this match?” with “is this product eligible at all?”. A Boolean query is the boundary between those concerns. Misplacing a clause can change ranking, matching, cache eligibility and even the meaning of should. This lesson makes that logic explicit and testable.

01

Explain scoring query context versus binary filter context and place clauses according to product intent.

02

Compose must, should, filter and must_not without confusing Boolean eligibility with score contribution.

03

Predict the default minimum_should_match rule and set it explicitly when API behavior should not depend on surrounding clauses.

04

Use named clauses and deterministic hit sets to debug Boolean logic before tuning relevance.

05

Treat cacheability as an engine optimization measured under workload, not as a contractual guarantee that every filter is cached.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Translate requirements into two questions

For each AtlasMart criterion, ask: must this affect relevance? If yes, it belongs in query context such as must/should. If it is a binary eligibility rule such as inventory availability, category, tenant, price or policy status, place it in filter context when semantics allow. Filters skip relevance scoring and may benefit from caching/bitset reuse according to engine heuristics.

Requirement Clause Why
Text “noise cancelling” must: multi_match Lexical evidence should affect ranking
available=true filter: term Binary eligibility; no score needed
price <= 150 filter: range Binary structured constraint
prefer Auralux should: term on brand Optional preference can add score
exclude discontinued tag must_not: term Binary exclusion

2. One Boolean request, with clause names

bool · scored text plus explicit eligibility
GET atlasmart-products-query-v1/_search
{
  "query": {
    "bool": {
      "must": [
        { "multi_match": {
          "query": "noise cancelling",
          "fields": ["name^3", "description"],
          "_name": "lexical"
        }}
      ],
      "filter": [
        { "term":  { "available": { "value": true, "_name": "in_stock" } } },
        { "range": { "price": { "lte": 150, "_name": "budget" } } }
      ],
      "should": [
        { "term": { "brand": { "value": "Auralux", "_name": "preferred_brand" } } }
      ]
    }
  }
}

For the fixture, p4 is the deterministic eligible lexical hit: it matches “noise cancelling”, is available, and costs at most 150. The brand preference can affect score but is not required because a must/filter exists and no explicit minimum_should_match was set.

3. The should default changes with surrounding clauses

If a Boolean query has one or more should clauses and no must or filter, the default minimum_should_match is 1. If a must or filter clause exists, the default becomes 0. This is convenient for relevance boosts but dangerous when an API designer meant “at least one selected preference is required.” Make the requirement explicit.

explicit minimum_should_match · stable API contract
GET atlasmart-products-query-v1/_search
{
  "query": {
    "bool": {
      "filter": [ { "term": { "available": true } } ],
      "should": [
        { "term": { "brand": "Auralux" } },
        { "term": { "category": "office/accessories" } }
      ],
      "minimum_should_match": 1
    }
  }
}

Conditional and percentage forms are available, but complexity should follow a tested business rule. “75% sounds relevant” is not a requirement.

4. must_not is filter context, not negative scoring

must_not excludes documents that match its clauses; it does not merely decrease their score. If you need a soft penalty, that is a relevance-design problem for Chapter 07. Confusing exclusion with demotion can silently remove valid products.

exclusion · binary policy
GET atlasmart-products-query-v1/_search
{
  "query": {
    "bool": {
      "filter": { "term": { "available": true } },
      "must_not": { "term": { "category": "sports/running" } }
    }
  }
}

5. Cacheability is observable, not promised per clause

Filter context avoids score calculation and can enable reusable cached structures, but engines decide what to cache based on query shape, segment state, frequency and other heuristics. Do not write a design document that says “all filters are cached.” Use node/index query-cache statistics and representative repeated workloads if cache behavior matters. Separate the shard-level request cache from query/filter caching concepts, and do not benchmark with one warmed query then claim general p99 improvement.

Common failure

Moving a full-text match clause into filter preserves yes/no matching but discards its score contribution. Moving exact availability/category checks into must can compute scores that the application does not use. Correctness first; then measure performance.

6. Debug Boolean logic with names before score tuning

Name important clauses uniquely. For each expected hit, record which names matched and which eligibility rules were applied. This catches logic errors such as a should clause accidentally becoming optional after a new filter was added. If a query behaves unexpectedly, simplify it to one clause at a time before opening Profile.

Check your understanding

  1. When is filter context the better fit?
  2. What is the default minimum_should_match for only should clauses?
  3. What happens to that default after adding a must or filter?
  4. Does must_not lower a score?
  5. Why should cache behavior be measured rather than assumed?
Review the answers

1. When the criterion is binary eligibility and should not contribute to relevance scoring, such as availability, tenant/category or a price window.

2. 1 when there is at least one should and no must/filter clause.

3. It becomes 0, making should clauses optional unless minimum_should_match is set explicitly.

4. No. It excludes documents matching the clause in filter context.

5. Caching is heuristic and depends on query/segment/workload state; filter context makes reuse possible but does not guarantee every clause is cached.

Production judgment

Encode eligibility and relevance separately in application code. Make minimum_should_match explicit when business meaning depends on it. Give each externally visible filter a typed allowlist and bounds, and monitor clause growth because deeply nested Boolean trees and large term lists can become resource problems even if each leaf looks harmless.

Summary and next step

Boolean Query DSL is now a translation of product intent rather than a bag of clauses. Next, preserve object correlation with nested queries, add geographic constraints, and isolate scripts/expensive queries behind stronger guardrails.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.