Chapter 11 · Ingest Pipelines, Processors, Enrichment, Failure Handling, and Data Quality
Grok/Dissect, Date, Convert, Rename, Set, Remove, Script, GeoIP/User-Agent-Like Enrichment Patterns
Choose parsing and transformation processors from input structure and cost: contrast dissect with grok, normalize dates/types/names, use scripts sparingly, and treat GeoIP/user-agent enrichment as versioned dependencies rather than magic.
Learning outcomes
AtlasMart receives several log shapes: fixed delimiter access lines, legacy free-form messages, timestamps in old formats, string-encoded numbers, browser user-agent strings and client IPs. The correct processor is determined by the input contract. This lesson deliberately avoids the anti-pattern of using an expensive regex parser for every event.
Choose dissect for stable delimiter layouts and grok when pattern matching is genuinely required.
Normalize timestamp and numeric types before mapping validation makes failures harder to diagnose.
Use rename/set/remove for schema ownership and scripts only when built-in processors are insufficient.
Treat GeoIP and user-agent parsing as enrichment dependencies whose databases/rules can change.
Benchmark processor mixes and keep the raw evidence needed for replay and parser-regression tests.
The reproducible examples target self-managed Elasticsearch 9.5.3 and OpenSearch 3.8.0 using the established AtlasMart lab conventions: Elasticsearch on https://localhost:9200 with the copied CA certificate, OpenSearch on https://localhost:9201 with the disposable demo certificate explicitly treated as local-only, pinned server versions, and no moving latest tags. OpenSearch 3.8 currently documents grok, dissect, date, convert, rename, set, remove, script, geoip, ip2geo and user_agent among its core processors. Elastic exposes analogous common processors plus its own evolving catalog. Exact parameters and enrichment databases remain product/version dependent.
The generation environment does not run the two search servers. Commands are reviewed deterministic lab specifications and expected invariants, not fabricated captured output. Execute them against disposable AtlasMart resources and record your own processor timings, CPU, throughput, error counts, response bodies, and security behavior before making production decisions.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Dissect and grok solve different parsing problems
| Processor | Best fit | Risk signal |
|---|---|---|
| dissect | Stable delimiter or positional text | Format drift breaks extraction; cheap structure should be tested explicitly. |
| grok | Known text patterns that require regex-like recognition | Pattern complexity and backtracking can make ingest CPU unpredictable. |
| date | Convert known timestamp representations to a date field | Ambiguous formats/time zones can produce semantic errors even when parsing succeeds. |
| convert | Create numeric/boolean/IP-compatible source values | Conversion does not rescue a fundamentally wrong schema. |
| script | Logic unavailable in built-in processors | More code, compilation/cache considerations and larger security/performance review surface. |
Dissect works from delimiter structure. Grok matches named patterns. If your line is always nine pipe-separated fields, use dissect; adding a broad grok pattern “just in case” increases maintenance and cost without increasing correctness.
2. Compare token-by-token parser evidence
POST _ingest/pipeline/_simulate?verbose=true
{
"pipeline":{"processors":[
{"dissect":{"field":"message","pattern":"%{ts}|%{service}|%{level}|%{duration}"}},
{"convert":{"field":"duration","type":"double"}}
]},
"docs":[{"_source":{"message":"2026-09-11T07:30:00Z|catalog-api|INFO|12.4"}}]
}
POST _ingest/pipeline/_simulate?verbose=true
{
"pipeline":{"processors":[
{"grok":{"field":"message","patterns":["%{TIMESTAMP_ISO8601:ts} %{WORD:level} %{NUMBER:duration:float}ms %{GREEDYDATA:text}"]}}
]},
"docs":[{"_source":{"message":"2026-09-11T07:30:00Z WARN 12.4ms cache miss"}}]
}
Run malformed boundary cases too: missing delimiter, embedded delimiter, decimal with comma, impossible date, huge line, unexpected Unicode and empty field. A parser is not validated by one happy-path sample.
3. Normalize into the Chapter 10 telemetry contract
PUT _ingest/pipeline/atlasmart-normalize-v1
{
"processors":[
{"rename":{"field":"ts","target_field":"event_time","ignore_missing":true}},
{"date":{"field":"event_time","target_field":"@timestamp","formats":["ISO8601","yyyy-MM-dd HH:mm:ss"]}},
{"convert":{"field":"status_code","type":"integer","ignore_missing":true}},
{"convert":{"field":"duration_ms","type":"double","ignore_missing":true}},
{"set":{"field":"schema.version","value":"telemetry-v1"}},
{"remove":{"field":["event_time","debug_blob"],"ignore_missing":true}}
]
}
Normalization should end in one governed target schema. Duplicating equivalent transformations in the producer, shipper, Data Prepper/Logstash and ingest pipeline creates disagreement during incident response because two paths may produce different field names or types.
4. Scripts are an escape hatch, not the default language
POST _ingest/pipeline/_simulate
{
"pipeline":{"processors":[
{"script":{
"tag":"latency-band",
"lang":"painless",
"source":"if (ctx.duration_ms != null) { ctx.latency_band = ctx.duration_ms >= 500 ? 'slow' : 'normal'; }"
}}
]},
"docs":[{"_source":{"duration_ms":512.0}}]
}
This script is deterministic and bounded, but the same logic may be cheaper and easier to own upstream. Never do network calls from scripts; do not turn ingest into a general ETL runtime. Keep parameters explicit, avoid unbounded loops, and watch script compilations/cache metrics when scripts are unavoidable.
5. GeoIP and user-agent enrichment are versioned reference transformations
GET _nodes/ingest?filter_path=nodes.*.ingest.processors
POST _ingest/pipeline/_simulate
{
"pipeline":{"processors":[
{"geoip":{"field":"client_ip","target_field":"client_geo","ignore_missing":true}},
{"user_agent":{"field":"user_agent","target_field":"client_user_agent","ignore_missing":true}}
]},
"docs":[{"_source":{
"client_ip":"8.8.8.8",
"user_agent":"Mozilla/5.0"
}}]
}
GeoIP output depends on the installed/reference database; user-agent parsing depends on parser rules. These are not immutable truths. Store an enrichment version or deployment version when reproducibility matters, and do not use geolocation derived from IP as a security authorization decision.
Client IPs and user-agent strings can be personal/security-sensitive telemetry. Minimize collection, redact secrets and session identifiers before indexing, scope read access, and set retention from legal/business requirements rather than technical convenience.
6. Deliberately wrong approach: one giant grok for everything
A broad grok pattern that attempts every source format often becomes an expensive, fragile hidden programming language. A single new log format can create CPU spikes or mis-parsed fields. Repair by classifying sources first, route by a trusted source/service field, use simple dissect pipelines for stable formats, reserve grok for bounded pattern families, and attach deterministic fixtures to each parser.
7. Processor cost is measured as a pipeline
GET _nodes/stats/ingest?filter_path=nodes.*.ingest
# Run the same fixed fixture N times with:
# A) direct index, no pipeline
# B) dissect + date + convert
# C) grok + date + convert
# D) parser + geoip + user_agent
# Record client elapsed time, docs/s, node CPU, ingest time delta,
# processor time delta, failures, indexing lag and output correctness.
Do not publish the fastest one-run number. Warm up the JVM, keep the fixture and refresh policy identical, repeat trials, report distribution/variance and make correctness a prerequisite before throughput.
Check your understanding
- When is dissect preferable to grok?
- Why convert before relying on the destination mapping?
- What is the primary risk of scripts?
- Why version GeoIP/user-agent enrichment?
- What makes a parser regression test useful?
Review the answers
1. When the source has a stable delimiter/positional structure and regex recognition is unnecessary.
2. It makes type normalization explicit and lets parser/conversion failures be observed at the transformation boundary.
3. They add code and CPU/security/maintenance surface that built-in processors may avoid.
4. Reference databases and parsing rules evolve, so the same raw event may enrich differently later.
5. It includes happy paths and malformed/boundary fixtures with exact expected transformed fields and failure classes.
Production judgment
Pick processors from data structure, not familiarity. Watch p95/p99 ingest latency, processor-time deltas, CPU saturation, failure rate and quarantine rate by source. If parsing consumes material cluster headroom or requires joins, state, fan-out or complex branching, move it upstream to a purpose-built stream/ETL layer and keep the search cluster focused on indexing and retrieval.
Summary and next step
You can now choose and validate common parsing/normalization/enrichment processors. Next we focus on lookup enrichment freshness: Elastic enrich policies, their snapshot-like execution model, and the OpenSearch alternatives that must not be mislabeled as the same feature.
Authoritative references
- Elastic ingest pipelines — Pipeline execution, conditionals, failure handling, node statistics, default/final pipeline concepts and operational limitations.
- Elastic simulate pipeline API — Test an existing or inline pipeline; verbose mode exposes per-processor intermediate results.
- Elastic ingest error handling — Processor-level and pipeline-level on_failure semantics and ingest failure metadata.
- Elastic enrich processor setup — Source data, enrich policy execution, generated enrich index and ingest processor workflow.
- Elastic enrich processor reference — policy_name, match field, target field, max_matches and lookup behavior.
- Elastic script processor — Painless ingest context and script processor behavior.
- OpenSearch ingest pipelines — Pipeline model, ingest node prerequisite and guidance on using Data Prepper for larger or more complex preprocessing.
- OpenSearch ingest processors — Current core processor inventory and Nodes Info inspection endpoint.
- OpenSearch pipeline failures — Failure handling and ingest-pipeline node statistics.
- OpenSearch access data in a pipeline — Source, metadata such as _index/_routing, and _ingest.timestamp access.
- OpenSearch grok processor — Pattern-based parsing and debug controls.
- OpenSearch dissect processor — Delimiter-based parsing for stable log formats.
- OpenSearch user-agent processor — User-agent parsing and target field configuration.
- OpenSearch Nodes Stats API — Per-node, pipeline and processor ingest counts, time, current work and failures.