Turn normalized security events into reproducible detections, findings, alerts, and investigations.

Security Analytics: Event Normalization, Detection Queries, Alert Volume, Long Retention, and Investigative Search

Show that product search, logs, security analytics, and observability require different schemas, shard/lifecycle/search patterns even when the same search engine can host them.

Intermediate → Advanced145–195 minutesSecurity detection & investigation lab · Chapter 28 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · ECS 9.5.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

01

Normalize security events so detections survive source-specific field names and log formats.

02

Distinguish raw events, detection rules, findings, alerts, and correlations instead of treating every match as an incident.

03

Control alert volume with rule scope, schedule, severity, grouping, and investigation context.

04

Separate security-analytics credentials and retention from general product/log search access.

05

Build a portable deterministic detection baseline and map it to Elastic/OpenSearch security features without assuming feature parity.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

Pinned platform baseline. Examples are reviewed against Elasticsearch/Kibana 9.5.3 (released 2026-09-03) and OpenSearch/OpenSearch Dashboards 3.8.0 (released 2026-08-04), using their bundled JVMs. The current Elastic Common Schema reference is ECS 9.5.0. The mandatory AtlasMart labs use free/local HTTP APIs and deterministic fixtures. The generation environment did not execute live clusters, so numeric latency, throughput, shard growth, and storage values shown as acceptance criteria are measurement instructions—not fabricated captured results.

1. AtlasMart problem: the same login event arrives in five vendor dialects

Security analytics starts with event normalization. A detection cannot be trusted if source.ip, user identity, action, outcome, process name, or file hash change field names by source. ECS helps Elastic normalize security data; OpenSearch Security Analytics ships detector workflows and maps supported log types to detection rules, including Sigma-based content. Neither removes the need to verify the actual event-to-rule mapping.

Define the layers: a raw event is observed activity; a rule encodes suspicious conditions; a finding is a rule match or detector result; an alert is a routed state that warrants attention; a correlation links multiple findings/events into a higher-level scenario. Treating each failed login as a page is an alert-volume design bug.

2. Portable security-event fixture

All Chapter 28 labs keep the existing local endpoints and security assumptions: Elasticsearch at https://localhost:9200 with ELASTIC_PASSWORD and the copied CA file atlasmart-es-http-ca; OpenSearch at https://localhost:9201 with OPENSEARCH_INITIAL_ADMIN_PASSWORD. OpenSearch's demo certificate trust bypass (-k) is acceptable only for this disposable local lab, never production. The shared Docker network remains atlasmart-search. Lab indices use one primary and zero replicas so a single-node workstation can complete the exercises; production redundancy decisions are deliberately separate.

Create the security-event index
PUT atlasmart-security-events-v1
{
  "settings":{"number_of_shards":1,"number_of_replicas":0},
  "mappings":{
    "dynamic":"strict",
    "properties":{
      "@timestamp":{"type":"date"},
      "event.category":{"type":"keyword"},
      "event.action":{"type":"keyword"},
      "event.outcome":{"type":"keyword"},
      "source.ip":{"type":"ip"},
      "user.name":{"type":"keyword"},
      "host.name":{"type":"keyword"},
      "rule.id":{"type":"keyword"},
      "tenant.id":{"type":"keyword"},
      "message":{"type":"text"}
    }
  }
}
Index a deterministic authentication sequence
POST _bulk?refresh=wait_for
{"index":{"_index":"atlasmart-security-events-v1","_id":"S-001"}}
{"@timestamp":"2026-09-12T12:10:00Z","event.category":"authentication","event.action":"login","event.outcome":"failure","source.ip":"203.0.113.50","user.name":"alex","host.name":"portal-1","tenant.id":"tenant-a","message":"Login failed"}
{"index":{"_index":"atlasmart-security-events-v1","_id":"S-002"}}
{"@timestamp":"2026-09-12T12:10:20Z","event.category":"authentication","event.action":"login","event.outcome":"failure","source.ip":"203.0.113.50","user.name":"alex","host.name":"portal-1","tenant.id":"tenant-a","message":"Login failed"}
{"index":{"_index":"atlasmart-security-events-v1","_id":"S-003"}}
{"@timestamp":"2026-09-12T12:10:40Z","event.category":"authentication","event.action":"login","event.outcome":"failure","source.ip":"203.0.113.50","user.name":"alex","host.name":"portal-1","tenant.id":"tenant-a","message":"Login failed"}
{"index":{"_index":"atlasmart-security-events-v1","_id":"S-004"}}
{"@timestamp":"2026-09-12T12:11:00Z","event.category":"authentication","event.action":"login","event.outcome":"success","source.ip":"203.0.113.50","user.name":"alex","host.name":"portal-1","tenant.id":"tenant-a","message":"Login succeeded"}

3. Portable detection baseline: prove the event evidence first

Count repeated failures from one source/user
POST atlasmart-security-events-v1/_search
{
  "size":0,
  "query":{
    "bool":{
      "filter":[
        {"term":{"tenant.id":"tenant-a"}},
        {"term":{"event.category":"authentication"}},
        {"term":{"event.action":"login"}},
        {"term":{"event.outcome":"failure"}},
        {"range":{"@timestamp":{"gte":"2026-09-12T12:10:00Z","lte":"2026-09-12T12:11:00Z"}}}
      ]
    }
  },
  "aggs":{
    "by_source_user":{
      "composite":{
        "size":100,
        "sources":[
          {"source":{"terms":{"field":"source.ip"}}},
          {"user":{"terms":{"field":"user.name"}}}
        ]
      }
    }
  }
}

The fixture should show three failures for the same source/user before a success. That is evidence, not an incident verdict. A production detector might include impossible travel, device history, MFA state, threat intelligence, or an account lockout signal. The portable baseline intentionally avoids pretending one simple threshold equals a mature detection.

4. OpenSearch Security Analytics: detector → finding → alert → correlation

OpenSearch 3.8 includes the Security Analytics plugin in the standard distribution. Current documentation describes detectors that apply rules to mapped log types, generate findings, optionally generate alerts, and correlate findings across log types. Sigma rules are a key content format, but the detector still depends on correct field mapping and permissions. Security Analytics system indexes are protected; users should receive scoped Security Analytics roles/actions rather than unrestricted cluster credentials.

For the lab, first validate the raw-query baseline above. Then, if the plugin is enabled in your local distribution, map the normalized fields and create a detector for the appropriate log type. The plugin result is an additional implementation surface—not the definition of the suspicious behavior itself.

5. Elastic Security: normalize first, then select solution-specific detections

Elastic Security similarly depends on normalized event fields and supports rule/detection workflows whose availability and advanced capabilities vary by deployment and subscription. Do not write the course as though OpenSearch Sigma detector JSON can be pasted into Elastic, or vice versa. Preserve the intent—data source, condition, schedule, severity, suppression, investigation fields, and response—then implement with the product’s current rule surface.

6. Alert-volume engineering

Control Question it answers Failure if ignored
Schedule/window How often and over what event window do we evaluate? Duplicate or missed detections.
Grouping Which entities form one alert instance? One page per event.
Severity How urgent is this scenario? Critical queue flooded by low-value findings.
Suppression/dedup When should repeats collapse? Pager fatigue and hidden real incidents.
Investigation context Which raw events and entities are attached? Alert cannot be verified.
Ownership Who acknowledges and tunes the rule? Stale detections with no operational response.

7. Retention and investigative search

Investigations often require long lookback windows and joins/correlations that differ from real-time alerting. Retention must follow legal, business, and forensic requirements, not a copied “30-day logs” policy. Long retention may justify lower-cost tiers, snapshots, or a separate security cluster. The important design variable is recovery/investigation time under real incident conditions.

Wrong approach. Give the detection engine the same broad credentials used by an interactive product-search application, and route every rule match directly to paging. Repair: isolate credentials and indexes, normalize fields, distinguish findings from alerts, add grouping/suppression/ownership, and retain the raw evidence needed to reproduce an alert.

8. Mini lab: portable finding plus optional platform detector

  1. Run the baseline failure query and store the source/user/count/window as a deterministic “finding.”
  2. Query the subsequent success event. Decide whether the combination should increase severity; document the rule logic.
  3. Create a reader role that can inspect atlasmart-security-events-v1 but cannot write/delete it. Use product-specific security APIs from Chapters 19–20.
  4. If using OpenSearch Security Analytics, map the fields and create a lab detector; compare its finding to the raw-query evidence.
  5. Measure how many alerts would result from 100 repeated failure events under no suppression versus grouped suppression. Report counts from your run, not invented values.

Check your understanding

  1. What is the difference between a finding and an alert?
  2. Why can a valid Sigma rule still fail?
  3. Should security analytics share product-search credentials?
  4. Why can security retention be longer than application-log retention?
  5. What is the bridge to Lesson 4?
Review the answers

1. A finding is detector/rule evidence; an alert is a routed operational state created under additional alerting conditions.

2. If source fields are not normalized/mapped to the fields expected by the rule, the intended condition may never match or may match incorrectly.

3. No. Security data and detector actions require their own least-privilege roles and audit boundary.

4. Forensics, legal obligations, and delayed investigations can require longer historical evidence.

5. Observability also correlates multiple event types, but its primary goal is service health and causality across logs, metrics, and traces rather than threat detection.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References and current-version checks

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.