Chapter 02 · Requirements Engineering: Business Processes, Questions, Metrics, Dimensions, and Grain

Identify Business Processes and Events Before Designing Tables

Identify AtlasMart business processes and measurement events before designing tables, keeping system names and source gaps separate from analytical semantics.

Intermediate → Advanced90–110 minutesProcess/event classification labPython standard library · synthetic metadataLast reviewed: September 2026

Learning outcomes

Requirements are not yet dimensional models. AtlasMart must now identify the operational business processes that generate measurement events or capture measurable state. The process is the stable analytical target; source tables are implementation artifacts that can change, split, or merge.

01

Identify business processes from real-world activities and state-capture events instead of from source table names.

02

Distinguish transaction events, milestone events, state snapshots, reference changes, and derived analytical calculations.

03

Separate order capture, payment, return, inventory, fulfillment, and customer lifecycle even when one source system stores several of them.

04

Document source events and event keys required to support later grain declarations.

05

Reject process boundaries that combine unrelated measurement events merely because they share customer or product identifiers.

Executed baseline for this chapter

The local verification scripts in this generated chapter were executed with Python 3.13.5 and SQLite 3.46.1. They use only stable standard-library/SQL features. Learners should still record the versions printed on their machine; the exercises prove semantic contracts on the synthetic fixture, not performance of a production warehouse.

1. A business process is the measurement-producing activity

In dimensional modeling, a business process is an operational activity such as taking an order, receiving a payment, processing a return, shipping a package, or snapshotting inventory. It is not a department (“Sales”), a source database (“ERP”), or a dashboard (“Executive Overview”). A single ERP can contain several processes, while one process can be assembled from more than one source.

The process boundary matters because each process has its own event semantics, natural identifiers, time semantics, facts, and potentially different grain. “Sales” is too broad if it silently merges orders, payments, returns, and inventory into one table.

2. Event taxonomy for AtlasMart

Process Event/state Candidate event key Measurement time Chapter 01 evidence
Order capture Committed order line order_id + line_no order_ts orders + order_lines
Payment Authorization/capture/refund event payment_event_id payment event time Authority named, event rows not yet in tiny fixture
Return Returned line/item event return_event_id return event time Authority named, event rows not yet in tiny fixture
Inventory Product inventory state at snapshot snapshot_ts + product_id (+ location later) snapshot_ts inventory
Fulfillment Shipment/milestone event shipment/milestone ID milestone time Process named; source gap remains
Customer lifecycle Profile/segment change customer_id + effective change identity change/effective time CRM authority named; only current tiny fixture attributes

3. Tables are clues, not process definitions

The source table orders includes status and channel. It does not follow that every status transition is an order-capture fact, nor that payment settlement should be inferred from a string value if the payment process has its own authoritative events. Likewise, inventory contains one row per product in the tiny fixture, but the business process is a snapshot of inventory state at a particular time—not simply “the inventory table.”

Process discovery therefore asks: what happened or what state was observed, when, with which durable identity, which source owns that meaning, and what measurement was generated?

4. Hands-on lab — classify source evidence by process and event type

Save and run the following script. It records event semantics independently from eventual table names.

classify_process_events.py
from dataclasses import dataclass@dataclass(frozen=True)class Evidence:    name: str    process: str    event_kind: str    key: tuple[str, ...]    source_available: boolitems = [    Evidence("erp_order_line", "order_capture", "transaction", ("order_id","line_no"), True),    Evidence("inventory_snapshot", "inventory", "periodic_state", ("snapshot_ts","product_id"), True),    Evidence("payment_capture", "payment", "transaction", ("payment_event_id",), False),    Evidence("shipment_milestone", "fulfillment", "milestone", ("shipment_event_id",), False),    Evidence("customer_segment_change", "customer_lifecycle", "attribute_change", ("customer_id","effective_ts"), False),]for e in items:    assert e.process not in {"erp", "sales_department", "dashboard"}    assert e.key, e.name    print(e.name, "=>", e.process, e.event_kind, "available=", e.source_available)

Expected evidence lists five event/state surfaces. The order line and inventory snapshot are available in the tiny Chapter 01 fixture; payment, fulfillment, and historical customer-change rows are known requirements but not yet materialized in that fixture. Deliberately replace order_capture with erp; the assertion should fail because a system name is not a business process.

Cleanup: delete only classify_process_events.py.

5. Controlled failure: one universal “sales” fact

A team may propose one table with columns for order amount, payment amount, return amount, inventory on hand, and shipment time because every row has a product and customer. That table has no single measurement event: an order line occurs at purchase, a payment may happen later, a return may happen days later, inventory is a state snapshot, and fulfillment has its own milestones.

Even if the table can be populated with nulls, its semantics are unstable. Aggregating rows can mix incomparable event populations and dates. The repair is to keep business processes separate at their natural grain and integrate them later through conformed dimensions and drill-across patterns when appropriate.

6. Process boundary checklist

Question Acceptable evidence
What happened or what state was captured? A business-language event/state, not a source table name
What uniquely identifies the event at the source? Natural event key or explicit source position
Which timestamp means business occurrence? Event time / snapshot time, not merely ingestion time
Which source owns the meaning? Named source authority; gaps remain explicit
Can every candidate measure be true for that event? If not, it may belong to another process/grain
Can reruns identify the same event? Stable identity or deterministic deduplication policy

7. Production judgment and bridge

A good process list is small enough to name concrete operational activities but rich enough to cover the organization’s measurement events. Process boundaries should survive application refactoring: renaming an ERP table should not change what “order capture” means.

Now that AtlasMart has process/event boundaries, the next step is to declare one precise grain for each fact candidate. That declaration is a binding contract: every future dimension and fact must be true at that grain.

Knowledge check

Check your understanding

  1. Why is “ERP” not a business process?
  2. Why should payment and order capture remain separate even if the order row contains a paid status?
  3. What kind of process is the inventory example?
  4. What does an event key contribute to requirements engineering?
  5. Why is a table full of nullable columns not a valid solution to incompatible processes?
Review the answers

1. ERP names a source/system boundary. A business process names the real activity or state capture that generates measurements.

2. They are different operational events with different timing, identifiers, retry/correction behavior, and potentially different authoritative evidence.

3. A periodic state/snapshot process: inventory is measured as of a particular snapshot time for a product (and later possibly a location).

4. It identifies the repeatable unit of business evidence and helps later grain, deduplication, rerun, and source-to-target decisions.

5. Nullability does not create one coherent measurement event. The rows would still have mixed semantics and unsafe aggregation.

Summary and next step

AtlasMart’s core process candidates are now distinct from systems and departments. The tiny fixture directly supports order capture and an inventory snapshot; other named processes remain explicit source gaps.

Next: declare grain before selecting dimensions or facts.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.