Chapter 02 · Requirements Engineering: Business Processes, Questions, Metrics, Dimensions, and Grain
Identify Business Processes and Events Before Designing Tables
Identify AtlasMart business processes and measurement events before designing tables, keeping system names and source gaps separate from analytical semantics.
Learning outcomes
Requirements are not yet dimensional models. AtlasMart must now identify the operational business processes that generate measurement events or capture measurable state. The process is the stable analytical target; source tables are implementation artifacts that can change, split, or merge.
Identify business processes from real-world activities and state-capture events instead of from source table names.
Distinguish transaction events, milestone events, state snapshots, reference changes, and derived analytical calculations.
Separate order capture, payment, return, inventory, fulfillment, and customer lifecycle even when one source system stores several of them.
Document source events and event keys required to support later grain declarations.
Reject process boundaries that combine unrelated measurement events merely because they share customer or product identifiers.
The local verification scripts in this generated chapter were executed with Python 3.13.5 and SQLite 3.46.1. They use only stable standard-library/SQL features. Learners should still record the versions printed on their machine; the exercises prove semantic contracts on the synthetic fixture, not performance of a production warehouse.
1. A business process is the measurement-producing activity
In dimensional modeling, a business process is an operational activity such as taking an order, receiving a payment, processing a return, shipping a package, or snapshotting inventory. It is not a department (“Sales”), a source database (“ERP”), or a dashboard (“Executive Overview”). A single ERP can contain several processes, while one process can be assembled from more than one source.
The process boundary matters because each process has its own event semantics, natural identifiers, time semantics, facts, and potentially different grain. “Sales” is too broad if it silently merges orders, payments, returns, and inventory into one table.
2. Event taxonomy for AtlasMart
| Process | Event/state | Candidate event key | Measurement time | Chapter 01 evidence |
|---|---|---|---|---|
| Order capture | Committed order line | order_id + line_no | order_ts | orders + order_lines |
| Payment | Authorization/capture/refund event | payment_event_id | payment event time | Authority named, event rows not yet in tiny fixture |
| Return | Returned line/item event | return_event_id | return event time | Authority named, event rows not yet in tiny fixture |
| Inventory | Product inventory state at snapshot | snapshot_ts + product_id (+ location later) | snapshot_ts | inventory |
| Fulfillment | Shipment/milestone event | shipment/milestone ID | milestone time | Process named; source gap remains |
| Customer lifecycle | Profile/segment change | customer_id + effective change identity | change/effective time | CRM authority named; only current tiny fixture attributes |
3. Tables are clues, not process definitions
The source table orders includes
status and channel. It does not follow
that every status transition is an order-capture fact, nor that
payment settlement should be inferred from a string value if the
payment process has its own authoritative events. Likewise,
inventory contains one row per product in the tiny
fixture, but the business process is a
snapshot of inventory state at a particular time—not
simply “the inventory table.”
Process discovery therefore asks: what happened or what state was observed, when, with which durable identity, which source owns that meaning, and what measurement was generated?
4. Hands-on lab — classify source evidence by process and event type
Save and run the following script. It records event semantics independently from eventual table names.
from dataclasses import dataclass@dataclass(frozen=True)class Evidence: name: str process: str event_kind: str key: tuple[str, ...] source_available: boolitems = [ Evidence("erp_order_line", "order_capture", "transaction", ("order_id","line_no"), True), Evidence("inventory_snapshot", "inventory", "periodic_state", ("snapshot_ts","product_id"), True), Evidence("payment_capture", "payment", "transaction", ("payment_event_id",), False), Evidence("shipment_milestone", "fulfillment", "milestone", ("shipment_event_id",), False), Evidence("customer_segment_change", "customer_lifecycle", "attribute_change", ("customer_id","effective_ts"), False),]for e in items: assert e.process not in {"erp", "sales_department", "dashboard"} assert e.key, e.name print(e.name, "=>", e.process, e.event_kind, "available=", e.source_available)
Expected evidence lists five event/state surfaces. The order
line and inventory snapshot are available in the tiny Chapter 01
fixture; payment, fulfillment, and historical customer-change
rows are known requirements but not yet materialized in that
fixture. Deliberately replace order_capture with
erp; the assertion should fail because a system
name is not a business process.
Cleanup: delete only
classify_process_events.py.
5. Controlled failure: one universal “sales” fact
A team may propose one table with columns for order amount, payment amount, return amount, inventory on hand, and shipment time because every row has a product and customer. That table has no single measurement event: an order line occurs at purchase, a payment may happen later, a return may happen days later, inventory is a state snapshot, and fulfillment has its own milestones.
Even if the table can be populated with nulls, its semantics are unstable. Aggregating rows can mix incomparable event populations and dates. The repair is to keep business processes separate at their natural grain and integrate them later through conformed dimensions and drill-across patterns when appropriate.
6. Process boundary checklist
| Question | Acceptable evidence |
|---|---|
| What happened or what state was captured? | A business-language event/state, not a source table name |
| What uniquely identifies the event at the source? | Natural event key or explicit source position |
| Which timestamp means business occurrence? | Event time / snapshot time, not merely ingestion time |
| Which source owns the meaning? | Named source authority; gaps remain explicit |
| Can every candidate measure be true for that event? | If not, it may belong to another process/grain |
| Can reruns identify the same event? | Stable identity or deterministic deduplication policy |
7. Production judgment and bridge
A good process list is small enough to name concrete operational activities but rich enough to cover the organization’s measurement events. Process boundaries should survive application refactoring: renaming an ERP table should not change what “order capture” means.
Now that AtlasMart has process/event boundaries, the next step is to declare one precise grain for each fact candidate. That declaration is a binding contract: every future dimension and fact must be true at that grain.
Knowledge check
Check your understanding
- Why is “ERP” not a business process?
- Why should payment and order capture remain separate even if the order row contains a paid status?
- What kind of process is the inventory example?
- What does an event key contribute to requirements engineering?
- Why is a table full of nullable columns not a valid solution to incompatible processes?
Review the answers
1. ERP names a source/system boundary. A business process names the real activity or state capture that generates measurements.
2. They are different operational events with different timing, identifiers, retry/correction behavior, and potentially different authoritative evidence.
3. A periodic state/snapshot process: inventory is measured as of a particular snapshot time for a product (and later possibly a location).
4. It identifies the repeatable unit of business evidence and helps later grain, deduplication, rerun, and source-to-target decisions.
5. Nullability does not create one coherent measurement event. The rows would still have mixed semantics and unsafe aggregation.
Summary and next step
AtlasMart’s core process candidates are now distinct from systems and departments. The tiny fixture directly supports order capture and an inventory snapshot; other named processes remain explicit source gaps.
Next: declare grain before selecting dimensions or facts.
Authoritative references
- Kimball Group — Four-Step Dimensional Design Process — Select business process, declare grain, identify dimensions, then identify facts.
- Kimball Group — Business Processes — Business processes are measurement-generating operational activities and define a design target.
- Kimball Group — Grain — Grain is the binding statement of what one fact row represents and must precede dimensions/facts.
- Kimball Group — 10 Essential Rules of Dimensional Modeling — Reinforces process-centric dimensional models and measurement events.
- Kimball Group — Enterprise Data Warehouse Bus Matrix — Rows represent business processes and later coordinate shared dimensions.
- Python documentation — dataclasses — Standard-library construct used by the process-classification lab.