Chapter 02 · Documents, Collections, Fields, References, Maps, Arrays, Timestamps, GeoPoints, and Limits

Auto IDs vs Semantic IDs, Hotspot Risks, Tenant Prefixes, and External Identity Mapping

Choose document IDs from access, security, and scale requirements rather than aesthetics: compare scattered auto IDs, semantic IDs, tenant prefixes, and external identity mappings without creating hotspot or privacy debt.

Beginner → Advanced90–120 minutesAtlasMart emulator-first labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

Document IDs are visible in paths, participate in lexicographic key distribution, often leak into URLs/logs, and frequently become application identifiers. Choosing them only because they look convenient can create hotspots, expose business sequence information, or couple Firestore to an external system that later changes.

01

Explain why Firestore auto IDs are scattered and why they do not encode creation time or sort order.

02

Identify hotspot risk from high-rate writes to lexicographically close or monotonically increasing document IDs.

03

Use semantic IDs only when deterministic lookup/idempotency benefits outweigh coupling and scale risks.

04

Evaluate tenant-prefixed IDs against alternative path/field designs instead of assuming prefixes are good partitioning.

05

Map external identities through opaque, privacy-aware keys without storing secrets or mutable authorization state in document IDs.

Chapter 02 baseline reviewed 14 September 2026

The lab continues Chapter 01 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native semantics are the default. Enterprise-only differences are labeled explicitly.

Execution and evidence note

These lessons are authored against current official documentation and the deterministic emulator design from Chapter 01. This generation environment does not run the Firebase Emulator Suite or install the pinned npm dependencies, so command output is described by invariant and expected shape rather than presented as captured execution. The emulator is evidence for local application behavior, not proof of production quotas, regional latency, backend index topology, or billing.

1. Auto IDs are scattered identifiers, not chronological keys

Firestore can generate document IDs for you. Current best-practice guidance explains that automatic IDs use a scatter strategy that helps avoid hotspotting when documents are created at high rates. They are not Realtime Database push IDs and they do not encode creation time. If AtlasMart needs chronological ordering, it stores an explicit timestamp field and queries/orders by that field with the appropriate index.

This separation is healthy: identity answers “which document?” while a timestamp answers “when?”. Encoding both into one monotonically increasing ID can make application code look simple but concentrate writes into a narrow key range.

generate IDs without writing documents · Web SDK
import { collection, doc } from "firebase/firestore";const ids = Array.from({ length: 12 }, () => doc(collection(db, "orders")).id);console.log(ids);console.assert(new Set(ids).size === ids.length);// Do not sort these IDs and infer creation order.// Store createdAt explicitly when temporal ordering is required.
What this proves

Seeing non-monotonic IDs demonstrates the client-generated identifier shape. It does not measure backend key-range distribution or guarantee throughput for your production workload; that requires production-scale evidence and Chapter 15 load testing.

2. Sequential semantic IDs can turn identity into a hotspot

IDs such as order-0000001, order-0000002, and order-0000003 are operationally attractive because people can read them, but sustained high-rate creation places writes into lexicographically close key ranges. Firestore's scaling best practices explicitly warn against monotonically increasing document IDs and against high traffic to narrow document ranges.

The correct alternative depends on the requirement. If the business needs a human-facing sequential order number, store it as a field and use a separate scattered document ID. If the external system already provides a stable opaque identifier with good distribution and AtlasMart needs idempotent upsert by that ID, a semantic document ID can be reasonable—but the distribution and privacy properties must be reviewed.

Requirement Bad shortcut Safer pattern
Human-readable order number Use sequential order number as Firestore document ID Use auto/scattered document ID; store orderNumber as a field.
Idempotent import from external system Blindly use raw external ID in path Validate format/privacy/distribution; map or hash when appropriate.
Per-tenant grouping Prefix every high-rate ID with tenantA-... Model tenant ownership in path/field/rules; evaluate key-range concentration.
Chronological order Sort by auto ID Store createdAt and query by timestamp.

3. Tenant prefixes are naming conventions, not automatic sharding

A path such as orders/tenantA-20260914-000001 embeds tenant, time, and sequence into one lexicographically clustered key. That can be useful for human debugging but it is not a free partitioning strategy. At high write rates it may put a tenant's traffic into a narrow range—the opposite of what you wanted.

Alternatives include a root orders collection with opaque IDs plus tenantId as an indexed authorization/query field, or a hierarchy such as tenants/{tenantId}/orders/{opaqueId} when ownership/lifecycle/rules benefit from the path. The hierarchy changes collection-group query, rules, deletion, and cross-tenant operational workflows, so choose it for those semantics—not because the prefix “looks partitioned.”

Security warning

Document IDs and paths are not secret. They appear in logs, URLs, error messages, references, exports, and telemetry. Do not place access tokens, emails, payment data, or other secrets/PII in IDs merely because they make lookup convenient.

4. External identity mapping decouples Firestore from another system's key

AtlasMart integrates sellers, payment providers, and legacy order systems. External IDs can be mutable, globally unique only within a provider, numeric beyond JavaScript's exact range, or privacy-sensitive. A dedicated mapping document lets the database use an opaque internal identity while preserving deterministic external lookup.

external identity mapping · trusted server example
import { createHash } from "node:crypto";function mappingId(provider, externalId) {  return createHash("sha256")    .update(`${provider}\0${externalId}`, "utf8")    .digest("hex");}const provider = "legacy-erp";const externalId = "9223372036854775806"; // keep exact opaque ID as stringconst id = mappingId(provider, externalId);await db.doc(`externalIdMap/${id}`).set({  provider,  externalId,  targetRef: db.doc("products/p-1001"),  schemaVersion: 1});

Hashing does not make a secret safe if the input space is guessable; it mainly creates a bounded opaque path key and avoids illegal characters/oversized IDs. Protect the mapping document with server-side authorization and store only the external fields required for reconciliation.

5. Hands-on lab: compare auto IDs, sequential IDs, and mapped identities

The emulator cannot reproduce production hotspot physics faithfully, so this lab separates identity evidence from throughput claims. Generate ID samples locally, write a small deterministic fixture, and record which properties are proven. Chapter 15 will perform workload-oriented hotspot testing.

scripts/ch02-ids.mjs · identity behavior only
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";import { initializeApp } from "firebase/app";import { getFirestore, connectFirestoreEmulator, collection, doc, setDoc, serverTimestamp } from "firebase/firestore";const app = initializeApp({ apiKey: "demo-key", projectId: "demo-atlasmart-firestore", appId: "1:123:web:demo" });const db = getFirestore(app);connectFirestoreEmulator(db, "127.0.0.1", 8080);const autoRefs = Array.from({ length: 5 }, () => doc(collection(db, "idDemos")));console.log("auto IDs", autoRefs.map(r => r.id));const sequential = [1,2,3,4,5].map(n => `order-${String(n).padStart(8, "0")}`);console.log("sequential IDs", sequential);for (const ref of autoRefs) {  await setDoc(ref, { kind: "auto", createdAt: serverTimestamp() });}// Small lab writes are safe evidence of correctness only; do not infer hotspot throughput.

Verification checklist: IDs are unique; auto IDs are not monotonic; chronological ordering is represented by createdAt; no secret or PII is encoded in the path; any semantic ID has a documented source, immutability rule, distribution assumption, and collision behavior.

6. Deliberately wrong approach: make the path the business database

Teams sometimes encode tenant, date, region, status, and sequence into a single document ID so they can “query from the name.” Firestore Core queries do not turn arbitrary document-ID substrings into a relational index, and changing a business attribute would require moving/copying documents because the path is the identity. This also bakes authorization and privacy-sensitive information into every reference/log.

The repair is to keep path identity stable and model queryable business dimensions as fields with explicit indexes. Use path hierarchy for ownership/lifecycle when it materially simplifies the design, not as a replacement for fields and indexes.

Production judgment

The best ID is the one whose uniqueness, immutability, distribution, privacy, idempotency, migration, and lookup properties match the workload. “Auto IDs everywhere” and “semantic IDs everywhere” are both oversimplifications.

Knowledge check

Check your understanding

  1. Do Firestore auto IDs provide chronological ordering?
  2. Why can monotonically increasing IDs cause problems at high write rates?
  3. When can a semantic ID still be reasonable?
  4. What is the risk of tenantA-000001-style IDs?
  5. Why might AtlasMart hash an external provider ID for a mapping document?
Review the answers

1. No. Store a timestamp field when ordering by creation time is required.

2. They concentrate writes into lexicographically close key ranges, increasing hotspot risk and latency/contention.

3. When deterministic lookup/idempotency is valuable and the ID is stable, privacy-safe, legal, and sufficiently distributed for the expected workload.

4. They can cluster one tenant’s traffic into a narrow range and leak tenant/business sequencing into paths.

5. To create a bounded opaque legal document ID while retaining provider/externalId as explicit fields. Hashing is not authorization or encryption.

Summary and next step

AtlasMart now separates identity from ordering and business attributes. Auto IDs provide scattered opaque keys; semantic IDs require explicit distribution/privacy reasoning; tenant prefixes are not sharding; external identities can be mapped without coupling every document path to another system. The final lesson turns all Chapter 02 decisions into a cross-SDK schema test so Web, mobile, and trusted server code observe the same meaning.

Next: Inspect and Validate a Document Schema Across Web/Mobile/Server SDK Serialization Boundaries.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.