Chapter 08 · Importing Data: LOAD CSV, Data Importer, Bulk Import, Transformation, and Validation
Bulk/Offline Import Concepts, Ordering, Memory/Storage Planning, and When LOAD CSV Is the Wrong Tool
Plan bulk/offline import with explicit headers, dry-run evidence, resource headroom and destructive-boundary controls.
Learning outcomes
AtlasMart now has a migration large enough that online
transactional insertion may be the wrong tool.
neo4j-admin database import is a separate
store-building path with offline/direct-server assumptions.
Choose online LOAD CSV versus admin import from scale and availability requirements.
Explain full-import empty/non-existent destination expectations.
Prepare :ID/:START_ID/:END_ID/:TYPE headers and typed properties.
Use current dry-run and report evidence before a real import.
Plan admin heap/off-heap and disk headroom and protect destructive overwrite boundaries.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
authentication enabled, no mandatory APOC/GDS plugin, and
stable AtlasMart business identifiers from Chapters 01–07.
Neo4j 5.26.30 remains the LTS comparison line.
Local file:/// sources are a self-managed
server-filesystem feature and are not available on Aura.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. Expected outputs are fixture
invariants, not fabricated captures. Every write is scoped to
labTag='ch08', dedicated CSV files, or a
disposable database. Never use broad cleanup, store overwrite,
or offline import commands against unrelated or production
data.
1. Different mechanism, different contract
| Dimension | LOAD CSV | neo4j-admin import |
|---|---|---|
| Execution | Online transactional Cypher | Admin/store-building workflow |
| Destination | Existing database is normal | Full import targets non-existent/empty database |
| Identity | MATCH/MERGE by graph properties | Importer ID spaces/header semantics |
| Failure evidence | Committed batches + Cypher status | Dry-run/import reports/logs |
| Typical fit | Small/medium online transformation | Very large initial/bulk migration |
2. Header semantics define endpoints and types
# customers-header.csvcustomerId:ID(Customer),name,email,tier,:LABEL# orders-header.csvorderId:ID(Order),orderedAt:datetime,status,:LABEL# placed-header.csv:START_ID(Customer),:END_ID(Order),:TYPE# contains-header.csv:START_ID(Order),:END_ID(Product),quantity:int,unitPrice:double,:TYPE
ID spaces keep identifier domains explicit. Typed headers avoid storing numeric/temporal data as accidental strings.
3. Dry run before store creation
Current Neo4j supports --dry-run=true to validate
inputs and estimate import characteristics. The importer also
exposes report/log evidence. Record version, file sizes,
accepted/rejected rows, effective memory budget and resulting
estimates rather than copying a throughput number from another
machine.
docker exec atlasmart-neo4j neo4j-admin database import full \ --dry-run=true \ --nodes=Customer=/var/lib/neo4j/import/ch08/admin/customers-header.csv,/var/lib/neo4j/import/ch08/admin/customers.csv \ --nodes=Order=/var/lib/neo4j/import/ch08/admin/orders-header.csv,/var/lib/neo4j/import/ch08/admin/orders.csv \ --relationships=PLACED=/var/lib/neo4j/import/ch08/admin/placed-header.csv,/var/lib/neo4j/import/ch08/admin/placed.csv \ --path-pattern-style=none \ atlasmart_bulk_ch08
4. Memory/storage planning
The current importer has an admin JVM heap plus an off-heap
working-memory budget controlled by
--max-off-heap-memory when set. Import memory is
not simply the live server page-cache setting. Plan source
bytes, destination store, temporary/log space, disk bandwidth
and container volumes with failure headroom.
5. Deliberately unsafe: overwrite the live destination
--overwrite-destination deletes existing
database files.
The mandatory lab never applies it to neo4j. Use
a dedicated disposable destination and validate the dry run
first.
Also, a zero exit status does not prove the business graph is correct; wrong input semantics can still create a valid but incorrect store.
Hands-on bulk decision lab
Build importer headers for the corrected AtlasMart dataset, run
only a dry run against atlasmart_bulk_ch08, capture
warnings/reports and compare the operational assumptions with
the online LOAD CSV workflow.
Check your understanding
- Why is admin import not just faster LOAD CSV?
- What do :START_ID and :END_ID mean?
- Why run --dry-run=true?
- Which memory resources matter?
- Why is --overwrite-destination high risk?
Review the answers
1. It builds stores using an admin/offline mechanism with different destination and input semantics.
2. Importer identity references for relationship endpoints.
3. To validate data/specification and estimate import requirements before writing a store.
4. Admin JVM heap plus importer off-heap working memory, as well as disk/storage headroom.
5. It deletes existing destination database files and therefore requires explicit isolation/recovery planning.
Summary and next step
Bulk import solves throughput, not trust. The final lesson applies the same reconciliation gate regardless of ingestion tool.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for the chapter baseline.
- Cypher LOAD CSV — Current LOAD CSV parsing, typing and local/remote source semantics.
- CALL subqueries in transactions — Bounded inner transactions, default batch size and partial-commit behavior.
- Neo4j Admin import — Current full/incremental import, dry run, headers, memory and report behavior.
- Default file locations — Self-managed import-directory and store/log boundaries.
- Import operations manual — Current full/incremental import, dry run, memory, headers and report behavior.
- Docker operations — Running neo4j-admin/import inside Docker.
- 2025–2026 import changes — Recent import logging/options/default changes.