Chapter 12 · APOC, Procedures, Functions, Triggers-Like Workflows, and Extending Cypher

Data Integration and Export Procedures: Files, HTTP, JSON, and Operational Safety

Treat APOC file/HTTP/JSON integration as a filesystem/network security boundary, validate inputs before writes, and prefer auditable outbox workflows for external side effects.

Advanced160–205 minutesJSON integration + safety labNeo4j 2026.07.1 Community · Cypher 25APOC Core 2026.07.1 optionalLast reviewed: September 2026

Learning outcomes

AtlasMart needs to ingest a small supplier JSON feed and export a review snapshot. That requirement crosses the database process into the filesystem or network. A procedure that “loads JSON” is therefore also an SSRF, credential, timeout, file-permission, egress and provenance problem.

01

Use apoc.load.json only with explicit URL/file trust and timeout assumptions.

02

Explain why APOC file import/export is disabled by default and how directory restrictions work.

03

Prefer streaming exports when a managed/cloud environment does not expose server filesystem access.

04

Distinguish data integration from transactionally safe business side effects.

05

Design source validation, quarantine and reconciliation around APOC rather than trusting procedure completion.

Chapter 12 baseline · reviewed 9 September 2026

The mandatory lab continues Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 for version-sensitive examples, container atlasmart-neo4j, authentication enabled, loopback HTTP/Bolt endpoints, and the AtlasMart identifiers/model built in Chapters 01–11. Neo4j 5.26.30 remains the LTS comparison line. This chapter adds APOC Core 2026.07.1 as an optional extension; APOC Extended is never required.

Evidence and privilege note

This generation environment does not run the Neo4j container, so no plugin load, procedure list, filesystem read/write, HTTP request, trigger execution or custom-JAR output is invented. The lab tells you exactly what to inspect. Community Edition lacks Enterprise RBAC/load privileges, so the free path emphasizes configuration minimization, loopback-only access, file/network denial by default, and application-level authorization boundaries.

1. Loading bytes is not validating a source

Boundary Threat / failure Control
URL selection SSRF to metadata/private services Allow only approved hosts; Community uses network/IP blocking controls; Enterprise can add LOAD privileges/CIDR restrictions
Authentication Secrets leaked in query logs/history or URLs Keep credentials out of URLs/query text; use application/proxy/secret mechanisms appropriate to deployment
Timeouts Slow remote endpoint ties up server work Configure connect/read timeouts and application retry budgets
Payload size/compression Large or decompression-bomb payloads consume heap/disk Set source limits and APOC decompression ratio controls
Schema/identity Valid JSON can contain duplicate/orphan business records Validate stable keys and reconcile counts/invariants after ingest
Filesystem Broad path access exposes server files Keep file import/export disabled; if enabled, restrict to import directory

2. Secure defaults: no file read, no file write

apoc.conf · only if the lab genuinely needs local files
# Keep these false unless the exercise explicitly needs local files.apoc.import.file.enabled=falseapoc.export.file.enabled=falseapoc.import.file_use_neo4j_config=true# Optional HTTP guardrails; tune to your environment, not folklore.apoc.http.timeout.connect=10000apoc.http.timeout.read=60000

When file access is enabled and apoc.import.file_use_neo4j_config=true, APOC respects Neo4j file URL permission and server.directories.import. Turning that integration off can allow arbitrary filesystem paths and should be treated as a major security change.

3. Controlled JSON load: local deterministic source first

For a reproducible lab, prefer a tiny JSON file you own rather than a live third-party endpoint. Place atlasmart-suppliers.json in the configured import directory only after explicitly enabling APOC file import for the disposable instance.

JSON · deterministic input
[  {"supplierId":"SUP-1001","name":"Northwind Optics","country":"TR"},  {"supplierId":"SUP-1002","name":"Caspian Sensors","country":"AZ"}]
Cypher · load and validate before MERGE
CYPHER 25CALL apoc.load.json('file:///atlasmart-suppliers.json') YIELD valueUNWIND value AS rowWITH rowWHERE row.supplierId IS NOT NULL  AND row.name IS NOT NULLRETURN row.supplierId AS supplierId,row.name AS name,row.country AS countryORDER BY supplierId;// Only after the validation query matches the expected fixture:CALL apoc.load.json('file:///atlasmart-suppliers.json') YIELD valueUNWIND value AS rowWITH row WHERE row.supplierId IS NOT NULL AND row.name IS NOT NULLMERGE (s:Supplier {supplierId:row.supplierId})SET s.name=row.name,s.country=row.country;

The first query is deliberately read-only. Procedure completion proves that parsing succeeded; it does not prove source completeness, uniqueness or business validity.

4. HTTP is a server-side outbound request

Cypher · shape of a controlled HTTPS load
CYPHER 25:param supplierUrl => 'https://approved.example.invalid/atlasmart/suppliers.json';CALL apoc.load.json($supplierUrl, '$[*]', {failOnError:true}) YIELD valueRETURN value.supplierId AS supplierId, value.name AS nameORDER BY supplierId;
Do not run this placeholder URL as evidence.

Replace it only with an endpoint you own/control and whose egress is explicitly allowed. Community and Enterprise have different policy controls; neither tier makes arbitrary URL loading safe by default.

5. Export: prefer stream when server files are undesirable

Cypher · stream JSON to the client
CYPHER 25MATCH (p:Product)WITH collect(p) AS productsCALL apoc.export.json.data(products, [], null, {stream:true})YIELD data, nodes, relationships, propertiesRETURN data, nodes, relationships, properties;

Streaming avoids granting broad server write access and is often the only sensible choice in managed/cloud environments. If a file export is truly required on self-managed Neo4j, enable apoc.export.file.enabled=true for the controlled window and keep the output directory restricted.

6. Deliberately wrong: load directly from an arbitrary user URL

An API endpoint that accepts ?url=... and passes it to apoc.load.json creates an SSRF primitive. Parameterization prevents Cypher injection but does not make the destination trustworthy. The repair is destination allowlisting/proxying, network-level blocking, bounded timeouts/size, application authorization and source lineage checks.

Evidence after integration Why it matters
Source checksum/version/request ID Reproducibility and lineage
Rows parsed / rejected / merged Detect silent partial processing
Duplicate/orphan reports Protect graph identity/referential invariants
HTTP/file error classification Separate source failure from database failure
Post-import business reconciliation Procedure success is not business correctness

7. Trigger-like workflows: do not hide external side effects inside database callbacks

APOC Core includes trigger procedures when apoc.trigger.enabled=true; current apoc.trigger.install is admin-only and the older apoc.trigger.add is removed in Cypher 25. Triggers execute server-side and increase coupling/observability burden. For “send message after order commit,” prefer an OutboxEvent node written in the same transaction and an application worker that performs the external HTTP/email side effect idempotently.

Cypher · transactional outbox pattern
CYPHER 25MATCH (o:Order {orderId:$orderId})SET o.status='READY_TO_NOTIFY'MERGE (e:OutboxEvent {eventId:$eventId})ON CREATE SET e.kind='ORDER_READY',e.orderId=o.orderId,e.createdAt=datetime(),e.status='PENDING'MERGE (o)-[:EMITTED]->(e);

Check your understanding

  1. Why does parameterizing a URL not solve SSRF?
  2. Why are APOC local file reads/writes disabled by default?
  3. What does a successful apoc.load.json invocation prove?
  4. Why can streaming export be safer than file export?
  5. Why prefer an outbox for external notifications?
Review the answers

1. Parameters prevent query-text injection but do not restrict where the server is allowed to connect.

2. Filesystem access increases sensitive-data and path-exposure risk and should be enabled only when needed.

3. That the procedure parsed/fetched enough to return; not that business keys, completeness or invariants are correct.

4. It avoids granting the Neo4j server write access to persistent filesystem paths and returns data through the client boundary.

5. The graph commit records a durable idempotent event; a separate worker can retry external side effects without hiding them inside database trigger execution.

Summary and next step

File and HTTP procedures cross trust boundaries. Lesson 4 examines the even stronger coupling of custom Java procedures/functions and the deployment risks of putting proprietary logic inside the Neo4j JVM.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.