Chapter 26 · Sharding, Distributed Database Concepts, and Global Data Placement
Compare Sharding with RAC, Data Guard, Partitioning, and Application-Level Distribution
Select among sharding, RAC, Data Guard, partitioning and application-owned distribution by dataset placement, transaction scope, failure isolation, latency, consistency, licensing, operational skill and measurable workload/SLO evidence.
Learning outcomes
An architecture review says “we need sharding for availability,” but the actual requirement is one database surviving a server failure. Another team says “use RAC for global data sovereignty,” but RAC instances still access the same database. Oracle has several scale/availability techniques whose names are often conflated. The final design step is to map each requirement to the mechanism that actually changes data placement or failure scope.
Compare sharding, RAC, Data Guard, table partitioning and application-level distribution by physical data placement and transaction scope.
Separate scale-out throughput, node HA, site DR, manageability, sovereignty and partition pruning as different requirements.
Apply the current 26ai licensing matrix for Free, RAC, Data Guard, Active Data Guard and sharding/Raft limits.
Use a measurable ServiceHub decision table rather than architecture fashion.
Identify combinations such as sharded databases whose individual shards also use RAC/Data Guard and understand the multiplied operational cost.
Mandatory examples target Oracle AI Database Free 26ai and were reviewed against the current August 2026 licensing documentation, RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1.222.1617. Oracle documentation now uses Oracle Globally Distributed AI Database / Oracle Globally Distributed Database for the technology formerly called Oracle Sharding. The 26ai licensing matrix marks Oracle Globally Distributed Database available in Free with use limited to three shards. Raft replication is also marked available in Free with use limited to three nodes, and each Free node remains limited to 2 CPU cores, 2 GB RAM, and 12 GB user data. However, the traditional shard-catalog creation guide still documents the shard catalog as an Enterprise Edition PDB and explicitly disallows CDB$ROOT as the catalog. Therefore this chapter does not claim that a complete catalog + shard directors + multi-host shard topology is a zero-cost single-machine lab. Mandatory work uses one Free FREEPDB1 database to simulate placement, routing, fan-out and failure semantics while all GDSCTL/SHARD DDL examples are clearly labeled topology-dependent. Shard directors may run on separate servers without a separate shard-director server license. Full Data Guard, Active Data Guard, RAC, GoldenGate, Exadata and multi-host production topologies retain their own entitlement/operational requirements. No lab raises COMPATIBLE, changes hidden parameters, or modifies the GitHub repository.
1. The fastest way to choose wrong is to compare product names instead of failure domains
| Mechanism | Where rows live | Main boundary changed |
|---|---|---|
| Partitioning | Partitions inside one database | Manageability/pruning/storage layout |
| RAC | One database/storage, multiple instances/nodes | Instance/node availability and scale-up/scale-out access |
| Data Guard | Primary database plus standby database copies | Site/database disaster recovery and protection |
| Sharding / Globally Distributed DB | Different row subsets in independent databases | Dataset/failure/geographic placement and horizontal scale |
| Application-owned distribution | Whatever databases/services the application maps manually | Custom placement/failure/routing under application responsibility |
2. Partitioning is not sharding
Partitioning can prune scans, simplify lifecycle operations and spread segments/tablespaces, but the partitions still belong to one database. A database outage affects every partition. Oracle Sharding itself uses partitioning concepts internally, but places partitions/chunks across independent shard databases.
CREATE TABLE sh26_partition_demo ( tenant_id NUMBER NOT NULL, work_order_id NUMBER NOT NULL, status_code VARCHAR2(12) NOT NULL)PARTITION BY HASH (tenant_id)PARTITIONS 4;-- Four partitions; still one FREEPDB1 database/failure domain.
3. RAC is not sharding
Oracle Real Application Clusters (RAC) runs multiple database instances against the same database storage. Cache Fusion/GCS/GES coordinate block ownership across instances. RAC can protect against instance/node failure and scale some workloads, but it does not place Tenant A's database rows in one sovereign region and Tenant B's rows in a physically independent database.
4. Data Guard is replication, not horizontal partitioning
Data Guard maintains standby copies of a database through redo transport/apply. It is a high availability/disaster recovery/data-protection mechanism. A standby generally contains the same database data rather than a disjoint tenant subset. Active Data Guard can open standby workloads for read and adds separately licensed capabilities on EE/EE-ES.
5. Sharding can combine with RAC/Data Guard
Each shard is still an Oracle database. A high-end deployment can therefore protect each shard with Data Guard and/or run a shard on RAC. That combines row-set distribution with local node/site HA—but multiplies databases, services, backups, patching, replication and licensing. Architecture diagrams should show both levels explicitly.
6. Raft is a sharding-native replication alternative in 26ai
26ai introduces native Raft replication for globally distributed deployments, using consensus-based replicated units and automatic redistribution/failover management. The licensing matrix marks Raft replication available in Free up to three nodes; each Free node keeps the Free resource limits. In EE/EE-ES the current matrix includes additional RAC requirements/rights for participating nodes, so exact production entitlements must be checked.
7. Application-level distribution buys flexibility by moving responsibility outward
An application can maintain its own tenant-to-database directory and route JDBC connections manually. That can work, especially for a small number of isolated tenants, but then schema rollout, global reporting, tenant moves, retry semantics, failover, backups, credentials and directory consistency are custom application/platform responsibilities. Oracle sharding exists partly to provide those database-level control-plane/routing functions.
8. Free decision-matrix lab
BEGIN EXECUTE IMMEDIATE 'DROP TABLE sh26_arch_eval PURGE';EXCEPTION WHEN OTHERS THEN IF SQLCODE != -942 THEN RAISE; END IF; END;/CREATE TABLE sh26_arch_eval ( criterion VARCHAR2(50) PRIMARY KEY, weight NUMBER NOT NULL CHECK (weight BETWEEN 1 AND 5), partitioning_score NUMBER CHECK (partitioning_score BETWEEN 0 AND 5), rac_score NUMBER CHECK (rac_score BETWEEN 0 AND 5), dataguard_score NUMBER CHECK (dataguard_score BETWEEN 0 AND 5), sharding_score NUMBER CHECK (sharding_score BETWEEN 0 AND 5), app_distribution_score NUMBER CHECK (app_distribution_score BETWEEN 0 AND 5), evidence VARCHAR2(500));INSERT INTO sh26_arch_eval( criterion,weight, partitioning_score,rac_score,dataguard_score, sharding_score,app_distribution_score,evidence) VALUES( 'DATA_SOVEREIGNTY',5, 1,0,2,5,5, 'Need tenant subsets physically placed by legal region');INSERT INTO sh26_arch_eval VALUES( 'NODE_FAILURE_WITHOUT_DB_FAILOVER',4, 1,5,3,3,2, 'Requirement is service continuity after one compute node fails');INSERT INTO sh26_arch_eval VALUES( 'SITE_DR',5, 1,2,5,4,2, 'Need tested remote recovery with explicit RPO/RTO');INSERT INTO sh26_arch_eval VALUES( 'SINGLE_TENANT_OLTP_SCALE',4, 2,4,3,2,2, 'Hot single tenant may not split naturally by tenant sharding key');INSERT INTO sh26_arch_eval VALUES( 'GLOBAL_DATA_SCALE',5, 2,3,3,5,4, 'Dataset/workload can be partitioned by many tenant keys');COMMIT;
The numeric scores are illustrative architecture workshop inputs, not Oracle product ratings. Replace them with evidence specific to the deployment.
9. Compute weighted totals but keep the evidence visible
SELECT architecture, SUM(weight*score) AS weighted_scoreFROM ( SELECT criterion,weight,'PARTITIONING' architecture, partitioning_score score FROM sh26_arch_eval UNION ALL SELECT criterion,weight,'RAC',rac_score FROM sh26_arch_eval UNION ALL SELECT criterion,weight,'DATA_GUARD',dataguard_score FROM sh26_arch_eval UNION ALL SELECT criterion,weight,'SHARDING',sharding_score FROM sh26_arch_eval UNION ALL SELECT criterion,weight,'APP_DISTRIBUTION', app_distribution_score FROM sh26_arch_eval)GROUP BY architectureORDER BY weighted_score DESC;
A weighted score starts a conversation; it cannot hide a hard requirement. If sovereignty is legally mandatory, an architecture that cannot physically satisfy it is disqualified even if its average score is high.
10. Explicit workload/SLO criteria
| Question | Likely implication |
|---|---|
| Can 95%+ of OLTP be keyed to one tenant/customer? | Sharding locality is plausible. |
| Is the problem only one-node failure? | RAC/service HA may fit better than sharding. |
| Need remote copy/RPO/RTO? | Data Guard/replication is central even if sharding is also used. |
| Need partition pruning/lifecycle but one failure domain is acceptable? | Partitioning may be enough. |
| Need per-tenant bespoke database versions/schema? | Application/federated distribution may fit, but loses uniform sharding assumptions. |
| Many cross-tenant atomic transactions? | Sharding key/model may fight the workload. |
11. Current licensing snapshot
The current 26ai matrix is not historical folklore:
- Globally Distributed Database: Free = yes, limited to three shards; EE/EE-ES = yes with shard-count conditions tied to RAC/Active Data Guard/other rights; BaseDB/ExaDB vary.
- Raft replication: Free = yes up to three nodes, each with Free limits; cloud/EE rules differ.
- RAC: not available in Free; on EE/EE-ES it is an extra-cost option where offered.
- Data Guard redo apply: not available in Free in the current matrix; available in EE/EE-ES and applicable cloud offerings.
- Active Data Guard: separately licensed/entitled depending on offering.
- Partitioning: included in Free; extra-cost on EE/EE-ES, with cloud offering differences.
Always re-read the current licensing manual before freezing a production design.
12. Deliberately wrong: choose sharding only because the table has billions of rows
A very large table whose hot transactions constantly join/update global rows may be a poor sharding candidate, while a smaller multi-tenant dataset with strong regional sovereignty may be an excellent one. Scale is necessary context, not sufficient justification. Partitioning, indexing, RAC, read replicas or application redesign may solve the actual bottleneck with much less distributed complexity.
13. Deliberately wrong: call three application databases “Oracle Sharding”
If the application manually maintains three unrelated databases with no shard catalog, shard directors, global services, SHARD DDL, chunk placement or Oracle routing metadata, that is application-level distribution. It may be valid, but it does not inherit Oracle sharding's management/routing/DDL/query-coordinator semantics.
14. Combined topology runbook checklist
- Document the unit of data ownership: tenant, region, account, time bucket or another domain key.
- Document the unit of failure: RAC instance/node, shard database, shard replica, region/site, catalog/director.
- Document the transaction boundary: local shard versus distributed 2PC/workflow.
- Document routing: global service/shard key versus application directory/host.
- Document backup/restore and point-in-time recovery for every data copy.
- Document patch/upgrade sequencing across catalog, directors, shards and replication.
- Load-test local and fan-out APIs separately.
- Run shard, catalog, director, region/network failure drills—not just happy-path scale tests.
15. Cleanup
DROP TABLE sh26_arch_eval PURGE;DROP TABLE sh26_partition_demo PURGE;
16. Production judgment and chapter close
Sharding is the mechanism that changes row ownership across independent databases. RAC changes instance availability around one database; Data Guard changes replica/site protection; partitioning changes table organization inside one database; application distribution moves the control plane into application/platform code. They can be combined, but every added layer multiplies monitoring, recovery, patching, security and skill requirements.
Current baseline is Oracle AI Database 26ai RU 23.26.3, SQL Developer 26.2 and SQLcl 26.2.1.222.1617. The current licensing matrix allows Globally Distributed Database in Free up to three shards and Raft up to three Free nodes, yet the traditional catalog guide documents an EE catalog PDB—so production topology/licensing must be validated for the exact deployment. The mandatory chapter remains a one-Free-instance simulation and makes no unsupported HA/performance promise. Chapter 27 can now approach migration/upgrade/patching with a clearer understanding that a distributed estate multiplies every lifecycle operation.
Check your understanding
- Which mechanism actually splits row ownership across independent databases?
- What does RAC change compared with sharding?
- What does Data Guard primarily solve?
- When is ordinary partitioning sufficient?
- Why can a combined sharding + RAC + Data Guard design be much harder to operate?
Review the answers
Oracle Globally Distributed Database / Sharding.
RAC provides multiple instances/nodes accessing one database rather than placing different row subsets in independent databases.
Replication-based database/site availability, disaster recovery and data protection.
When one database failure domain is acceptable and the need is pruning/manageability/storage lifecycle rather than geographic/failure-domain distribution.
Every shard can itself have clustered/replicated components, multiplying services, backups, patch sequencing, monitoring, failure modes and licensing.
Authoritative references
- Oracle Globally Distributed AI Database Overview — distributed partitioning/failure isolation/replication
- Licensing Information — Free/RAC/Data Guard/Partitioning/Sharding matrix
- Real Application Clusters Guide — RAC instance/database model
- Data Guard Concepts and Administration — primary/standby data protection
- Data Distribution Methods — sharding placement methods