Managed changes the boundary of operations; it does not transfer responsibility for correctness.
Managed Control Planes vs Self-Managed Data Planes: Responsibility Boundaries
Draw an explicit shared-responsibility boundary for a managed NoSQL service so provider-operated infrastructure is not confused with customer-owned data modeling, IAM, recovery validation, or incident decisions.
Distinguish provider control-plane/data-plane operation from AtlasMart responsibilities for application correctness, IAM, tenant isolation, schemas, queries, and recovery acceptance.
Build a shared-responsibility matrix covering patching, replication, backups, capacity, encryption, networking, indexes, monitoring, and incident response.
Observe how an over-privileged application identity can cause a destructive managed-service incident even when provider infrastructure is healthy.
Define evidence and runbooks for the responsibilities that remain with the customer after moving to a managed or serverless database.
1. A managed database removes toil, not accountability
AtlasMart is considering a managed NoSQL service because the platform team no longer wants to replace failed hosts, patch database binaries, or hand-operate replica membership. Those are valid reasons. The dangerous leap is to translate managed into “the provider owns the database outcome.” A control plane is the API/service layer that creates, configures, scales, backs up, and observes database resources. The data plane is the path that stores and serves application reads and writes. A managed provider can operate both underlying planes while AtlasMart still owns the application’s data model, access policy, request routing, business invariants, and acceptance criteria for recovery.
A responsibility boundary must be expressed as an operational contract. “Encryption is managed” is incomplete: who selects customer-managed keys, who can disable them, what happens if key access is lost, and which backups depend on them? “Backups are automatic” is incomplete: what retention is configured, who has restore permission, whether restore creates a new resource, and who verifies tenant counts and invariants? The same questions apply to autoscaling, global replication, IAM, network endpoints, indexes, schema changes, and incident escalation.
2. Build the shared-responsibility matrix before the first production write
Start from work, not vendor labels. For every capability record the provider duty, AtlasMart duty, shared hand-off, observable evidence, and failure owner. The provider may patch the service fleet while AtlasMart owns client-library compatibility. The provider may replicate bytes while AtlasMart owns whether a partition key creates a hot tenant. The provider may expose point-in-time recovery while AtlasMart owns the restore point, dependency order, validation queries, and traffic cutover.
| Capability | Managed/provider side | AtlasMart/customer side | Evidence |
|---|---|---|---|
| Host/engine operation | Physical fleet, service binaries, failure replacement according to service contract | Review maintenance/version notices and client compatibility | service health, maintenance events, client errors |
| Replication | Operate configured replication topology | Choose region/consistency mode; design conflict-safe workload | replication lag, replica status, client history |
| Capacity | Execute provisioned/autoscale/serverless mechanism | Keys, traffic shape, quotas, budgets, throttling policy | request units, throttles, hot-key metrics, bill |
| Security | Service infrastructure controls and supported encryption primitives | IAM, tenant isolation, network policy, secrets, key policy where customer-managed | audit trail, denied/allowed actions, endpoint policy |
| Recovery | Operate configured backup/PITR mechanism | RPO/RTO, retention selection, restore drill, validation and cutover | restorable point, restore test, invariants |
“No servers to manage” says who handles infrastructure provisioning. It does not imply infinite throughput, zero throttling, global strong consistency, automatic tenant isolation, free egress, or a tested restore path.
3. Deliberately wrong approach: give the application the service administrator role
A common shortcut is to create one credential for both application and operations. That works until an application bug, leaked secret, or compromised runtime calls a destructive control-plane API. The provider can authenticate and correctly authorize that request; from the provider’s perspective the managed service behaved as configured. The failure is AtlasMart’s role design. The safer pattern gives the request path only data-plane permissions it needs, separates backup/monitoring identities, and reserves destructive administration for audited break-glass roles.
4. AtlasMart lab: make the responsibility boundary executable
Python 3.13+ standard library only. No cloud account, credentials, provider SDK, or destructive resource is used. The role names and responsibility rows are a teaching model, not a provider IAM policy.
from dataclasses import dataclass
@dataclass(frozen=True)
class Request:
actor: str
action: str
# Managed service: provider operates physical hosts and service replication.
# AtlasMart still owns application identity, data access, query design, and recovery validation.
wrong_role = {
"catalog-api": {"read", "write", "drop_table", "change_replication"},
"backup-job": {"read", "export"},
}
fixed_role = {
"catalog-api": {"read", "write"},
"backup-job": {"read", "export"},
"break-glass-admin": {"drop_table", "change_replication"},
}
requests = [
Request("catalog-api", "read"),
Request("catalog-api", "drop_table"), # compromised app credential
]
def authorize(roles, request):
return request.action in roles.get(request.actor, set())
print("WRONG: 'managed means the provider owns everything'")
for req in requests:
print(req, "allowed=", authorize(wrong_role, req))
print("result: compromised application identity can execute a destructive admin action")
print("\nFIXED SHARED-RESPONSIBILITY BOUNDARY")
for req in requests:
print(req, "allowed=", authorize(fixed_role, req))
responsibility = {
"physical host replacement": "provider",
"service engine patching": "provider/service-specific",
"replication service operation": "provider/service-specific",
"partition key and query shape": "AtlasMart",
"application IAM": "AtlasMart",
"tenant isolation": "AtlasMart",
"backup policy selection": "shared/service-specific",
"restore testing": "AtlasMart",
"incident response": "shared",
}
for item, owner in responsibility.items():
print(f"{item:30} -> {owner}")
print("\nverification: app credential is no longer an implicit cluster administrator")
Under the broken role, catalog-api can execute
drop_table. Under the repaired role it can only
read/write application data; destructive control moves to an
explicit break-glass identity. This proves that managed
infrastructure does not compensate for an over-privileged
customer identity.
5. Production judgment: manage the hand-offs as interfaces
A managed database is valuable when the provider can perform undifferentiated operations more reliably or economically than the team, but responsibility moves rather than disappears. Record escalation paths, regional dependencies, service quotas, maintenance/update behavior, client compatibility, backup permissions, and evidence required before declaring an incident recovered. For regulated or multi-tenant data, the customer must still prove authorization and residency behavior end to end.
Operationally, the strongest shared-responsibility matrix is testable: a role-policy test rejects forbidden actions; a restore drill proves backup use; a load test proves throttling behavior; a private-endpoint test proves traffic cannot escape the expected boundary. The next lesson adds the other managed abstraction people often over-trust—automatic capacity—and shows that billing and hot-partition mechanics remain workload dependent.
Check your understanding
- What does a managed database provider usually remove from the customer’s day-to-day work?
- Why does managed service health not prove application correctness?
- Why is a shared superuser dangerous even on a serverless database?
- Who owns restore acceptance?
- What should a responsibility matrix contain besides security?
Review the answers
1. It can remove or automate parts of host, engine, replication, patching, and service-fleet operation; the exact boundary is service-specific.
2. The customer still controls identities, tenant filters, schemas/indexes, query behavior, data semantics, client configuration, and many recovery decisions.
3. A compromised application credential can inherit destructive administrative capabilities that the application never needed.
4. The customer must define recovery objectives and prove restored data/app behavior, even if the provider operates the backup mechanism.
5. Capacity, backup/restore, schema/index design, observability, incident escalation, networking, upgrades, and ownership of business invariants.
References
Provider-specific claims below are current implementation anchors reviewed in August 2026. They are not universal NoSQL definitions, and mandatory labs do not require the services.
- AWS — Shared Responsibility Model — Official statement separating AWS infrastructure responsibilities from customer responsibilities.
- AWS Well-Architected — Shared responsibility — Current managed-service examples, including patch-management responsibility differences.
- Microsoft Azure — Shared responsibility in the cloud — Official PaaS/IaaS/SaaS responsibility framing.
- Azure Reliability — Shared responsibility — Official reliability-focused boundary between platform and workload responsibilities.