Chapter 26 · Security: Authentication, Authorization, Roles, TLS, and Secrets
Operational Security Review: Patch Cadence, Audit Evidence, Incident Access, and Multi-Tenant Boundaries
Operate security continuously through patching, audit, incident access and tenant-isolation reviews.
Learning outcomes
Launch-day hardening decays: roles accumulate privileges, certificates approach expiry, one node misses audit configuration, and a stale image remains deployed. Chapter 26 ends by turning security into an evidence-driven operating loop.
Capture Cassandra/JVM/image/driver patch inventory and review cadence.
Enable/inspect per-node audit logging and understand coordinator-local records.
Rehearse break-glass access and post-incident cleanup/revocation.
Assess multi-tenant isolation beyond CQL permissions.
Produce a recurring control/evidence/owner/expiry register for operations.
The mandatory work uses a separate disposable security cluster
so hardening experiments cannot lock the shared course
cluster. Pin Apache Cassandra 5.0.9 with the
official cassandra:5.0.9 image, Java 17 from that
image, Docker network
atlasmart-cassandra-security, cluster
atlasmart-security, nodes
atlasmart-sec-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and LOCAL_QUORUM for application-style
verification. Tables use UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
The lab intentionally starts with Cassandra's compatible
security defaults so the migration is observable:
AllowAllAuthenticator,
AllowAllAuthorizer, client TLS disabled,
internode encryption none, and JMX local-only. No
Cassandra/JMX/storage ports are published to the host. This is
a local learning boundary, not a production posture. Never
place a real password, token, private key, keystore password,
backup key, or break-glass credential into Git, lesson HTML,
screenshots, chat, shell history, or process arguments.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms before the security mechanisms
CQL is Cassandra Query Language. A coordinator is the node accepting a client request; replicas store copies of the target partition. A partition is the row group located by a partition key; its hash maps to a token. vnodes are multiple token ranges assigned per node. A keyspace defines replication. RF is replication factor and CL is consistency level. An SSTable is Cassandra's immutable on-disk table file set; compaction merges SSTables; repair compares replica ranges and streams differences. SAI is Storage-Attached Indexing. A driver is the client library speaking Cassandra's native protocol.
Authentication proves an identity; Cassandra's
authenticator implements that decision.
Authorization decides what an authenticated
identity may do; the authorizer checks
permissions on Cassandra resources. A role is
Cassandra's principal/inheritance object;
LOGIN=true makes it directly authenticatable. A
superuser bypasses normal authorization and
therefore belongs only in tightly controlled administration or
break-glass recovery. Least privilege grants
only the permissions needed.
TLS (Transport Layer Security) protects network
traffic and authenticates endpoints with certificates.
Client TLS protects application/cqlsh ↔
Cassandra native protocol;
internode TLS protects Cassandra node ↔ node
gossip/replication/streaming. mTLS (mutual TLS)
means both sides present certificates. A
keystore holds private key/certificate material
and a truststore holds trust anchors.
JMX (Java Management Extensions) is Cassandra's
management interface used by nodetool.
Defense in depth layers identity, permissions,
encryption, network isolation, secrets, backups, audit, patching
and recovery instead of trusting one control.
1. Patch cadence is a security control
The course pins Cassandra 5.0.9 for deterministic reproduction. Production should remain on a maintained release line and evaluate Apache Cassandra advisories/release notes, Java runtime/base-image fixes, driver compatibility and dependency/license boundaries. Rolling patch tests must include authentication, authorization, TLS, native protocol, repair/streaming, backup/restore and application SLOs.
docker exec atlasmart-sec-1 nodetool versiondocker exec atlasmart-sec-1 java -versiondocker exec atlasmart-sec-1 cqlsh --versiondocker inspect cassandra:5.0.9 --format '{{json .RepoDigests}}'# Record the deployed digest and separately review ASF/JDK/base-image/driver advisories.
2. Audit logging is per-node coordinator evidence
Cassandra audit logging records security/CQL activity on the coordinator that handled the request. It is enabled/configured per node and is not replicated like application data. Central compliance evidence therefore requires every intended node plus durable off-node collection, retention and tamper controls. It also does not replace configuration-management history or records of JMX/nodetool actions.
for n in 1 2 3; do docker exec atlasmart-sec-$n nodetool enableauditlog --included-categories AUTH,DCL,DDL,DML docker exec atlasmart-sec-$n nodetool getauditlogdone# Generate one controlled failed login and one allowed/denied query, then inspect# the coordinator's binary audit files with the documented auditlogviewer.# Exact file names and records are learner-captured runtime evidence.
for n in 1 2 3; do docker exec atlasmart-sec-$n nodetool disableauditlog docker exec atlasmart-sec-$n nodetool getauditlogdone
3. Break-glass incident game day
PRECONDITIONS- documented incident trigger/approval- break-glass credential stored in approved vault- audit collection healthy- exact task, stop condition and rollback knownDRILL1. Retrieve the credential through approved workflow.2. Connect only from approved management path.3. Start with read-only diagnostics.4. Record any mutation and its rollback.5. End the session and rotate/revoke temporary access as policy requires.6. Verify audit/config-change evidence.7. Remove temporary secret material.8. Review whether normal least-privilege roles can replace future superuser use.
It is neither emergency-only nor attributable. Use named/time-bounded approval where possible, vault retrieval, session/audit evidence and post-use cleanup, plus a documented recovery path for identity-provider/secret-store outages.
4. Multi-tenant boundaries exceed permissions
Roles/keyspaces are necessary, but mutually untrusted tenants can still share CPU, heap/off-heap, disk, compaction/repair, caches, queues, network, JMX/admin blast radius and failure domains. SAI/vector indexes add shared index/resource pressure and may expose sensitive derived retrieval patterns if modeled/authorized poorly. Evaluate separate keyspaces plus resource/network controls, separate clusters/accounts, or managed-service tenancy guarantees according to threat/compliance needs.
| Domain | Evidence | Red flag |
|---|---|---|
| roles | LIST ROLES/PERMISSIONS + denial tests | superuser/ALL grant sprawl |
| TLS | clients view + cert inventory + internode state | plaintext/expiring/untrusted path |
| JMX/network | listeners + policy + negative reachability | broad remote management access |
| audit | every node + central collection | one node disabled or local-only ephemeral logs |
| backups | vault IAM/KMS/integrity/restore test | backup reader broader than DB reader |
| patching | Cassandra/JDK/image/driver inventory | unsupported/stale release |
| tenancy | permission + load/failure isolation tests | one tenant exhausts shared resources/admin blast radius |
5. Chapter 26 acceptance register
CONTROL EXPECTED EVIDENCEAuthentication anonymous denied; intended roles succeedAuthorization least-privilege allow + explicit denial testsBootstrap admin default bootstrap login disabledsystem_auth RF/topology/repair healthClient TLS ssl_enabled/protocol/cipher + trust validationInternode TLS encrypted peer state + restart/repair/streaming testJMX/network local-only or authenticated+TLS+network-limitedSecrets source/owner/rotation; no Git/history leakageBackups restricted/encrypted/audited + restore drillAudit every intended node + durable off-node collectionPatching maintained Cassandra/JDK/image/driverIncident access break-glass game day + cleanup/revocation evidenceTenant boundary permission + resource/failure isolation evidence
Check your understanding
- Why is audit on only two nodes incomplete?
- Why is patch cadence part of security?
- Why can keyspace permissions be insufficient for hostile tenants?
- What must a break-glass drill prove after use?
- What is defense in depth across this chapter?
Review the answers
1. Requests coordinated by the third node generate their audit record there; audit is not replicated.
2. Security fixes arrive through maintained software; TLS/authorization do not remove vulnerable code.
3. They still share resources, management plane and operational blast radius.
4. Audit/change evidence, secret cleanup/rotation or revocation, rollback and return to least privilege.
5. Verified independent controls for identity, authorization, encryption, network/management, secrets, backups, audit, patching and recovery.
Production judgment
Security controls affect latency, availability and operator skill. Record Cassandra/JVM/driver versions, DC/racks/RF/CL, auth/TLS/JMX/network state, role inheritance, credential/certificate expiry, audit retention, backup encryption/access, patch status, repair/restore credentials and the application routing/retry/idempotency model. Re-run p50/p95/p99 latency and replica-failure tests after encryption/authorization changes. Data modeling still matters: partition/cardinality/TTL/tombstone/compaction/repair/SAI/vector choices can change sensitive-data exposure and resource-denial risk.
Do not treat a private subnet, one superuser password, client TLS or one audit stream as sufficient. Managed Cassandra may operate some server certificates, JMX, upgrades or backups, but application identities, permissions, client trust, data classification, secret use and incident response still have explicit owners. Migration/rollback must retain a tested administrative path and prevent coordinator-dependent security drift. Chapter 27 turns these controls and all prior Cassandra mechanisms into observability, capacity/guardrails, rolling-upgrade/multi-DC game days and the final production capstone.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Metrics, Logs, Tracing, nodetool tablestats/tablehistograms/proxyhistograms, and Alerting Signals.
Authoritative references
Security defaults, algorithms, tooling and deprecations are version-sensitive. Re-check the target Cassandra/JVM/driver distribution and organizational cryptographic policy before production rollout.