Chapter 26 · Security: Authentication, Authorization, Roles, TLS, and Secrets
Enable Authentication and Authorization Safely with a Bootstrap / Recovery Plan
Stage authentication/authorization with a tested recovery identity instead of risking cluster-wide lockout.
Learning outcomes
AtlasMart must stop accepting anonymous CQL before production
launch. A rushed all-node edit can lock every application and
administrator out if system_auth, credentials or
permissions are wrong. Security rollout therefore includes
recovery design.
Explain AllowAllAuthenticator/AllowAllAuthorizer versus PasswordAuthenticator/CassandraAuthorizer.
Increase and repair system_auth replication before security depends on it.
Canary authentication on one drained node and prove anonymous failure/authenticated success.
Create and test an alternate break-glass administrator without plaintext secret leakage.
Disable the bootstrap login, stage authorization, roll all nodes and verify consistent behavior.
The mandatory work uses a separate disposable security cluster
so hardening experiments cannot lock the shared course
cluster. Pin Apache Cassandra 5.0.9 with the
official cassandra:5.0.9 image, Java 17 from that
image, Docker network
atlasmart-cassandra-security, cluster
atlasmart-security, nodes
atlasmart-sec-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes (vnodes) per
node, NetworkTopologyStrategy, replication factor
(RF) 3, and LOCAL_QUORUM for application-style
verification. Tables use UnifiedCompactionStrategy (UCS),
default_time_to_live=0, and
gc_grace_seconds=864000.
The lab intentionally starts with Cassandra's compatible
security defaults so the migration is observable:
AllowAllAuthenticator,
AllowAllAuthorizer, client TLS disabled,
internode encryption none, and JMX local-only. No
Cassandra/JMX/storage ports are published to the host. This is
a local learning boundary, not a production posture. Never
place a real password, token, private key, keystore password,
backup key, or break-glass credential into Git, lesson HTML,
screenshots, chat, shell history, or process arguments.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms before the security mechanisms
CQL is Cassandra Query Language. A coordinator is the node accepting a client request; replicas store copies of the target partition. A partition is the row group located by a partition key; its hash maps to a token. vnodes are multiple token ranges assigned per node. A keyspace defines replication. RF is replication factor and CL is consistency level. An SSTable is Cassandra's immutable on-disk table file set; compaction merges SSTables; repair compares replica ranges and streams differences. SAI is Storage-Attached Indexing. A driver is the client library speaking Cassandra's native protocol.
Authentication proves an identity; Cassandra's
authenticator implements that decision.
Authorization decides what an authenticated
identity may do; the authorizer checks
permissions on Cassandra resources. A role is
Cassandra's principal/inheritance object;
LOGIN=true makes it directly authenticatable. A
superuser bypasses normal authorization and
therefore belongs only in tightly controlled administration or
break-glass recovery. Least privilege grants
only the permissions needed.
TLS (Transport Layer Security) protects network
traffic and authenticates endpoints with certificates.
Client TLS protects application/cqlsh ↔
Cassandra native protocol;
internode TLS protects Cassandra node ↔ node
gossip/replication/streaming. mTLS (mutual TLS)
means both sides present certificates. A
keystore holds private key/certificate material
and a truststore holds trust anchors.
JMX (Java Management Extensions) is Cassandra's
management interface used by nodetool.
Defense in depth layers identity, permissions,
encryption, network isolation, secrets, backups, audit, patching
and recovery instead of trusting one control.
1. Why bootstrap order matters
| Stage | Mechanism | Evidence | Rollback |
|---|---|---|---|
| 0 | isolated cluster, no host-published ports | anonymous works only inside lab network | untouched source config |
| 1 | system_auth RF=3 | NTS dc1=3 + repair | all nodes UN |
| 2 | PasswordAuthenticator on canary | anonymous denied; bootstrap login succeeds | local console/JMX + config copy |
| 3 | new break-glass role | new admin LIST ROLES/query succeeds | vaulted secret |
| 4 | bootstrap login disabled | old login denied; alternate admin succeeds | tested alternate role |
| 5 | CassandraAuthorizer | allowed and denied operations match plan | pre-created grants |
| 6 | rolling rollout | same behavior on all coordinators | per-node config evidence |
PasswordAuthenticator stores password hashes in
system_auth.roles;
CassandraAuthorizer stores permission state in
system_auth.role_permissions. Cassandra's own
configuration warns to increase
system_auth replication when using these
mechanisms. A schema RF change does not retroactively populate
existing replicas, so repair is part of the rollout.
2. Build the sandbox and protect system_auth first
docker network inspect atlasmart-cassandra-security >/dev/null 2>&1 || docker network create atlasmart-cassandra-securityfor n in 1 2 3; do docker volume create atlasmart-sec-$n-data; donedocker run -d --name atlasmart-sec-1 --hostname atlasmart-sec-1 --network atlasmart-cassandra-security \ -e CASSANDRA_CLUSTER_NAME=atlasmart-security -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-sec-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only when node 1 is UN.docker exec atlasmart-sec-1 nodetool statusdocker run -d --name atlasmart-sec-2 --hostname atlasmart-sec-2 --network atlasmart-cassandra-security \ -e CASSANDRA_CLUSTER_NAME=atlasmart-security -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-sec-1 -v atlasmart-sec-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-sec-3 --hostname atlasmart-sec-3 --network atlasmart-cassandra-security \ -e CASSANDRA_CLUSTER_NAME=atlasmart-security -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-sec-1 -v atlasmart-sec-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-sec-1 nodetool versiondocker exec atlasmart-sec-1 java -versiondocker exec atlasmart-sec-1 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_secureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_secure.orders_by_customer ( customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction={'class':'UnifiedCompactionStrategy'} AND default_time_to_live=0 AND gc_grace_seconds=864000;CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_secure.orders_by_customer(customer_id,order_month,order_time,order_id,status,total,note)VALUES ('cust-42','2026-09-01','2026-09-08T10:00:00Z',00000000-0000-0000-0000-000000000261,'PAID',129.90,'security-fixture');SELECT customer_id,order_month,order_time,status,totalFROM atlasmart_secure.orders_by_customerWHERE customer_id='cust-42' AND order_month='2026-09-01';
ALTER KEYSPACE system_authWITH replication={'class':'NetworkTopologyStrategy','dc1':3};DESCRIBE KEYSPACE system_auth;
docker exec atlasmart-sec-1 nodetool repair --full system_authdocker exec atlasmart-sec-1 nodetool statusdocker exec atlasmart-sec-1 nodetool describecluster
3. Canary PasswordAuthenticator on one drained node
The official container's topology environment variables do not replace a reviewed security configuration. In this disposable lab, keep a rollback copy and edit only the canary node. Production would drain/remove it from client traffic and use configuration management.
docker exec atlasmart-sec-1 cp /etc/cassandra/cassandra.yaml /etc/cassandra/cassandra.yaml.pre-authdocker exec atlasmart-sec-1 sed -ri \ 's/^authenticator: AllowAllAuthenticator$/authenticator: PasswordAuthenticator/' \ /etc/cassandra/cassandra.yamldocker restart atlasmart-sec-1docker exec atlasmart-sec-1 nodetool status# Anonymous should now fail on the canary.docker exec atlasmart-sec-1 cqlsh --disable-history -e "SELECT release_version FROM system.local;" || true# Connect interactively as the documented bootstrap role. Enter the documented# bootstrap password at the prompt; never put it in a process argument.docker exec -it atlasmart-sec-1 cqlsh -u cassandra --disable-history
Expected evidence: the anonymous request is rejected with an authentication error; the bootstrap role succeeds interactively. This proves authentication on that coordinator, not authorization and not cluster-wide rollout.
4. Establish alternate administration before disabling bootstrap access
if (-not $env:ATLASMART_BREAKGLASS_PASSWORD) { throw "Set the secret from an approved source." }$hash = docker exec -e ATLASMART_BREAKGLASS_PASSWORD="$env:ATLASMART_BREAKGLASS_PASSWORD" ` atlasmart-sec-1 hash_password --environment-var ATLASMART_BREAKGLASS_PASSWORD@"CREATE ROLE IF NOT EXISTS atlasmart_breakglassWITH LOGIN=true AND SUPERUSER=true AND HASHED PASSWORD='$hash';LIST ROLES OF atlasmart_breakglass;"@ | Set-Content .\ch26-breakglass.cqldocker cp .\ch26-breakglass.cql atlasmart-sec-1:/tmp/ch26-breakglass.cqlRemove-Item .\ch26-breakglass.cql# SOURCE /tmp/ch26-breakglass.cql from an authenticated bootstrap cqlsh session.
ALTER ROLE cassandra WITH SUPERUSER=false AND LOGIN=false;LIST ROLES;
If role state or configuration is wrong, every path can be locked out at once. The safe mechanism is canary → alternate admin → prove login → disable bootstrap → authorization → rolling rollout. Keep local management/recovery access and reviewed rollback config throughout.
5. Stage authorization and remove coordinator-dependent drift
docker exec atlasmart-sec-1 sed -ri \ 's/^authorizer: AllowAllAuthorizer$/authorizer: CassandraAuthorizer/' \ /etc/cassandra/cassandra.yamldocker restart atlasmart-sec-1# Create/test Lesson 2 application grants on the canary before routing traffic.for n in 2 3; do docker exec atlasmart-sec-$n cp /etc/cassandra/cassandra.yaml /etc/cassandra/cassandra.yaml.pre-auth docker exec atlasmart-sec-$n sed -ri \ -e 's/^authenticator: AllowAllAuthenticator$/authenticator: PasswordAuthenticator/' \ -e 's/^authorizer: AllowAllAuthorizer$/authorizer: CassandraAuthorizer/' \ /etc/cassandra/cassandra.yaml docker restart atlasmart-sec-$n docker exec atlasmart-sec-1 nodetool statusdonefor n in 1 2 3; do docker exec atlasmart-sec-$n sh -lc \ "grep -E '^(authenticator|authorizer|role_manager|network_authorizer):' /etc/cassandra/cassandra.yaml"done
Check your understanding
- Why raise system_auth RF first?
- Why canary one drained node?
- Why use HASHED PASSWORD?
- When can the bootstrap login be disabled?
- Does authenticated access imply table access?
Review the answers
1. Authentication/authorization role state lives in system_auth; RF=1 is a fragile dependency.
2. It limits a lockout/configuration mistake to one coordinator path.
3. It keeps plaintext out of CQL files/history; the original secret still needs safe handling.
4. Only after an alternate administrator and recovery path are independently proven.
5. No. CassandraAuthorizer separately evaluates permissions.
Production judgment
Security controls affect latency, availability and operator skill. Record Cassandra/JVM/driver versions, DC/racks/RF/CL, auth/TLS/JMX/network state, role inheritance, credential/certificate expiry, audit retention, backup encryption/access, patch status, repair/restore credentials and the application routing/retry/idempotency model. Re-run p50/p95/p99 latency and replica-failure tests after encryption/authorization changes. Data modeling still matters: partition/cardinality/TTL/tombstone/compaction/repair/SAI/vector choices can change sensitive-data exposure and resource-denial risk.
Do not treat a private subnet, one superuser password, client TLS or one audit stream as sufficient. Managed Cassandra may operate some server certificates, JMX, upgrades or backups, but application identities, permissions, client trust, data classification, secret use and incident response still have explicit owners. Migration/rollback must retain a tested administrative path and prevent coordinator-dependent security drift. Lesson 2 replaces superuser-centric operation with reusable inherited roles, narrow grants and explicit denial tests.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Roles, LOGIN, Permissions, GRANT/REVOKE, Role Inheritance, and Least Privilege.
Authoritative references
Security defaults, algorithms, tooling and deprecations are version-sensitive. Re-check the target Cassandra/JVM/driver distribution and organizational cryptographic policy before production rollout.