Chapter 03 · Partitioners, Tokens, Vnodes, Token Rings, and Data Distribution

Virtual Nodes: Many Tokens per Node, Distribution, Streaming, and Operational Tradeoffs

See why vnodes improve incremental distribution, why modern Cassandra uses far fewer than the historical 256 by default, and what that flexibility costs operationally.

Intermediate115–140 minutesVnode allocation and bootstrap-streaming labApache Cassandra 5.0.9 · Murmur3 · 16 vnodes/nodeLast reviewed: September 2026

Learning outcomes

A single token per node makes consistent hashing easy to draw but awkward to expand incrementally. Virtual nodes (vnodes) give each physical node many token positions, allowing one joining node to receive many smaller ranges from multiple peers. Cassandra 5.0 also makes the vnode count an explicit availability/elasticity tradeoff rather than “more is always better.”

01

Explain a vnode as a token position owned by a physical host ID, not a separate Cassandra process.

02

Inspect 16 tokens per course node and connect them to multiple token ranges.

03

Explain how vnodes spread bootstrap/decommission streaming across more peers.

04

Compare modern low-token guidance with historical 256-vnode defaults without mixing versions.

05

Identify availability, repair, metadata, and range-operation costs that increase with token count.

Chapter baseline reviewed 7 September 2026

Apache Cassandra 5.0.9 is the current GA 5.0 patch on the official download page. The labs pin cassandra:5.0.9 and explicitly set CASSANDRA_NUM_TOKENS=16 so vnode behavior is reproducible instead of inheriting an unnoticed image/configuration default. Current Cassandra 5.0 documentation uses num_tokens: 16 as the modern baseline; older Cassandra material often mentions 256 random vnodes, so this chapter treats token count as a version- and deployment-sensitive design choice rather than folklore.

Execution and evidence note

The generation environment does not contain Docker or Cassandra, so Cassandra commands were checked against current official documentation but were not executed here. Exact token values, IP addresses, host IDs, ownership percentages, load, stream sizes, latency, and hot-partition samples must be captured on the learner's machine. Expected output is described as invariants or shapes, never presented as measured output.

1. A vnode is an ownership position, not another server

A virtual node is one token on the ring assigned to a physical Cassandra node. Sixteen vnodes do not launch sixteen JVMs or create sixteen host IDs. They give one endpoint sixteen positions, and therefore sixteen primary range endpoints spread through the hash space.

This changes incremental scale-out. With one token per host, adding a single machine to an evenly spaced small ring can leave the cluster imbalanced unless token positions are carefully designed. With multiple token positions, the new node can take small ranges distributed across the ring. That reduces the need to double cluster size merely to keep a simple single-token ring balanced.

bash / PowerShell · reproducible 3-node, 16-vnode lab
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all nodes to become Up/Normal.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool version
evidence · one host ID, many token positions
docker exec atlasmart-cass-1 sh -lc "grep '^num_tokens:' /etc/cassandra/cassandra.yaml"docker exec atlasmart-cass-1 cqlsh -e "SELECT host_id, tokens FROM system.local;"docker exec atlasmart-cass-1 nodetool status# Verbose, but useful once in this lesson:docker exec atlasmart-cass-1 nodetool ring | head -80

Record the host ID and count the token values returned by system.local.tokens. The expected course invariant is one host identity with sixteen tokens.

2. Why current Cassandra guidance moved away from 256 as a universal default

Older Cassandra vnode implementations used random token selection and historically defaulted to 256 tokens per node to smooth imbalance statistically. Current Cassandra documentation describes a deterministic allocator and recommends much smaller counts, with num_tokens: 16 as the documented configuration baseline. Production recommendations explicitly describe tradeoffs among 1, 4, 8, and 16 tokens and warn that more tokens can reduce availability in larger clusters.

There is no contradiction once version and allocator are named. The correct lesson is not “256 is wrong” or “16 is always right”; it is that token allocation algorithms and operational guidance evolved. Re-check the exact stable release and workload before freezing cluster-wide token policy.

Token count effect Benefit Cost/risk
More vnode positions Finer-grained incremental rebalancing More neighbor relationships and metadata/range work
More streaming peers Bootstrap can draw smaller ranges from many sources More concurrent disk/network relationships to manage
More token ranges per node Random/allocated imbalance can be reduced Repair and range-oriented operations become more fragmented
Fewer tokens Simpler neighbor/availability relationships Expansion can become less flexible if too coarse

3. Token allocation should understand replication, not only raw ring spacing

Modern cassandra.yaml exposes allocate_tokens_for_local_replication_factor and an optional allocate_tokens_for_keyspace. The allocator can choose token positions with replicated load in mind rather than blindly sampling random points. This matters because balanced primary token ranges do not automatically imply balanced replicated load under rack-aware RF.

configuration evidence · vnode count and replication-aware allocation
docker exec atlasmart-cass-1 sh -lc "grep -E '^(num_tokens:|allocate_tokens_for_local_replication_factor:|# allocate_tokens_for_keyspace:)' /etc/cassandra/cassandra.yaml"# Inspect actual keyspace replication before reasoning about effective load.docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_tokens';"

Do not casually set initial_token by hand in a new vnode cluster. Current configuration guidance says manual initial tokens are primarily for legacy/special cases, while the generatetokens tool exists for edge cases where positions must be known before bootstrap. Automatic allocation is the ordinary path.

4. Streaming is the data-plane consequence of changing ownership

When a new node bootstraps, it receives tokens and becomes responsible for their ranges. Cassandra must then stream the corresponding replicated data from existing nodes. Streaming is not a metadata-only event: it consumes disk reads/writes, network bandwidth, CPU, compaction headroom, and time. nodetool netstats and bootstrap status are key evidence.

With vnodes, one new host usually obtains many ranges from multiple existing hosts. That spreads the transfer, but it can also create broad resource pressure. Operators therefore throttle/plan streaming and observe tail latency rather than assuming scale-out is free because the final state has more capacity.

controlled scale-out · observe bootstrap/streaming
# Before a topology change, capture a baseline.docker exec atlasmart-cass-1 nodetool status atlasmart_tokensdocker exec atlasmart-cass-1 nodetool netstats -Hdocker volume create atlasmart-cass-4-datadocker run -d --name atlasmart-cass-4 --hostname atlasmart-cass-4 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-4-data:/var/lib/cassandra cassandra:5.0.9# While node 4 joins, sample from node 4 and an existing peer.docker exec atlasmart-cass-4 nodetool netstats -Hdocker exec atlasmart-cass-1 nodetool status atlasmart_tokens# After bootstrap completes, record the final state.docker exec atlasmart-cass-4 nodetool status atlasmart_tokens

The exact stream totals depend on how much data exists and which token ranges node 4 receives. A tiny fixture may stream very little; this lab is about the state transition and evidence path, not a throughput benchmark.

5. Production judgment: vnode count is an availability and operations choice

More vnode positions make a physical-node failure intersect more independent token ranges and peer relationships. The official architecture documentation notes that token count can increase combinations of node failures that make portions of the ring unavailable and can slow cluster-wide maintenance such as repair. Therefore vnode count belongs in architecture review alongside RF, rack count, cluster size, expansion pattern, repair strategy, and failure targets.

AtlasMart uses 16 in this chapter because it matches current 5.0 guidance and makes the mechanics visible. That is a lab baseline, not a universal production tuning value. If a real cluster is expected to exceed the scale ranges called out by current production guidance, the team should model alternative token counts before rollout.

Verification checklist

  • You proved one node owns multiple tokens without confusing vnodes with processes.
  • You recorded the chapter's explicit num_tokens=16.
  • You can explain why historical 256-vnode advice is version-contextual.
  • You observed or prepared the evidence path for bootstrap streaming with netstats.
  • You can name at least two costs that increase as token count rises.

Check your understanding

  1. Is a vnode a separate Cassandra JVM?
  2. Why did older deployments often use many more tokens?
  3. Why can more tokens reduce availability?
  4. What evidence shows bootstrap is moving data?
  5. Should you copy 16 into every production cluster?
Review the answers

1. No. It is a token position/range owned by a physical node/host ID.

2. Random token allocation needed a high count to smooth imbalance; modern deterministic allocation supports lower counts.

3. A node participates in more range/neighborhood relationships, creating more combinations where concurrent failures affect a range.

4. nodetool netstats/bootstrap state and resulting token/ownership changes, plus measured disk/network activity.

5. No. It is current guidance/baseline, but cluster size, RF, elasticity and availability goals must be evaluated for the actual release and workload.

Cleanup

cleanup · Bash; remove only course-owned resources
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
cleanup · PowerShell; remove only course-owned resources
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Hot Partitions, Skewed Keys, Monotonic Workloads, and Why More Nodes Do Not Fix Bad Keys.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.