Cockroach Labs Estate economics What idle nodes, provisioned bytes and peak-sized bills cost, worked out from your estate.

Estate summary

  1. Nodes 3 baseline
  2. Avg CPU utilization 15.0% idle hardware
  3. Storage provisioned 60 GB plus 10 GB in S3
  4. Est. monthly saving 12% $1,000/m becomes $880/m

Step 3 of 8 - Storage

Pay for bytes, not boxes

One change drives the next three pages: the persistent state stops living on the node.

Capacity provisioned to hold 10 GB of dataBlock storage on each node60 GB on disk3 copies of the datafree space, bought and idle6xPlenum60 GB on disk + 10 GB in object storage3 copies of the datain S3free space, not on the bill7xbilled: 16 GB, 40% of an untiered used meter24 GB the cold tier never billsuntiered, on a used meter: 40 GBCapacity provisioned per byte of dataBlock storage on each node provisions 6 bytes of capacity per byte of data, all of it on disk. Plenum provisions 7, of which 1 is a copy in object storage. Billed as used, 1.6 of those 7 bytes reach the bill: the copies on disk and the object storage backstop, but not the free space. Tiering 80% of the data down to 1 copy in object storage takes the bill from the 40 GB an untiered used meter would charge to 16 GB.
Block storage on each noden1n2n3diskdiskdiskThe bytes belong to the node:add, drain or lose one and they move.Plenumn1n2n3PlenumS3 backstop copyThe bytes belong to the layer:a node attaches to a store, then lets go.Where the bytes live under each storage modelWith Block storage on each node, each node owns the disk holding its store, so scaling and failure both copy data. With Plenum, the stores sit in a shared layer with a copy in object storage, and a node attaches to a store rather than holding it.

1 cluster x 10 GB = 10 GB of logical data becomes 60 GB of provisioned capacity, plus 10 GB held in object storage.

StepWorkingResult
Logical data1 cluster x 10 GB10 GB
Provisioned capacity10 GB x 6 bytes on disk per logical byte60 GB
Backstop in object storage10 GB x 1 copy10 GB
Metered as used10 GB x 4 copies on disk per logical byte40 GB
Cold tier80% of the data keeps 1 copy in object storage instead of 4 on disk8 GB
Still on disk20% keeps all 48 GB
Metered after tiering8 GB cold + 8 GB hot16 GB
Not billedcapacity the estate holds and nobody pays for: 54 GB of 70 GB54 GB
Storage bill16 GB metered against 40 GB kept hot, at the same rate either way40.0%

3 shared copies at 50% utilization (3/0.5 = 6), plus 1 copy in S3. Consolidate to lift utilization to 75%. Billed as used, and tiered: 16 GB reaches the bill, out of 70 GB of capacity.

What reaches the bill

Plenum bills what you hold: what reaches the bill is bytes the estate is actually holding, and the 54 GB of capacity that isn't holding any doesn't. The capacity is unchanged - what's changed is who carries the idle part of it.

Metered by the hour, not by the month: the figure above is a month of holding this much, and a volume that arrives halfway through the month is charged for half of it. Each window is reduced to its lowest reading, so compaction and rebalancing don't land on the bill.

You won't see a storage estimate for Plenum in the console: usage-based bytes aren't known when a cluster is sized, so what's quoted there is compute and storage arrives on the invoice. The storage figures here are this demo's model of what the meter would read.

Single region, and not on a Private Host Cluster: a host provisions the capacity its virtual clusters draw from, so the host itself isn't a Plenum cluster. Metered but not charged while it's in preview.

Tiered: 80% of the data stops keeping copies on disk and lives in object storage instead, so it bills as one copy rather than four. Nothing is deleted and no capacity comes back - the meter simply has less to count. 16 GB reaches the bill instead of 40 GB, which is 40% of keeping the whole estate on disk and takes the estimated cost above from 89% of the baseline to 88%.

The rate doesn't change: a byte in object storage is metered at what a copy on disk is metered at. Tiering is worth something because it removes bytes from the meter, not because it makes them cheaper.

The ratio is the point, not the size: the volume is the one you described on the front page, and every figure on the following pages follows it, including what a rebalance moves and what a dead node has to re-read.

Where the saving actually comes from
  1. Double replication eliminated. Hard links let all Raft replicas of a range share one copy, so Raft's 3x stops stacking on the storage layer's own replication. Plenum still replicates across zones.
  2. Local NVMe instead of network block storage. Reserved EC2 discounts run around 50%, EBS around 25%. Instance-local NVMe takes the better rate.
  3. Storage utilization. Provisioning storage independently of compute: about 50% to about 75%.
  4. Compute auto-scaling. Stateless KV nodes come and go freely: 10-15% average CPU to 30-45%.

On a used meter the third of these has already happened as far as the bill is concerned: lifting utilization fills capacity nobody was charged for, so it improves what the estate costs to run rather than what it costs to buy.