Cockroach Labs Estate economics What idle nodes, provisioned bytes and peak-sized bills cost, worked out from your estate.

Estate summary

  1. Nodes 31 down from 75
  2. Avg CPU utilization 36.3% up from 15.0%
  3. Storage provisioned 1.5 TB 250 GB of data
  4. Est. monthly saving 58% $1,000/m becomes $420/m

Step 1 of 8 - Virtualization

You're running an estate, not a cluster

Every cluster below was sized for its own peak, so every one of them idles.

Deployment model
1 private host cluster, 25 virtual clustersPrivate Host cluster31 nodes on average, 38 at peakvc-01own connection stringvc-02own connection stringvc-03own connection stringvc-04own connection stringvc-05own connection stringvc-06own connection stringvc-07own connection stringvc-08own connection stringvc-09own connection stringvc-10own connection stringvc-11own connection stringvc-12own connection stringvc-13own connection stringvc-14own connection stringvc-15own connection stringvc-16own connection stringvc-17own connection stringvc-18own connection stringvc-19own connection stringvc-20own connection stringvc-21own connection stringvc-22own connection stringvc-23own connection stringvc-24own connection stringvc-25own connection stringShared, auto-scaled node pool31 of 38 nodes at peakEstate after consolidationOne private host cluster carrying 25 virtual clusters on 31 nodes on average and 38 at peak, averaging 36.3% CPU utilization.
36%Average CPUAverage CPU CPU utilization36.3% of provisioned CPU is in use.

31 nodes running 90 vCPU of actual work.

Pooled and sized once for the peak, 38 nodes would average 29.6%. Auto-scaling brings it to 31 nodes at 36.3%, moving between 29.5% and 39.2% across the day against a 38% setpoint.

The arithmetic, in full
StepWorkingResult
Provisioned today25 clusters of 3 nodes each, 75 in total x 8 vCPU600 vCPU
Mean demand600 vCPU x 15.0% estate average utilization (per-cluster 5.4% to 24.3%, sd 4.6%)90 vCPU
Aggregate peak90.0 vCPU x (1 + (2.23 - 1) / sqrt(25)), a pooled peak-to-mean of 1.246112 vCPU
Consolidated, sized once for the peakmax(3, ceil(112.1 vCPU / (8 vCPU x 37.5% target))) = max(3, ceil(37.37)); the mean sits well under that peak38 nodes, averaging 29.6%
Auto-scaled through the daythe pool follows the pooled curve in whole nodes, 30 minutes behind it, holding its size for 2 hours before shrinking31 nodes on average, at 36.3% CPU
Utilization across the day29.5% to 39.2% against a 37.5% setpoint: the pool runs hot while it's a bucket behind a rise, and cool while the cooldown is still holding capacity after a peak29.5%-39.2%

Assumptions

  • 8 vCPU per node. Standalone clusters are drawn around 3 nodes, snapped to an odd number for quorum and held between 3 and 15; in this estate they're all 3, because a spread centred on the end of that range has nowhere to go, for 75 nodes in total.
  • Each standalone cluster's CPU is a draw from a normal distribution centred on 15.0% and held between 0% and 100%; in this estate they run 5.4% to 24.3%, a 4.6-point standard deviation. The published band is 10-15% today against a 30-45% auto-scaled target, and this Private Host Cluster is set to 37.5%.
  • The Private Host Cluster floors at 3 nodes and must reach 38 at peak.
  • The auto-scaler sizes for what it saw 30 minutes ago and waits 2 hours before shrinking, so the pool averages 36.3% rather than sitting on its 37.5% setpoint.
  • Workloads are independent; consolidation moves where work runs, not how much there is.
What a virtual cluster gives an app team

Its own connection string, databases, schemas, users, roles and backups, and no visibility into the Private Host Cluster or any other virtual cluster.