Capacity overview
What the capacity module measures, which data it rests on, and how the overview ranks what is tight right now.
Every number in this module is allocation as configured — vCPU, memory and disk as assigned to a machine, plus physical placement facts — and not actual utilization. There is no CPU-busy measurement here, no active memory and no IOPS, and there will not be without an agent on the machine. The Model: allocation chip sits in the header of every screen in the module and cannot be dismissed, so an allocation number is never read as a load measurement.
Where the numbers come from
| Source | What it contributes |
|---|---|
| The inventory | Host names, systems and owners |
| Physical placement fields | Cluster, the host a virtual machine runs on, rack, rack unit, physical memory and average power draw |
| The rack register | Rack height, amps, volts and weight limit per rack |
| The maintenance calendar | Windows already scheduled, for the next 30 days |
| The business impact register | Business criticality per system |
Physical placement fields have no automatic source: a person puts a machine in a cabinet, and nothing on the network knows which room it is in. They arrive from a rack-list import or from manual entry on the host record. Where they are missing, the screen says they are missing rather than filling in a guess.
The tiles at the top
| Tile | What it counts |
|---|---|
| Runway | Days until the nearest threshold crossing across the estate, on the observed trend |
| Clusters at risk after a failure | Clusters whose memory utilization after a host failure crosses the warning threshold |
| Evacuations blocked in coming windows | Maintenance windows on a hypervisor host where the rest of the cluster cannot absorb its guests |
| Racks over the threshold | Racks crossing the threshold on either U utilization or circuit utilization |
| Reclaimable capacity | The number of distinct machines that can be given back without buying anything |
| Total power | The sum of average draw across physical servers, with the derived cooling tonnage beside it |
A tile whose value is zero renders a healthy state rather than disappearing. A metric that vanishes when it is healthy teaches nobody that it exists, and on the day it is not zero the reader has no baseline for it.
What binds, and when
For each cluster the table ranks the resource closest to its own threshold, out of four candidates: memory after a host failure, committed memory, the rack's electrical circuit and rack units. The last two come from the cabinets the cluster's hypervisor hosts physically sit in — a cluster at a third of its memory whose hosts sit in a cabinet that has finished its circuit does not have a memory problem, and a screen that leads with the memory bar sends someone to buy the wrong thing.
Two resources are reported and deliberately kept out of the ranking, each stating why:
| Resource | Why it is not ranked |
|---|---|
| Provisioned disk | Reported as growth only — there is no array capacity to divide it by, so it has no percentage |
| vCPU:pCPU ratio | A planning target the organization sets for itself, not a threshold |
The When column does not always aim at the same threshold. A resource that has already passed its warning edge is asked about its critical edge instead: "you crossed it in June" is a finding and "you cross it in November" is a plan, and answering the first with 0 days buries the second.
Capacity alerts
The exception list is one row per threshold crossed, critical before warning. Every row carries the affected systems and their criticality, not just a percentage: a cluster running out of memory under a payments system and a cluster running out of memory under a DR environment are the same arithmetic and completely different sentences. A rack breach names the systems in that cabinet too.
An empty list is not an empty screen: it prints the thresholds in force at that moment, so an empty screen reads as healthy rather than broken.
The same breaches, and an infeasible evacuation in a scheduled window, also reach the notification bell. They are visible only to someone holding the capacity view permission.
Thresholds
Thresholds are editable in settings, and each row on the settings screen prints what its number rests on and grades that source — standard for a binding source, convention for accepted practice with no source behind it.
| Threshold | What it rests on |
|---|---|
| Memory utilization after failover | The M/M/1 load knee. The mechanism is documented; the round numbers themselves are in no standard |
| Memory utilization | The same load knee, with no binding source for the numbers |
| Rack utilization | Disputed — one school plans to 80–85% for airflow and growth, another rejects any blanket percentage |
| Circuit utilization | NFPA 70 §210.20(A) — a continuous load is planned at 80% of the circuit rating |
| Runway | Inverted: fewer days is worse. Rests on the defaults capacity vendors publish |
| Provisioned-disk growth | No default is shipped |
| vCPU:pCPU ratio | Deliberately has no threshold |
The last two rows arrive with no number, and that is not an omission. A settings screen that offers a default for everything teaches its reader that every default is equally well founded.
Trends
Trends and forecasts rest on a daily measurement series written by a nightly job — one point per day, per scope, per metric. It cannot be backfilled: an inventory read gives today and only today, so the series begins on the night the job is first scheduled.
Every screen header in the module shows History: N days, and History: none yet when there is nothing. Instead of four empty charts, the screen renders one panel stating how many points have been collected and how many are needed. The forecast rules themselves are described under Capacity planning.
Permissions
The module has three permissions: view, plan and export. The last two imply the first, so a holder of the planning permission sees the other screens too. Editing thresholds is not a fourth key — it rests on the system settings permission, because the thresholds tab is one more card on a screen that already exists.
None of the three is part of an operator's default permissions; they are granted through a role. Without the view permission the module's rows do not appear in the sidebar, and any read of capacity data is refused at the server rather than only on screen.
The screens
| Screen | The question it answers |
|---|---|
| Compute clusters | If a host dies tonight, does anything stop? |
| Racks and sites | Where is there physical room, and for what |
| Reclaimable capacity | What can be given back without buying anything |
| Capacity planning | Is there room for this addition? |
Updated
This page is the file content/docs/en/v1/capacity/overview.mdx