What you get
- GPU inventory: every card in one list, with its model, where it runs, and when it last reported. Cards on Kubernetes nodes and on virtual machines appear side by side.
- GPU spend: estimated month-to-date cost for your GPUs, with a projected monthly figure per card.
- Idle GPUs: a count of cards and MIG partitions that are doing no work, so you can find capacity to release.
- Utilization verdict: a plain label for each card, from critically under-used to saturated, with the reasons behind it.
- Usage history: each measurement summarized over the latest reading, the past week, and the past month.
- Detailed metrics: charts for utilization, memory, power and energy, thermals and clocks, engines, interconnect, and hardware faults.
- MIG partitions: for a partitioned card, each partition with its own utilization and whether it is idle.
- Workload links: on Kubernetes, the node and cluster each card belongs to, and the model server holding it when CostGraph can identify one.
- Container GPU usage: on a container’s usage page, GPU measurements sit next to CPU and memory.
Find your GPUs
Open GPUs in the sidebar to see your whole fleet. The same list, narrowed to one place, appears on the GPUs tab of a cluster and of a virtual machine. Filter the list by:- Provider
- Model
- Utilization verdict
- Runs on: Kubernetes or a virtual machine
- Reported within: hide cards that have stopped reporting
- Partitioning: MIG-enabled cards or whole cards
Read the utilization verdict
Each card gets one verdict, based on how much of it your workloads use over time.
The Verdict and evidence panel on the card’s overview explains the label.
It lists the reasons behind the verdict and the signals it rests on. It also
shows how complete the data is:
- Complete: every signal was present for the whole period.
- Partial signal: some of the period has no data, so the figures cover less time than stated.
- Stale: nothing has reported this card recently.
- No signal: the signals the verdict needs haven’t been collected.
Act on idle and under-used GPUs
An idle or under-used card is a candidate to release, share, or replace with a smaller one:- Idle card: if nothing needs it, remove it from the node pool or virtual machine.
- Under-used card: move the work onto fewer cards, or onto a smaller GPU model.
- Idle MIG partition: change the partition layout, or schedule work onto the free partitions before you add cards.
- Over-used or saturated card: check whether jobs are waiting on it before you remove capacity elsewhere.
How GPU cost is estimated
GPU cost is an estimate from public list prices, not your invoice:- Month to date covers the hours since the start of the month that the card has been running.
- Projected month is what a full month at the current rate comes to.
Supported environments
- Kubernetes clusters running the CostGraph operator
- Linux virtual machines and bare-metal hosts running the CostGraph agent
- NVIDIA GPUs, including MIG-partitioned cards
- Google Cloud
- STACKIT
Set up GPU collection
CostGraph reads GPU measurements from NVIDIA’s DCGM exporter. Every GPU host needs the NVIDIA driver installed and the DCGM exporter running.Kubernetes
The operator chart can deploy the DCGM exporter for you. It is off by default. Turn it on, and turn on the matching scrape target, in your operator values:DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE to match your cluster:
device-nameon GKE.uuidon clusters that use the standard NVIDIA device plugin, such as EKS.
helm upgrade. The exporter runs only on GPU nodes and
tolerates the usual nvidia.com/gpu taint. For the full set of options,
including the optional profiling exporter that adds engine-level measurements,
see Configuration and
Accelerators.
Virtual machines and bare metal
Install the CostGraph agent with the install script. On a host wherenvidia-smi is available, the script also installs and
starts the DCGM exporter, and the agent reads from it automatically. The host
must use systemd for the script to set up the exporter.
If the agent is already installed, run the install script again on the GPU
host to add the exporter.
Already run a DCGM exporter? Point the installer at it instead:
COSTGRAPH_DCGM_ENDPOINT on the agent to its
metrics URL.
To skip the exporter on a GPU host, set COSTGRAPH_INSTALL_GPU_EXPORTER=0 when
you run the installer.
Fix missing GPU data
A GPU host shows no GPUs
A GPU host shows no GPUs
Check that the NVIDIA driver is installed:
nvidia-smi must list the card.
Then check that the DCGM exporter is running on that host or node. On
Kubernetes, confirm dcgm-exporter.enabled and its scrape target are both
true, and that an exporter pod is running on the GPU node. On a virtual
machine, check the exporter service with systemctl status nvidia-dcgm-exporter.A measurement says No data reached CostGraph
A measurement says No data reached CostGraph
The card is known, but no measurements arrived for it. The DCGM exporter is
usually not running or not scheduled on that node. Start it, and the
measurements appear on the next collection.
A measurement says Not reported by this card
A measurement says Not reported by this card
The GPU model or its driver doesn’t publish that measurement. No setting
changes this. The other measurements for the card are unaffected.
GPUs appear but aren't linked to pods
GPUs appear but aren't linked to pods
DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE doesn’t match your cluster’s device
plugin. Use device-name on GKE and uuid elsewhere, then upgrade the
release.The verdict says Not measured or Stale
The verdict says Not measured or Stale
New cards need some history before CostGraph can judge them, so verdicts
appear later than live metrics. A stale verdict means the card stopped
reporting. Check that the host or node is running and the exporter is up.
A GPU shows no cost
A GPU shows no cost
CostGraph has no list price for that GPU model on that provider and machine.
Usage and verdicts still work. The card is left out of the GPU spend total.
Next steps
Operator configuration
Every option for the DCGM exporter on Kubernetes.
Agent
Install and configure the agent on virtual machines.