Skip to main content
GPUs are often the most expensive line on a compute bill, and the easiest to waste. A card reserved for a training job that finished last week costs the same as one running flat out. The GPUs page in CostGraph shows every NVIDIA GPU across your Kubernetes clusters and virtual machines. It shows where each card runs, how much of it is used, and what it costs.

What you get

  • GPU inventory: every card in one list, with its model, where it runs, and when it last reported. Cards on Kubernetes nodes and on virtual machines appear side by side.
  • GPU spend: estimated month-to-date cost for your GPUs, with a projected monthly figure per card.
  • Idle GPUs: a count of cards and MIG partitions that are doing no work, so you can find capacity to release.
  • Utilization verdict: a plain label for each card, from critically under-used to saturated, with the reasons behind it.
  • Usage history: each measurement summarized over the latest reading, the past week, and the past month.
  • Detailed metrics: charts for utilization, memory, power and energy, thermals and clocks, engines, interconnect, and hardware faults.
  • MIG partitions: for a partitioned card, each partition with its own utilization and whether it is idle.
  • Workload links: on Kubernetes, the node and cluster each card belongs to, and the model server holding it when CostGraph can identify one.
  • Container GPU usage: on a container’s usage page, GPU measurements sit next to CPU and memory.

Find your GPUs

Open GPUs in the sidebar to see your whole fleet. The same list, narrowed to one place, appears on the GPUs tab of a cluster and of a virtual machine. Filter the list by:
  • Provider
  • Model
  • Utilization verdict
  • Runs on: Kubernetes or a virtual machine
  • Reported within: hide cards that have stopped reporting
  • Partitioning: MIG-enabled cards or whole cards
Select a card to open it. The card page has three tabs: Overview, Metrics, and Usage.

Read the utilization verdict

Each card gets one verdict, based on how much of it your workloads use over time. The Verdict and evidence panel on the card’s overview explains the label. It lists the reasons behind the verdict and the signals it rests on. It also shows how complete the data is:
  • Complete: every signal was present for the whole period.
  • Partial signal: some of the period has no data, so the figures cover less time than stated.
  • Stale: nothing has reported this card recently.
  • No signal: the signals the verdict needs haven’t been collected.
Treat a verdict with partial or stale data with care. Fix the data gap before you act on it.

Act on idle and under-used GPUs

An idle or under-used card is a candidate to release, share, or replace with a smaller one:
  • Idle card: if nothing needs it, remove it from the node pool or virtual machine.
  • Under-used card: move the work onto fewer cards, or onto a smaller GPU model.
  • Idle MIG partition: change the partition layout, or schedule work onto the free partitions before you add cards.
  • Over-used or saturated card: check whether jobs are waiting on it before you remove capacity elsewhere.
On Kubernetes, the container view shows the GPUs each container holds. GPUs are requested as whole cards, so the only sizing change is the number of cards a container asks for. CostGraph shows whether a container has its GPU to itself or shares it. When a card is shared, or sharing is unknown, CostGraph makes no sizing recommendation for it. The usage it sees isn’t that container’s alone.

How GPU cost is estimated

GPU cost is an estimate from public list prices, not your invoice:
  • Month to date covers the hours since the start of the month that the card has been running.
  • Projected month is what a full month at the current rate comes to.
Where a provider prices the GPU separately, CostGraph uses that price. Where it doesn’t, the card’s cost is its share of the machine it is attached to. A card CostGraph has no price for shows no cost instead of a guess. The fleet’s GPU spend total covers only the cards that have a price.

Supported environments

  • Kubernetes clusters running the CostGraph operator
  • Linux virtual machines and bare-metal hosts running the CostGraph agent
  • NVIDIA GPUs, including MIG-partitioned cards
Inventory, utilization, and verdicts work on any provider. Separate GPU list prices are available for:
  • Google Cloud
  • STACKIT
On other providers, a card’s cost comes from the price of the machine it is attached to, where CostGraph has one.

Set up GPU collection

CostGraph reads GPU measurements from NVIDIA’s DCGM exporter. Every GPU host needs the NVIDIA driver installed and the DCGM exporter running.

Kubernetes

The operator chart can deploy the DCGM exporter for you. It is off by default. Turn it on, and turn on the matching scrape target, in your operator values:
Set DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE to match your cluster:
  • device-name on GKE.
  • uuid on clusters that use the standard NVIDIA device plugin, such as EKS.
With the wrong value, GPUs still appear but aren’t linked to the pods using them. Apply the change with helm upgrade. The exporter runs only on GPU nodes and tolerates the usual nvidia.com/gpu taint. For the full set of options, including the optional profiling exporter that adds engine-level measurements, see Configuration and Accelerators.

Virtual machines and bare metal

Install the CostGraph agent with the install script. On a host where nvidia-smi is available, the script also installs and starts the DCGM exporter, and the agent reads from it automatically. The host must use systemd for the script to set up the exporter. If the agent is already installed, run the install script again on the GPU host to add the exporter. Already run a DCGM exporter? Point the installer at it instead:
On a host where the agent runs in Docker, or without systemd, run the DCGM exporter yourself. Then set COSTGRAPH_DCGM_ENDPOINT on the agent to its metrics URL. To skip the exporter on a GPU host, set COSTGRAPH_INSTALL_GPU_EXPORTER=0 when you run the installer.

Fix missing GPU data

Check that the NVIDIA driver is installed: nvidia-smi must list the card. Then check that the DCGM exporter is running on that host or node. On Kubernetes, confirm dcgm-exporter.enabled and its scrape target are both true, and that an exporter pod is running on the GPU node. On a virtual machine, check the exporter service with systemctl status nvidia-dcgm-exporter.
The card is known, but no measurements arrived for it. The DCGM exporter is usually not running or not scheduled on that node. Start it, and the measurements appear on the next collection.
The GPU model or its driver doesn’t publish that measurement. No setting changes this. The other measurements for the card are unaffected.
DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE doesn’t match your cluster’s device plugin. Use device-name on GKE and uuid elsewhere, then upgrade the release.
New cards need some history before CostGraph can judge them, so verdicts appear later than live metrics. A stale verdict means the card stopped reporting. Check that the host or node is running and the exporter is up.
CostGraph has no list price for that GPU model on that provider and machine. Usage and verdicts still work. The card is left out of the GPU spend total.

Next steps

Operator configuration

Every option for the DCGM exporter on Kubernetes.

Agent

Install and configure the agent on virtual machines.