> ## Documentation Index
> Fetch the complete documentation index at: https://docs.costgraph.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# GPUs

> See every GPU you pay for, how much of it is used, and which cards sit idle

GPUs are often the most expensive line on a compute bill, and the easiest to
waste. A card reserved for a training job that finished last week costs the
same as one running flat out. The **GPUs** page in CostGraph shows every NVIDIA
GPU across your Kubernetes clusters and virtual machines. It shows where each
card runs, how much of it is used, and what it costs.

## What you get

* **GPU inventory**: every card in one list, with its model, where it runs, and
  when it last reported. Cards on Kubernetes nodes and on virtual machines
  appear side by side.
* **GPU spend**: estimated month-to-date cost for your GPUs, with a projected
  monthly figure per card.
* **Idle GPUs**: a count of cards and MIG partitions that are doing no work, so
  you can find capacity to release.
* **Utilization verdict**: a plain label for each card, from critically
  under-used to saturated, with the reasons behind it.
* **Usage history**: each measurement summarized over the latest reading, the
  past week, and the past month.
* **Detailed metrics**: charts for utilization, memory, power and energy,
  thermals and clocks, engines, interconnect, and hardware faults.
* **MIG partitions**: for a partitioned card, each partition with its own
  utilization and whether it is idle.
* **Workload links**: on Kubernetes, the node and cluster each card belongs to,
  and the model server holding it when CostGraph can identify one.
* **Container GPU usage**: on a container's usage page, GPU measurements sit
  next to CPU and memory.

## Find your GPUs

Open **GPUs** in the sidebar to see your whole fleet. The same list, narrowed to
one place, appears on the **GPUs** tab of a cluster and of a virtual machine.

Filter the list by:

* **Provider**
* **Model**
* **Utilization** verdict
* **Runs on**: Kubernetes or a virtual machine
* **Reported within**: hide cards that have stopped reporting
* **Partitioning**: MIG-enabled cards or whole cards

Select a card to open it. The card page has three tabs: **Overview**,
**Metrics**, and **Usage**.

## Read the utilization verdict

Each card gets one verdict, based on how much of it your workloads use over
time.

| Verdict | What it means |
| - | - |
| Critically under-used | The card is close to idle. You pay for nearly all of it and use very little. |
| Under-used | The card does some work, but far less than it can. |
| Healthy | The card is used at a level that justifies its cost. |
| Over-used | The card runs hot for long stretches. Work may be queuing behind it. |
| Saturated | The card is at its limit. More capacity, or less work, is likely needed. |
| Not measured | CostGraph hasn't received enough data to judge the card. |

The **Verdict and evidence** panel on the card's overview explains the label.
It lists the reasons behind the verdict and the signals it rests on. It also
shows how complete the data is:

* **Complete**: every signal was present for the whole period.
* **Partial signal**: some of the period has no data, so the figures cover less
  time than stated.
* **Stale**: nothing has reported this card recently.
* **No signal**: the signals the verdict needs haven't been collected.

Treat a verdict with partial or stale data with care. Fix the data gap before
you act on it.

## Act on idle and under-used GPUs

An idle or under-used card is a candidate to release, share, or replace with a
smaller one:

* **Idle card**: if nothing needs it, remove it from the node pool or
  virtual machine.
* **Under-used card**: move the work onto fewer cards, or onto a smaller GPU
  model.
* **Idle MIG partition**: change the partition layout, or schedule work onto
  the free partitions before you add cards.
* **Over-used or saturated card**: check whether jobs are waiting on it before
  you remove capacity elsewhere.

On Kubernetes, the container view shows the GPUs each container holds. GPUs are
requested as whole cards, so the only sizing change is the number of cards a
container asks for. CostGraph shows whether a container has its GPU to itself
or shares it. When a card is shared, or sharing is unknown, CostGraph makes no
sizing recommendation for it. The usage it sees isn't that container's alone.

## How GPU cost is estimated

GPU cost is an estimate from public list prices, not your invoice:

* **Month to date** covers the hours since the start of the month that the card
  has been running.
* **Projected month** is what a full month at the current rate comes to.

Where a provider prices the GPU separately, CostGraph uses that price. Where it
doesn't, the card's cost is its share of the machine it is attached to. A card
CostGraph has no price for shows no cost instead of a guess. The fleet's
**GPU spend** total covers only the cards that have a price.

## Supported environments

* Kubernetes clusters running the [CostGraph operator](/costgraph/operator/overview)
* Linux virtual machines and bare-metal hosts running the
  [CostGraph agent](/costgraph/agent)
* NVIDIA GPUs, including MIG-partitioned cards

Inventory, utilization, and verdicts work on any provider. Separate GPU list
prices are available for:

* Google Cloud
* STACKIT

On other providers, a card's cost comes from the price of the machine it is
attached to, where CostGraph has one.

## Set up GPU collection

CostGraph reads GPU measurements from NVIDIA's DCGM exporter. Every GPU host
needs the NVIDIA driver installed and the DCGM exporter running.

### Kubernetes

The operator chart can deploy the DCGM exporter for you. It is off by default.
Turn it on, and turn on the matching scrape target, in your operator values:

```yaml theme={null}
dcgm-exporter:
  enabled: true
  extraEnv:
    - name: DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE
      value: device-name
operatorPrometheus:
  scrapeTargets:
    dcgm-exporter:
      enabled: true
```

Set `DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE` to match your cluster:

* `device-name` on GKE.
* `uuid` on clusters that use the standard NVIDIA device plugin, such as EKS.

With the wrong value, GPUs still appear but aren't linked to the pods using
them.

Apply the change with `helm upgrade`. The exporter runs only on GPU nodes and
tolerates the usual `nvidia.com/gpu` taint. For the full set of options,
including the optional profiling exporter that adds engine-level measurements,
see [Configuration](/costgraph/operator/configuration#bundled-exporters) and
[Accelerators](/costgraph/operator/accelerators).

### Virtual machines and bare metal

Install the [CostGraph agent](/costgraph/agent#installation) with the install
script. On a host where `nvidia-smi` is available, the script also installs and
starts the DCGM exporter, and the agent reads from it automatically. The host
must use systemd for the script to set up the exporter.

If the agent is already installed, run the install script again on the GPU
host to add the exporter.

Already run a DCGM exporter? Point the installer at it instead:

```bash theme={null}
curl -sSL https://setup.costgraph.ai/install.sh | \
  COSTGRAPH_API_KEY="bl_..." \
  COSTGRAPH_GPU_EXPORTER_ENDPOINT="http://<exporter-host>:<port>/metrics" sh
```

On a host where the agent runs in Docker, or without systemd, run the DCGM
exporter yourself. Then set `COSTGRAPH_DCGM_ENDPOINT` on the agent to its
metrics URL.

To skip the exporter on a GPU host, set `COSTGRAPH_INSTALL_GPU_EXPORTER=0` when
you run the installer.

## Fix missing GPU data

<AccordionGroup>
  <Accordion title="A GPU host shows no GPUs">
    Check that the NVIDIA driver is installed: `nvidia-smi` must list the card.
    Then check that the DCGM exporter is running on that host or node. On
    Kubernetes, confirm `dcgm-exporter.enabled` and its scrape target are both
    `true`, and that an exporter pod is running on the GPU node. On a virtual
    machine, check the exporter service with `systemctl status
            nvidia-dcgm-exporter`.
  </Accordion>

  <Accordion title="A measurement says No data reached CostGraph">
    The card is known, but no measurements arrived for it. The DCGM exporter is
    usually not running or not scheduled on that node. Start it, and the
    measurements appear on the next collection.
  </Accordion>

  <Accordion title="A measurement says Not reported by this card">
    The GPU model or its driver doesn't publish that measurement. No setting
    changes this. The other measurements for the card are unaffected.
  </Accordion>

  <Accordion title="GPUs appear but aren't linked to pods">
    `DCGM_EXPORTER_KUBERNETES_GPU_ID_TYPE` doesn't match your cluster's device
    plugin. Use `device-name` on GKE and `uuid` elsewhere, then upgrade the
    release.
  </Accordion>

  <Accordion title="The verdict says Not measured or Stale">
    New cards need some history before CostGraph can judge them, so verdicts
    appear later than live metrics. A stale verdict means the card stopped
    reporting. Check that the host or node is running and the exporter is up.
  </Accordion>

  <Accordion title="A GPU shows no cost">
    CostGraph has no list price for that GPU model on that provider and machine.
    Usage and verdicts still work. The card is left out of the GPU spend total.
  </Accordion>
</AccordionGroup>

## Next steps

<CardGroup cols={2}>
  <Card title="Operator configuration" icon="gear" href="/costgraph/operator/configuration#bundled-exporters">
    Every option for the DCGM exporter on Kubernetes.
  </Card>

  <Card title="Agent" icon="microchip" href="/costgraph/agent">
    Install and configure the agent on virtual machines.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.