> ## Documentation Index
> Fetch the complete documentation index at: https://docs.costgraph.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get accelerator metrics

> Returns time-series metrics for one accelerator held by the tenant selected by the X-CostGraph-Tenant-ID header, read from the metrics store rather than from the stored reports the detail endpoint returns. Pass mig_instance_id to read one partition of a carved-up card instead of the whole card. gpu_availability states each dimension separately: supported where the card measured it, capacity_unknown where the card measured it but reports no capacity to express it as a percentage, unsupported where the card cannot report it at all, and no_data where nothing was measured; none of the four is ever returned as zero. gpu_memory_used_mib and gpu_power_watts carry the raw measurement behind the memory and power ratios, in mebibytes and watts, whenever the card reported it, so a capacity_unknown dimension still arrives with its numbers rather than as an empty series. gpu_series carries every other signal the card reports as its own time series, keyed by the same names the stored reports use: gpu_temperature, gpu_memory_temperature, gpu_sm_clock, gpu_memory_clock, gpu_encoder_utilization, gpu_decoder_utilization, gpu_memory_copy_utilization, gpu_sm_active, gpu_sm_occupancy, gpu_pcie_tx_bytes, gpu_pcie_rx_bytes, gpu_energy_joules, gpu_xid_errors, gpu_remapped_rows_correctable, gpu_remapped_rows_uncorrectable and gpu_row_remap_failure. A signal the card does not report is left out of gpu_series and named unsupported in gpu_availability, which now states all of them; a card without a memory temperature sensor reports a literal 0, and that is dropped rather than charted as freezing. gpu_units gives the unit of every series, including gpu_memory_used_mib and gpu_power_watts, so a caller never has to infer one: percent, celsius, megahertz, bytes, joules, count, mebibytes or watts. The fault counters and energy are reported as what accrued in the five minutes before each point rather than as the driver's running total, which only ever climbs. gpu_identity_source names the label the series were matched on, since not every card carries a costgraph instance id. A uuid the tenant has never reported gets a 404, as does a mig_instance_id the card has never reported.



## OpenAPI

````yaml /api-reference/costgraph/openapi.json get /api/v1/tenant/gpus/{gpu_uuid}/metrics
openapi: 3.0.0
info:
  description: Read and manage your CostGraph organization, spend, alerts, and settings.
  title: CostGraph API
  contact: {}
  version: '1.0'
servers:
  - url: https://api.costgraph.ai
security: []
tags:
  - name: ai
    x-group: AI
  - name: ai-serving
    x-group: AI serving
  - name: anomalies
    x-group: Anomalies
  - name: auth
    x-group: Auth
  - name: billing
    x-group: Billing
  - name: billing-export
    x-group: Billing export
  - name: budgets
    x-group: Budgets
  - name: ci
    x-group: CI
  - name: compute-recommendations
    x-group: Compute recommendations
  - name: config
    x-group: Config
  - name: cost
    x-group: Cost
  - name: gpus
    x-group: GPUs
  - name: graphai
    x-group: Graph AI
  - name: infracost
    x-group: Infracost
  - name: integrations
    x-group: Integrations
  - name: invitations
    x-group: Invitations
  - name: kubernetes-clusters
    x-group: Kubernetes clusters
  - name: marketplace
    x-group: Marketplace
  - name: network-requests
    x-group: Network requests
  - name: notifications
    x-group: Notifications
  - name: oauth
    x-group: OAuth
  - name: oauth-clients
    x-group: OAuth clients
  - name: opencost
    x-group: OpenCost
  - name: organization
    x-group: Audit log
  - name: organizations
    x-group: Organizations
  - name: placement-alternatives
    x-group: Placement alternatives
  - name: reports
    x-group: Reports
  - name: service-map
    x-group: Service map
  - name: settings
    x-group: Settings
  - name: sso
    x-group: Single sign-on
  - name: tenants
    x-group: Tenants
  - name: user
    x-group: Users
  - name: virtual-machines
    x-group: Virtual machines
  - name: virtual-tags
    x-group: Virtual tags
  - name: workflows
    x-group: Workflows
paths:
  /api/v1/tenant/gpus/{gpu_uuid}/metrics:
    get:
      tags:
        - gpus
      summary: Get accelerator metrics
      description: >-
        Returns time-series metrics for one accelerator held by the tenant
        selected by the X-CostGraph-Tenant-ID header, read from the metrics
        store rather than from the stored reports the detail endpoint returns.
        Pass mig_instance_id to read one partition of a carved-up card instead
        of the whole card. gpu_availability states each dimension separately:
        supported where the card measured it, capacity_unknown where the card
        measured it but reports no capacity to express it as a percentage,
        unsupported where the card cannot report it at all, and no_data where
        nothing was measured; none of the four is ever returned as zero.
        gpu_memory_used_mib and gpu_power_watts carry the raw measurement behind
        the memory and power ratios, in mebibytes and watts, whenever the card
        reported it, so a capacity_unknown dimension still arrives with its
        numbers rather than as an empty series. gpu_series carries every other
        signal the card reports as its own time series, keyed by the same names
        the stored reports use: gpu_temperature, gpu_memory_temperature,
        gpu_sm_clock, gpu_memory_clock, gpu_encoder_utilization,
        gpu_decoder_utilization, gpu_memory_copy_utilization, gpu_sm_active,
        gpu_sm_occupancy, gpu_pcie_tx_bytes, gpu_pcie_rx_bytes,
        gpu_energy_joules, gpu_xid_errors, gpu_remapped_rows_correctable,
        gpu_remapped_rows_uncorrectable and gpu_row_remap_failure. A signal the
        card does not report is left out of gpu_series and named unsupported in
        gpu_availability, which now states all of them; a card without a memory
        temperature sensor reports a literal 0, and that is dropped rather than
        charted as freezing. gpu_units gives the unit of every series, including
        gpu_memory_used_mib and gpu_power_watts, so a caller never has to infer
        one: percent, celsius, megahertz, bytes, joules, count, mebibytes or
        watts. The fault counters and energy are reported as what accrued in the
        five minutes before each point rather than as the driver's running
        total, which only ever climbs. gpu_identity_source names the label the
        series were matched on, since not every card carries a costgraph
        instance id. A uuid the tenant has never reported gets a 404, as does a
        mig_instance_id the card has never reported.
      parameters:
        - description: Tenant ID
          name: X-CostGraph-Tenant-ID
          in: header
          required: true
          schema:
            type: string
        - description: Accelerator UUID as the device reports it
          name: gpu_uuid
          in: path
          required: true
          schema:
            type: string
        - description: Time window (e.g. 1h, 6h, 24h, 7d). Default 1h
          name: range
          in: query
          schema:
            type: string
        - description: Resolution step (e.g. 30s, 1m, 5m)
          name: step
          in: query
          schema:
            type: string
        - description: 'Metric groups to fetch (comma-separated): all,gpu. Default gpu'
          name: groups
          in: query
          schema:
            type: string
        - description: Restrict to one MIG partition of the card
          name: mig_instance_id
          in: query
          schema:
            type: string
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                allOf:
                  - $ref: '#/components/schemas/responses.SuccessResponse'
                  - type: object
                    properties:
                      data:
                        $ref: >-
                          #/components/schemas/costgraph_agent.VirtualMachineMetrics
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '403':
          description: Forbidden
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '404':
          description: Not Found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
      security:
        - BearerAuth: []
components:
  schemas:
    responses.SuccessResponse:
      type: object
      required:
        - message
        - status
      properties:
        data: {}
        message:
          type: string
          example: some message
        status:
          type: string
          example: success
    costgraph_agent.VirtualMachineMetrics:
      type: object
      required:
        - end
        - start
        - step_seconds
      properties:
        cpu_irq:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_nice:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_psi_some:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_softirq:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_steal:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_system:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_user:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        cpu_wait:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_busy_percent:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_read_iops:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_read_latency:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_read_throughput:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_write_iops:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_write_latency:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        disk_write_throughput:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        end:
          type: string
        gpu_availability:
          type: object
          additionalProperties:
            type: string
        gpu_identity_source:
          type: string
        gpu_memory_used_mib:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        gpu_memory_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        gpu_power_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        gpu_power_watts:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        gpu_series:
          type: object
          additionalProperties:
            type: array
            items:
              $ref: '#/components/schemas/metrics.MetricPoint'
        gpu_units:
          type: object
          additionalProperties:
            type: string
        gpu_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        io_psi_some:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        load_average_15m:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        load_average_1m:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        load_average_5m:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_available:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_buffers:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_buffers_percent:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_cached:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_cached_percent:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_free:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_psi_some:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_used:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        memory_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        network_receive:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        network_transmit:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        start:
          type: string
        step_seconds:
          type: integer
        swap_utilization:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        tcp_retransmits:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
    responses.ErrorResponse:
      type: object
      required:
        - message
        - status
      properties:
        message:
          type: string
          example: some message
        status:
          type: string
          example: error
    metrics.MetricPoint:
      type: object
      required:
        - t
        - v
      properties:
        t:
          type: integer
        v:
          type: number
  securitySchemes:
    BearerAuth:
      description: Enter "Bearer {token}"
      type: apiKey
      name: Authorization
      in: header

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.