Get accelerator metrics
Returns time-series metrics for one accelerator held by the tenant selected by the X-CostGraph-Tenant-ID header, read from the metrics store rather than from the stored reports the detail endpoint returns. Pass mig_instance_id to read one partition of a carved-up card instead of the whole card. gpu_availability states each dimension separately: supported where the card measured it, capacity_unknown where the card measured it but reports no capacity to express it as a percentage, unsupported where the card cannot report it at all, and no_data where nothing was measured; none of the four is ever returned as zero. gpu_memory_used_mib and gpu_power_watts carry the raw measurement behind the memory and power ratios, in mebibytes and watts, whenever the card reported it, so a capacity_unknown dimension still arrives with its numbers rather than as an empty series. gpu_series carries every other signal the card reports as its own time series, keyed by the same names the stored reports use: gpu_temperature, gpu_memory_temperature, gpu_sm_clock, gpu_memory_clock, gpu_encoder_utilization, gpu_decoder_utilization, gpu_memory_copy_utilization, gpu_sm_active, gpu_sm_occupancy, gpu_pcie_tx_bytes, gpu_pcie_rx_bytes, gpu_energy_joules, gpu_xid_errors, gpu_remapped_rows_correctable, gpu_remapped_rows_uncorrectable and gpu_row_remap_failure. A signal the card does not report is left out of gpu_series and named unsupported in gpu_availability, which now states all of them; a card without a memory temperature sensor reports a literal 0, and that is dropped rather than charted as freezing. gpu_units gives the unit of every series, including gpu_memory_used_mib and gpu_power_watts, so a caller never has to infer one: percent, celsius, megahertz, bytes, joules, count, mebibytes or watts. The fault counters and energy are reported as what accrued in the five minutes before each point rather than as the driver’s running total, which only ever climbs. gpu_identity_source names the label the series were matched on, since not every card carries a costgraph instance id. A uuid the tenant has never reported gets a 404, as does a mig_instance_id the card has never reported.
Authorizations
Enter "Bearer {token}"
Headers
Tenant ID
Path Parameters
Accelerator UUID as the device reports it
Query Parameters
Time window (e.g. 1h, 6h, 24h, 7d). Default 1h
Resolution step (e.g. 30s, 1m, 5m)
Metric groups to fetch (comma-separated): all,gpu. Default gpu
Restrict to one MIG partition of the card