Get AI serving deployment metrics
Returns latency quantile, capacity and throughput time-series per model for one serving deployment, sourced from the metrics store. Each series states its own unit because gateways disagree: vLLM publishes seconds and APISIX milliseconds, and a histogram cannot be rescaled after recording. Signals the deployment’s gateway does not publish are named in missing_signals rather than returned as zero. A range of 24h or less is answered from the per-replica series; a longer range is answered from the deployment-grain series, which carries no replica identity, so a replica filter beyond 24h is rejected. The response also states how the deployment’s pods map to accelerators, and never divides a card’s figure between pods. cost_per_mtok is the accelerator spend behind a million output tokens, at each point of the window; it is returned only where the deployment’s accelerator rate is known and the deployment is serving output tokens, and is named in missing_signals otherwise. gateway_connection says why a deployment has no series, since the three causes need three different actions: not_connected when no model server is connected for this namespace, selector_mismatch when one is but this deployment’s pods do not carry the labels it is read by, naming that selector, and connected when the deployment is being read and any silence is the gateway’s own. missing_signal_notes explains a missing signal the connected server only started publishing in a later release, so a request-stage breakdown an older build never emits reads as a version gap rather than as a broken scrape.
Authorizations
Enter "Bearer {token}"
Headers
Tenant ID
Path Parameters
AI serving deployment ID
Query Parameters
Time window (e.g. 1h, 6h, 24h, 7d). Default 1h
Resolution step (e.g. 30s, 1m, 5m)
Serving signals (comma-separated): all,time_to_first_token,inter_token,end_to_end,queue,prefill,decode,request_output_tokens,kv_cache_utilization,batch_size,queue_depth,preemptions,concurrency,tokens_per_second,cost_per_mtok. Default all
Quantiles between 0 and 1 (comma-separated). Default 0.5,0.95,0.99
Restrict to one model
Restrict to one replica, as namespace/pod