> ## Documentation Index
> Fetch the complete documentation index at: https://docs.costgraph.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get AI serving deployment metrics

> Returns latency quantile, capacity and throughput time-series per model for one serving deployment, sourced from the metrics store. Each series states its own unit because gateways disagree: vLLM publishes seconds and APISIX milliseconds, and a histogram cannot be rescaled after recording. Signals the deployment's gateway does not publish are named in missing_signals rather than returned as zero. A range of 24h or less is answered from the per-replica series; a longer range is answered from the deployment-grain series, which carries no replica identity, so a replica filter beyond 24h is rejected. The response also states how the deployment's pods map to accelerators, and never divides a card's figure between pods. cost_per_mtok is the accelerator spend behind a million output tokens, at each point of the window; it is returned only where the deployment's accelerator rate is known and the deployment is serving output tokens, and is named in missing_signals otherwise. gateway_connection says why a deployment has no series, since the three causes need three different actions: not_connected when no model server is connected for this namespace, selector_mismatch when one is but this deployment's pods do not carry the labels it is read by, naming that selector, and connected when the deployment is being read and any silence is the gateway's own. missing_signal_notes explains a missing signal the connected server only started publishing in a later release, so a request-stage breakdown an older build never emits reads as a version gap rather than as a broken scrape.



## OpenAPI

````yaml /api-reference/costgraph/openapi.json get /api/v1/tenant/ai/deployments/{id}/metrics
openapi: 3.0.0
info:
  description: Read and manage your CostGraph organization, spend, alerts, and settings.
  title: CostGraph API
  contact: {}
  version: '1.0'
servers:
  - url: https://api.costgraph.ai
security: []
tags:
  - name: ai
    x-group: AI
  - name: ai-serving
    x-group: AI serving
  - name: anomalies
    x-group: Anomalies
  - name: auth
    x-group: Auth
  - name: billing
    x-group: Billing
  - name: billing-export
    x-group: Billing export
  - name: budgets
    x-group: Budgets
  - name: ci
    x-group: CI
  - name: compute-recommendations
    x-group: Compute recommendations
  - name: config
    x-group: Config
  - name: cost
    x-group: Cost
  - name: gpus
    x-group: GPUs
  - name: graphai
    x-group: Graph AI
  - name: infracost
    x-group: Infracost
  - name: integrations
    x-group: Integrations
  - name: invitations
    x-group: Invitations
  - name: kubernetes-clusters
    x-group: Kubernetes clusters
  - name: marketplace
    x-group: Marketplace
  - name: network-requests
    x-group: Network requests
  - name: notifications
    x-group: Notifications
  - name: oauth
    x-group: OAuth
  - name: oauth-clients
    x-group: OAuth clients
  - name: opencost
    x-group: OpenCost
  - name: organization
    x-group: Audit log
  - name: organizations
    x-group: Organizations
  - name: placement-alternatives
    x-group: Placement alternatives
  - name: reports
    x-group: Reports
  - name: service-map
    x-group: Service map
  - name: settings
    x-group: Settings
  - name: sso
    x-group: Single sign-on
  - name: tenants
    x-group: Tenants
  - name: user
    x-group: Users
  - name: virtual-machines
    x-group: Virtual machines
  - name: virtual-tags
    x-group: Virtual tags
  - name: workflows
    x-group: Workflows
paths:
  /api/v1/tenant/ai/deployments/{id}/metrics:
    get:
      tags:
        - ai-serving
      summary: Get AI serving deployment metrics
      description: >-
        Returns latency quantile, capacity and throughput time-series per model
        for one serving deployment, sourced from the metrics store. Each series
        states its own unit because gateways disagree: vLLM publishes seconds
        and APISIX milliseconds, and a histogram cannot be rescaled after
        recording. Signals the deployment's gateway does not publish are named
        in missing_signals rather than returned as zero. A range of 24h or less
        is answered from the per-replica series; a longer range is answered from
        the deployment-grain series, which carries no replica identity, so a
        replica filter beyond 24h is rejected. The response also states how the
        deployment's pods map to accelerators, and never divides a card's figure
        between pods. cost_per_mtok is the accelerator spend behind a million
        output tokens, at each point of the window; it is returned only where
        the deployment's accelerator rate is known and the deployment is serving
        output tokens, and is named in missing_signals otherwise.
        gateway_connection says why a deployment has no series, since the three
        causes need three different actions: not_connected when no model server
        is connected for this namespace, selector_mismatch when one is but this
        deployment's pods do not carry the labels it is read by, naming that
        selector, and connected when the deployment is being read and any
        silence is the gateway's own. missing_signal_notes explains a missing
        signal the connected server only started publishing in a later release,
        so a request-stage breakdown an older build never emits reads as a
        version gap rather than as a broken scrape.
      parameters:
        - description: Tenant ID
          name: X-CostGraph-Tenant-ID
          in: header
          required: true
          schema:
            type: string
        - description: AI serving deployment ID
          name: id
          in: path
          required: true
          schema:
            type: string
        - description: Time window (e.g. 1h, 6h, 24h, 7d). Default 1h
          name: range
          in: query
          schema:
            type: string
        - description: Resolution step (e.g. 30s, 1m, 5m)
          name: step
          in: query
          schema:
            type: string
        - description: >-
            Serving signals (comma-separated):
            all,time_to_first_token,inter_token,end_to_end,queue,prefill,decode,request_output_tokens,kv_cache_utilization,batch_size,queue_depth,preemptions,concurrency,tokens_per_second,cost_per_mtok.
            Default all
          name: signals
          in: query
          schema:
            type: string
        - description: Quantiles between 0 and 1 (comma-separated). Default 0.5,0.95,0.99
          name: quantiles
          in: query
          schema:
            type: string
        - description: Restrict to one model
          name: model
          in: query
          schema:
            type: string
        - description: Restrict to one replica, as namespace/pod
          name: replica
          in: query
          schema:
            type: string
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                allOf:
                  - $ref: '#/components/schemas/responses.SuccessResponse'
                  - type: object
                    properties:
                      data:
                        $ref: '#/components/schemas/aigateway.ServingLatencyMetrics'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '403':
          description: Forbidden
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '404':
          description: Not Found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/responses.ErrorResponse'
      security:
        - BearerAuth: []
components:
  schemas:
    responses.SuccessResponse:
      type: object
      required:
        - message
        - status
      properties:
        data: {}
        message:
          type: string
          example: some message
        status:
          type: string
          example: success
    aigateway.ServingLatencyMetrics:
      type: object
      properties:
        end:
          type: string
        gateway_connection:
          $ref: '#/components/schemas/aigateway.ServingGatewayConnection'
        gpu_attribution:
          $ref: '#/components/schemas/aigateway.ServingGPUAttribution'
        missing_signal_notes:
          type: object
          additionalProperties:
            type: string
        missing_signals:
          type: array
          items:
            type: string
        rate_window_seconds:
          type: integer
        series:
          type: array
          items:
            $ref: '#/components/schemas/aigateway.ServingLatencySeries'
        start:
          type: string
        step_seconds:
          type: integer
    responses.ErrorResponse:
      type: object
      required:
        - message
        - status
      properties:
        message:
          type: string
          example: some message
        status:
          type: string
          example: error
    aigateway.ServingGatewayConnection:
      type: object
      properties:
        gateway:
          type: string
        label_selector:
          type: string
        reason:
          type: string
        state:
          type: string
    aigateway.ServingGPUAttribution:
      type: object
      properties:
        devices:
          type: integer
        pods:
          type: integer
        reason:
          type: string
        state:
          type: string
    aigateway.ServingLatencySeries:
      type: object
      properties:
        kind:
          type: string
        model:
          type: string
        points:
          type: array
          items:
            $ref: '#/components/schemas/metrics.MetricPoint'
        quantile:
          type: number
        signal:
          type: string
        unit:
          type: string
        usage_kind:
          type: string
    metrics.MetricPoint:
      type: object
      required:
        - t
        - v
      properties:
        t:
          type: integer
        v:
          type: number
  securitySchemes:
    BearerAuth:
      description: Enter "Bearer {token}"
      type: apiKey
      name: Authorization
      in: header

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.