Skip to main content
A bill tells you what you spent. It doesn’t tell you which of that spend is justified. Insights close that gap. CostGraph watches your connected clouds, Kubernetes clusters, virtual machines, and SaaS providers, and raises an insight when something needs your attention. That could be a cost that jumped or a resource you pay for but don’t use. It could also be a workload that keeps failing, or a cheaper way to run what you already run. Each insight names the resource, says what happened, puts a monthly dollar figure on it where one applies, and suggests where to start the fix.

Find your insights

Open Insights in the sidebar. The top of the page summarizes your environment:
  • Total spend for the selected period.
  • Potential cost impact: the monthly amount at stake across open insights.
  • Active anomalies: how many insights are open, and how many are new since yesterday.
  • Cluster efficiency and cost by provider, for context.
Under Needs attention is the list of insights. Each row shows its severity, what was found, the resource, the monthly cost impact, and how long it has been open. Search by name, or filter the list by:
  • Severity: critical, warning, or info
  • Status: active, resolved, or muted
  • Category: cost, waste, efficiency, reliability, or data quality
  • Resource type
  • Provider
Click Export to download the current list as a CSV file.

Read an insight

Select an insight to open it. The detail page shows:
  • Headline and summary: what happened, and why it usually happens.
  • Chart: the measurement that triggered the insight, against its normal level where one exists.
  • What CostGraph observed: the specific readings behind the finding.
  • Cost impact: the monthly amount at stake, as extra spend or as a saving.
  • Value vs baseline: how far the measurement moved from normal.
  • Recurrence: when it was first detected and how often it has come back.
  • At a glance: where the resource runs, such as its provider, region, cluster, namespace, or pod.
  • Root-cause breakdown: the resources or dimensions that drove a change, with each one’s share.
  • Cheaper placements: for compute, equivalent options on other providers or regions that cost less.
  • Related recommendations: rightsizing actions for the same resource, such as downsize, upsize, or terminate.
Not every section appears on every insight. CostGraph shows the ones that apply.

Severity

  • Critical: a large cost impact or a severe deviation. Look at these first.
  • Warning: a clear finding worth acting on soon.
  • Info: worth knowing, with little or no immediate cost.

Ask GraphAI

Every insight connects to GraphAI, the CostGraph assistant. Click Investigate with GraphAI to open a conversation about the insight, with its context already loaded. GraphAI can trace the resource, look at its usage and cost, and propose a fix. The detail page also lists suggested questions for that insight, such as what changed or what the resource has in common with others. Select one to ask it directly. You can read insights from your own tools as well, through the CostGraph MCP server.

How insights open and resolve

CostGraph checks for new insights every hour. When the same problem appears again, CostGraph updates the existing insight instead of opening a duplicate, and counts the recurrence. Insights resolve themselves. When the condition clears, for example after you delete an orphaned disk or a cost returns to normal, the insight moves to Resolved. The condition has to stay clear for a few hourly checks first. That way, a value that dips briefly doesn’t close and reopen an insight. Resolved insights stay in the list for 90 days, under the Resolved status filter.

Mute an insight

If an insight is expected, mute it. Click Mute on the insight and choose a scope:
  • This resource: mute this kind of insight on this resource only.
  • All of this kind: mute this kind of insight across every resource.
Add a reason if you want one, so your team knows why it was muted. Muted insights move to the Muted status. Future matches of the same scope stay muted, so they don’t return to Needs attention.

Insight types

Cost changes

  • Cost spiked: a resource’s cost rose well over its normal level.
  • Cost fell: a resource’s cost dropped well under its normal level.
  • Cost per unit of work rose: cost per core-hour climbed, so the same work now costs more.
  • Billed usage spiked: usage of a billed service, such as requests, storage, or tokens, rose well over normal.
  • Network traffic spiked: traffic from a resource rose well over normal, often a sign of new data transfer charges.
  • Spend is tracking over forecast: at the current rate, this month’s spend finishes well over last month’s.
  • Spend with no identified owner: a provider or service is costing money that no resource, tag, or team accounts for.

Waste

  • Idle VM: a virtual machine is running with little or no CPU or network activity.
  • Over-provisioned VM: a virtual machine is much larger than the workload on it.
  • Over-provisioned compute: Kubernetes workloads request far more CPU or memory than they use.
  • Idle GPU: a GPU on a virtual machine is doing no work.
  • Idle GPU partitions: MIG partitions on a GPU are allocated but unused.
  • Over-provisioned database: a managed database has far more compute or storage than it uses.
  • Idle volume: a mounted Kubernetes volume has no reads or writes.
  • Over-provisioned volume: a volume is much larger than the data on it.
  • Unattached volume: a Kubernetes persistent volume is not bound to any workload.
  • Orphaned disk: a cloud disk is not attached to anything and is still billing.
  • Stopped resource still billing: a stopped instance, powered-off server, stopped service, or detached volume is still on your bill.
  • Unused seats: you pay for seats, such as GitHub Copilot or Buildkite, that nobody has used recently.
  • Unused included capacity: a plan or commitment includes usage you are not consuming.
  • CI artifacts kept too long: a repository keeps build artifacts longer than it needs, and the storage adds up.

Savings opportunities

  • Run on fewer nodes: a cluster’s workloads fit on fewer nodes, with the monthly saving from removing the rest.
  • Cheaper placement: the same compute shape costs less on another provider or region.
  • Cheaper AI model or provider: the same model, or a comparable one, costs less elsewhere for your usage.
  • Prompt caching costs more than it saves: caching writes cost more than the cached reads save for this usage.
  • All spend is on-demand: steady spend on a provider has no commitment or reservation covering it.
  • Plan allowance regularly exceeded: you run over a plan’s allowance most months and pay overage rates, so a larger plan may cost less.
  • Paying for extended support: an old version of a managed service is billed at an extended support premium. Upgrading removes it.
  • Cheaper CI runners: a repository’s CI jobs cost less on another runner provider.

Reliability

  • Out of memory: a container keeps being killed for running out of memory.
  • Crash loop: a container keeps restarting.
  • Unstable workload: a workload keeps flapping between healthy and unhealthy.
  • Pod churn: a workload keeps replacing its pods.
  • Volume nearly full: a Kubernetes volume is close to running out of space.
  • Model server capacity mismatch: a model server is short of capacity, holds a GPU without serving, or reserves far more than it uses.

Usage changes

  • Usage spiked: CPU, memory, or network usage on a workload or VM rose well over normal, which can mean it needs more room.
  • Usage fell: CPU, memory, or network usage dropped well under normal, so the resource may be larger than it needs.

Data quality

  • Usage fell across many resources together: many resources dropped at the same moment. That usually means data stopped arriving or a shared dependency failed.