> ## Documentation Index
> Fetch the complete documentation index at: https://docs.costgraph.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Insights

> See what changed in your spend, what is wasted, and what to do about it

A bill tells you what you spent. It doesn't tell you which of that spend is
justified. **Insights** close that gap. CostGraph watches your connected clouds,
Kubernetes clusters, virtual machines, and SaaS providers, and raises an insight
when something needs your attention. That could be a cost that jumped or a
resource you pay for but don't use. It could also be a workload that keeps
failing, or a cheaper way to run what you already run.

Each insight names the resource, says what happened, puts a monthly dollar
figure on it where one applies, and suggests where to start the fix.

## Find your insights

Open **Insights** in the sidebar. The top of the page summarizes your
environment:

* **Total spend** for the selected period.
* **Potential cost impact**: the monthly amount at stake across open insights.
* **Active anomalies**: how many insights are open, and how many are new since
  yesterday.
* **Cluster efficiency** and **cost by provider**, for context.

Under **Needs attention** is the list of insights. Each row shows its severity,
what was found, the resource, the monthly cost impact, and how long it has been
open. Search by name, or filter the list by:

* **Severity**: critical, warning, or info
* **Status**: active, resolved, or muted
* **Category**: cost, waste, efficiency, reliability, or data quality
* **Resource type**
* **Provider**

Click **Export** to download the current list as a CSV file.

## Read an insight

Select an insight to open it. The detail page shows:

* **Headline and summary**: what happened, and why it usually happens.
* **Chart**: the measurement that triggered the insight, against its normal
  level where one exists.
* **What CostGraph observed**: the specific readings behind the finding.
* **Cost impact**: the monthly amount at stake, as extra spend or as a saving.
* **Value vs baseline**: how far the measurement moved from normal.
* **Recurrence**: when it was first detected and how often it has come back.
* **At a glance**: where the resource runs, such as its provider, region,
  cluster, namespace, or pod.
* **Root-cause breakdown**: the resources or dimensions that drove a change,
  with each one's share.
* **Cheaper placements**: for compute, equivalent options on other providers or
  regions that cost less.
* **Related recommendations**: rightsizing actions for the same resource, such
  as downsize, upsize, or terminate.

Not every section appears on every insight. CostGraph shows the ones that apply.

### Severity

* **Critical**: a large cost impact or a severe deviation. Look at these first.
* **Warning**: a clear finding worth acting on soon.
* **Info**: worth knowing, with little or no immediate cost.

## Ask GraphAI

Every insight connects to GraphAI, the CostGraph assistant. Click **Investigate
with GraphAI** to open a conversation about the insight, with its context
already loaded. GraphAI can trace the resource, look at its usage and cost, and
propose a fix.

The detail page also lists suggested questions for that insight, such as what
changed or what the resource has in common with others. Select one to ask it
directly.

You can read insights from your own tools as well, through the
[CostGraph MCP server](/costgraph/mcp/usage).

## How insights open and resolve

CostGraph checks for new insights every hour. When the same problem appears
again, CostGraph updates the existing insight instead of opening a duplicate,
and counts the recurrence.

Insights resolve themselves. When the condition clears, for example after you
delete an orphaned disk or a cost returns to normal, the insight moves to
**Resolved**. The condition has to stay clear for a few hourly checks first.
That way, a value that dips briefly doesn't close and reopen an insight.
Resolved insights stay in the list for 90 days, under the **Resolved** status
filter.

## Mute an insight

If an insight is expected, mute it. Click **Mute** on the insight and choose a
scope:

* **This resource**: mute this kind of insight on this resource only.
* **All of this kind**: mute this kind of insight across every resource.

Add a reason if you want one, so your team knows why it was muted. Muted
insights move to the **Muted** status. Future matches of the same scope stay
muted, so they don't return to **Needs attention**.

## Insight types

### Cost changes

* **Cost spiked**: a resource's cost rose well over its normal level.
* **Cost fell**: a resource's cost dropped well under its normal level.
* **Cost per unit of work rose**: cost per core-hour climbed, so the same work
  now costs more.
* **Billed usage spiked**: usage of a billed service, such as requests, storage,
  or tokens, rose well over normal.
* **Network traffic spiked**: traffic from a resource rose well over normal,
  often a sign of new data transfer charges.
* **Spend is tracking over forecast**: at the current rate, this month's spend
  finishes well over last month's.
* **Spend with no identified owner**: a provider or service is costing money
  that no resource, tag, or team accounts for.

### Waste

* **Idle VM**: a virtual machine is running with little or no CPU or network
  activity.
* **Over-provisioned VM**: a virtual machine is much larger than the workload
  on it.
* **Over-provisioned compute**: Kubernetes workloads request far more CPU or
  memory than they use.
* **Idle GPU**: a GPU on a virtual machine is doing no work.
* **Idle GPU partitions**: MIG partitions on a GPU are allocated but unused.
* **Over-provisioned database**: a managed database has far more compute or
  storage than it uses.
* **Idle volume**: a mounted Kubernetes volume has no reads or writes.
* **Over-provisioned volume**: a volume is much larger than the data on it.
* **Unattached volume**: a Kubernetes persistent volume is not bound to any
  workload.
* **Orphaned disk**: a cloud disk is not attached to anything and is still
  billing.
* **Stopped resource still billing**: a stopped instance, powered-off server,
  stopped service, or detached volume is still on your bill.
* **Unused seats**: you pay for seats, such as GitHub Copilot or Buildkite,
  that nobody has used recently.
* **Unused included capacity**: a plan or commitment includes usage you are not
  consuming.
* **CI artifacts kept too long**: a repository keeps build artifacts longer
  than it needs, and the storage adds up.

### Savings opportunities

* **Run on fewer nodes**: a cluster's workloads fit on fewer nodes, with the
  monthly saving from removing the rest.
* **Cheaper placement**: the same compute shape costs less on another provider
  or region.
* **Cheaper AI model or provider**: the same model, or a comparable one, costs
  less elsewhere for your usage.
* **Prompt caching costs more than it saves**: caching writes cost more than the
  cached reads save for this usage.
* **All spend is on-demand**: steady spend on a provider has no commitment or
  reservation covering it.
* **Plan allowance regularly exceeded**: you run over a plan's allowance most
  months and pay overage rates, so a larger plan may cost less.
* **Paying for extended support**: an old version of a managed service is
  billed at an extended support premium. Upgrading removes it.
* **Cheaper CI runners**: a repository's CI jobs cost less on another runner
  provider.

### Reliability

* **Out of memory**: a container keeps being killed for running out of memory.
* **Crash loop**: a container keeps restarting.
* **Unstable workload**: a workload keeps flapping between healthy and
  unhealthy.
* **Pod churn**: a workload keeps replacing its pods.
* **Volume nearly full**: a Kubernetes volume is close to running out of space.
* **Model server capacity mismatch**: a model server is short of capacity,
  holds a GPU without serving, or reserves far more than it uses.

### Usage changes

* **Usage spiked**: CPU, memory, or network usage on a workload or VM rose well
  over normal, which can mean it needs more room.
* **Usage fell**: CPU, memory, or network usage dropped well under normal, so
  the resource may be larger than it needs.

### Data quality

* **Usage fell across many resources together**: many resources dropped at the
  same moment. That usually means data stopped arriving or a shared dependency
  failed.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.