Interpret usage, cost and diagnostics
Section: DOC-IS-usage-cost#interpret-usage-cost-and-diagnostics.
Use usage evidence to explain activity, investigate a request and compare the cost of operating an agent. Start by identifying what each number measures: provider-reported usage, an estimate, an allocation and a customer charge answer different questions.
Generation responses/events and evaluation reads expose selected usage and cost information. Use the schema deployed for your project and keep unavailable values distinct from zero.
1. Compare the same application usage before and after the change
Section: DOC-IS-usage-cost#choose-a-comparable-population.
Use this recipe when the cost of your coaching assistant rose after a change. Start with the previous and current periods, the affected model/profile and the release or configuration change you want to investigate.
- Choose comparable periods and the same application/customer population.
- Separate higher usage from a changed model rate or a changed allocation method.
- Inspect representative runs for extra attempts, longer prompts, tool activity or failures.
- Account for missing or delayed observations before estimating the size of the change.
- If the invoice still differs, take the period, scope and rate basis into billing review.
The result is an explanation with a known measurement basis and remaining gaps, not a claim that a trace estimate is the final charge.
Before comparing two periods or models, keep the project, test/live environment, customer audience, timezone and time window explicit. Check whether a chart counts logical runs, provider attempts, messages or observations. Retries can add attempts and cost without adding a new user turn.
A useful report identifies its source, unit, time range, last refresh and known gaps. A missing observation is not a measured zero. Keep filters and time bounds fixed while paging or exporting, and discard results from a previous scope after changing project or audience.
Evaluation overview cards currently mix sampled coverage/profile values with wider volume metrics. Do not combine them as though they describe one complete cohort. For a controlled comparison, retain a versioned dataset and reconcile evaluated, failed and excluded cases.
2. Separate consumption, allocation and the amount billed
Section: DOC-IS-usage-cost#keep-four-kinds-of-cost-separate.
| Value | What it tells you | How to use it |
|---|---|---|
| Provider spend | Usage or cost reported by a model or another upstream provider | Check its coverage, currency and reporting basis |
| Allocated infrastructure cost | A modeled share of compute, storage, idle capacity or shared services | Compare methods and source windows; retain unallocated amounts |
| Customer price or charge | The amount determined by the applicable commercial terms and metering | Refer to the billing record and the pricing basis effective when usage occurred |
| Invoice amount | The billed financial statement, including applicable adjustments | Reconcile through billing; a dashboard estimate does not settle an invoice |
A price multiplied by usage is not automatically a measurement of provider cost. Likewise, request counts or tokens can be allocation weights without proving how many CPU cycles one customer consumed. Forecasts and allocations should name their method and uncertainty.
Historical source corrections may revise an estimate. They must not silently move old usage to a new payer or reprice it with today's rates. Retain the original currency, payer and commercial basis when investigating a discrepancy. Reporting also does not authorize additional spend: a stale or empty dashboard is not an available-balance check.
3. Look for longer prompts, repeated work or changed tool use
Section: DOC-IS-usage-cost#understand-tokens-and-tools.
Prefer provider-reported counters when available, keeping their field definitions and missing values intact. Input, output, cached, reasoning, audio and image fields can use different accounting bases. Do not assume they always form a disjoint sum or derive text usage by subtracting unrelated counters.
Tool evidence also needs a clear stage:
- Offered: the model request included that tool definition.
- Requested: the model asked to call it.
- Executed: the platform actually attempted the authorized operation.
- Succeeded: that execution reported success.
A requested call rejected by policy was not executed. “Never requested” is useful only across observations where that exact tool/version was offered. A rare tool is not automatically unnecessary.
Run usage is not necessarily conversation-lifetime usage. Compaction changes the active prompt; it does not erase previously incurred consumption. A branch or revert can remove visible messages without refunding prior calls. An unknown context-window size should remain unavailable, not appear as 0% used. See context management.
4. Follow a representative request to its outcome
Section: DOC-IS-usage-cost#investigate-a-request.
Keep the server request ID, approximate time, authorized project/environment, endpoint and observed outcome. Distinguish a network attempt from a retried business operation, and use available run/turn/trace links to follow what happened. An HTTP response or client disconnect does not, by itself, describe every later tool or generation outcome. Generation outcomes and trace inspection cover the existing documented surfaces.
Share safe request identifiers with support; never include API keys, authorization headers or signed download URLs in a diagnostic report.
Upcoming: consolidated usage views
Status: Upcoming — not yet available.
Section: DOC-IS-usage-cost#consolidated-insights
Investigate a cost or quality change in one view. Select the project, audience and time window, compare usage and performance, then open the related run or integration evidence. Each chart identifies what it counts, its source and unit, when it was refreshed and any gaps. Filtering and export preserve those choices.
Infrastructure-cost allocations will remain labeled models with their methods and uncertainty, separately from provider spend, customer prices and invoices. Correcting a report must not silently reprice prior usage or change its payer. The current generation and evaluation reads above remain the available sources. A dashboard query API is not yet available.
Example: Select a project, customer audience and a fixed seven-day window; compare provider spend and model usage, open a related run, then export the same filtered rows. The view labels estimates and incomplete periods.
Upcoming: prompt contribution estimates
Status: Upcoming — not yet available.
Section: DOC-IS-usage-cost#prompt-attribution
Find which parts of a prompt contribute to its size: conversation history, memory, tool definitions and prompt fragments. The attribution view distinguishes measured contributions from tokenizer estimates and estimates scaled to a known input total. A scaled breakdown can add up correctly while its per-source allocation remains modeled; it cannot establish a separate billable charge for a memory, tool definition or prompt fragment.
The view also separates tools offered to the model from tools requested, attempted and completed. Missing observations and an unknown context-window size will remain unavailable rather than becoming zero usage or 0% used. Compaction and history edits will not erase already incurred consumption.
Example: Inspect one generation’s prompt breakdown across history, memory, tool definitions and prompt fragments. Each contribution says whether it is measured, estimated or scaled; an unknown contribution stays unknown.
Upcoming: request capture and retention controls
Status: Upcoming — not yet available.
Section: DOC-IS-usage-cost#request-capture-and-retention
For recurring support investigations, choose which project/environment may retain request bodies and for how long. Review the effective capture policy, use a safe request ID to find a failing attempt, and inspect only its permitted content. After changing capture or retention, review the effect on new requests and any cleanup still pending for older copies.
Use request logs to investigate what an application sent and why a response failed. When content is missing, the log explains whether capture was unavailable, disabled, excluded, redacted, truncated, dropped or still pending.
Eligible request payloads are captured by default. Turn capture off for a project and environment when you do not want that content retained. Mandatory exclusions apply even when capture is enabled. Screenshots, computer-use recordings and retained voice recordings require separate choices. If policy checks or redaction fail, only permitted metadata and a safe failure reason remain. Redaction cannot guarantee that arbitrary free text contains no sensitive information.
Before changing retention, review its effective date, applicable price changes and effect on existing copies. Immediate access withdrawal and completed physical deletion are separate states; backups and external copies have their own consequences. These controls are not currently available through a public request-log or retention endpoint in this guide.
Example: Disable eligible payload capture for one project/environment while retaining permitted request metadata. Before shortening retention, review the effective date, price impact and cleanup consequences; the result distinguishes access removal from pending physical cleanup.
Upcoming: customer-managed telemetry export
Status: Upcoming — not yet available.
Section: DOC-IS-usage-cost#telemetry-export
To use your organization’s monitoring system, select a supported destination and the metadata records it needs, authorize that destination, then follow an initial delivery. Inspect rejected or retrying records before expanding the feed. When pausing an export, review its buffer and missing intervals so resume does not appear to recover history that expired.
Send telemetry to your own monitoring or storage system, starting with metadata-only records. Including payload content requires explicit permission for both the source and destination; permission to capture locally is not permission to export. Self-hosted options support OTLP, a standard telemetry protocol, and object-storage destinations. Each destination documents the records it accepts and how fields are mapped.
Delivery status will distinguish accepted, delivered, rejected, retrying and irrecoverably missing records. Pause and resume must explain whether accepted records remain buffered and which intervals were skipped. A finite buffer can expire, so resume is not a promise to recover all history. Copies already delivered remain subject to the destination's policy.
A public export-management endpoint is not yet available.
Example: Choose an OTLP or object-storage destination and start with metadata only. Including payloads requires a separate explicit choice. After pausing and resuming, the view lists buffered deliveries and any unrecoverable interval.
Next steps
Section: DOC-IS-usage-cost#next-steps.
- Read traces and scores for supported quality and usage evidence.
- Build comparable dataset runs with explicit failure accounting and a reproducibility manifest.
Document ID: DOC-IS-usage-cost. Section identities and revisions.