Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An LLM telemetry table has no single natural denominator. A row may count a model inference call, a broader GenAI operation, or a top-level application request—and one application request can contain several model and tool calls. Before comparing counts, token totals, latency, or cost, identify exactly what the table counts and what it includes.

What does one row or observation count?

OpenTelemetry uses “GenAI operation” broadly: an operation may be a request to a language model, a function call, or another distinct action in a larger workflow. Its inference span is narrower: a client call to a GenAI model or service that generates a response or requests a tool call. Those are different possible units, not interchangeable labels.

A trace can make the distinction visible. In OpenTelemetry’s walkthrough, a top-level invoke_agent span contains child chat spans for model calls and execute_tool spans for tool calls. A table that counts child model-call spans is therefore not counting top-level application requests. The aggregation from child operations to a request-level measure is an implementation choice, not a universal rule.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Illustrative example

Suppose one user request starts an agent workflow, which makes two model calls and invokes a tool once. A call-level table could show two model-call observations; a broader operation count could include the tool operation too; a request-level table could show one application request. These figures describe different units, not conflicting measurements. This is an illustrative workflow, not a reported statistic.

How should you compare counts and averages?

Read the metric name and aggregation together. A count of observations, a sum over a time window, a rate, and an average per call or per request answer different questions. For an average, check both the numerator and denominator: “tokens per request” is ambiguous unless “request” is defined as an application request, agent turn, or model inference call.

  • Unit and scope: Is the denominator an application request, an inference call, or any GenAI operation?
  • Aggregation and time window: Is the figure an observation count, total, rate, or average, and over what period?
  • Call composition: Are retries, tool calls, embeddings, or multiple model calls included? If child operations are rolled up to requests, how are they combined?
  • Grouping: Which provider and exact requested model are represented? OpenTelemetry notes that a provider attribute may identify the configured client or proxy rather than the ultimate upstream provider.

Retries and repeated prompts matter because a workflow may send input again on a later call. A token sum across model calls can therefore exceed the token count associated with any one user request. A request-level metric should state how those child-call values are combined.

What do token totals include?

Input and output are separate observations in OpenTelemetry’s walkthrough, and NVIDIA’s implementation reference likewise records two observations for each downstream LLM call, distinguished by token type. Do not treat a combined token total as self-explanatory: identify input versus output and any included categories, such as cached, image, or reasoning tokens.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenTelemetry says input-token totals should include all input types, including cached tokens, while detailed usage attributes are subsets of the total counts. When a provider exposes both billed and consumed token counts, its guidance recommends reporting billed counts so the measurement aligns with charged units. A dashboard should say which basis it uses; billed and consumed values can answer different questions.

For token efficiency, specify whether the denominator is a model call or an application request. A workflow-level “tokens per request” average may include several calls and their repeated inputs, while a per-call average does not describe the same workload.

What does missing streaming usage mean?

Missing usage data is not automatically zero usage. In NVIDIA’s implementation reference, streaming token usage is emitted only when the upstream provider returns a usage field. If that field is absent, no observation is recorded; that is deliberately different from recording a zero-token observation.

When interpreting a chart, distinguish an actual zero from an absent observation or a gap in provider reporting. Otherwise, an average or total may imply coverage that the underlying telemetry does not have.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which latency does the table measure?

OpenTelemetry defines inference duration as the time from issuing a model request until the response is fully received, or until the operation ends in error or cancellation. That is not automatically the duration of the entire agent workflow or application request. A broader workflow can include tool execution and other work beyond the model call.

Check the span or metric scope before calling a value “model latency.” If a number covers the top-level request, label it as request or workflow duration; use inference duration for the model-call boundary.

How should cost per request be interpreted?

“Cost per request” requires both a stated request boundary and a cost basis aligned with billing. Decide whether one request means an application request, an agent turn, or a model call, then state how costs from its child calls are attributed. Billed token counts can help align token-based estimates with charged units, but the telemetry conventions do not define one universal cost formula.

Token and duration metrics can support cost estimation, but they are not, by themselves, a complete bill. A table should name its billing or estimation basis rather than presenting an undefined cost-per-request figure as directly comparable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can telemetry measure usage without capturing prompts?

Yes. OpenTelemetry’s walkthrough describes default telemetry that can include metadata such as model names, token counts, and durations without recording prompt or tool content. Prompt and tool-argument capture is opt-in because that content can contain sensitive information. Enabling it adds message and tool details to spans.

For usage measurement and latency analysis, content capture is not required. If debugging calls for content, treat the decision as a separate privacy and data-handling choice rather than assuming that a token metric requires prompt logging.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should a well-labeled table disclose?

  • Counted unit: application request, inference call, or broader GenAI operation.
  • Metric meaning: observation count, token sum, rate, average, or duration.
  • Aggregation: time window and, for request-level values, how child calls are rolled up.
  • Token basis: input or output, included token categories, and billed or consumed basis.
  • Workflow coverage: treatment of retries, tools, embeddings, and multiple model calls.
  • Latency boundary: inference call versus full request or workflow.
  • Model attribution: provider and exact requested model, with any proxy or client attribution made clear.
  • Missing data: whether absent usage fields are omitted rather than recorded as zero.

OpenTelemetry’s GenAI conventions define attributes and metrics for operation, provider, model, and token usage, including separate usage metrics for input, output, cache-read input, cache-write input, and reasoning output. The conventions are actively developed, so check the documentation and stability status for the version you implement. The semantics help label a table; they do not choose its aggregation boundary for you.

Sources

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.