Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 9 additions & 3 deletions docs/observability/opentelemetry_v2.md
Original file line number Diff line number Diff line change
Expand Up @@ -278,9 +278,9 @@ Open your Arize project; the trace appears under the project named by `ARIZE_PRO
| `llm.model_name`, `llm.provider` | model, provider |
| `llm.token_count.prompt`, `completion`, `total` | usage split |
| `llm.invocation_parameters` | JSON blob of request params |
| `llm.input_messages.{idx}.message.role`, `content` | prompt (content capture on) |
| `llm.output_messages.{idx}.message.role`, `content` | response (content capture on) |
| `input.value`, `output.value` | JSON arrays of the same (content capture on) |
| `llm.input_messages.{idx}.message.role`, `content` | prompt (content capture on), [capped](#chat-messages-are-capped) |
| `llm.output_messages.{idx}.message.role`, `content` | response (content capture on), [capped](#chat-messages-are-capped) |
| `input.value`, `output.value` | JSON arrays of every message's role and text (content capture on) |
| `llm.tools.{idx}.tool.name`, `description`, `json_schema` | tool definitions, [capped](#tool-definitions-are-capped) |

See the full [OpenInference spec](https://github.com/Arize-ai/openinference/blob/main/spec/semantic_conventions.md) for the definitive vocabulary.
Expand Down Expand Up @@ -512,6 +512,12 @@ The ceiling is span-wide, not per vocabulary. Tool definitions may claim a quart

`litellm.request.tools.declared` always carries the true total, so you can tell when the per-tool detail was truncated. Requests declaring fewer tools than the allowance keep full detail.

#### Chat messages are capped

The `openinference` mapper's `llm.input_messages.{idx}.*` and `llm.output_messages.{idx}.*` keys are the other unbounded family: two attributes per message, for the prompt and the response alike, on the same span. Past a few dozen turns they alone exceed the 128-attribute default and evict the `gen_ai.*` model, usage, cost, and finish-reason attributes written before them. The per-index keys therefore share one span-wide allowance of an eighth of the budget, 8 messages total under the default limit, with at least half of it reserved for the response so a long prompt can never push the completion off the span. The prompt's share goes to message 0 and the most recent turns, under their original indices: a 60-turn conversation with one reply indexes `llm.input_messages.0`, `llm.input_messages.54` through `.59`, and `llm.output_messages.0`. The selection is by position, not role: message 0 and the newest prompt messages always keep a short key of their own, which matters when `OTEL_SPAN_ATTRIBUTE_VALUE_LENGTH_LIMIT` clips the `input.value` blob before it reaches the end of the conversation.

The cap only touches the per-index convenience keys. `input.value` and `output.value` still list every message's role and text, and the canonical `gen_ai.input.messages` and `gen_ai.output.messages` blobs carry the full message objects (tool calls and non-text parts included), so the whole conversation stays on the span and Arize keeps rendering it. Conversations shorter than the allowance are indexed in full. Phoenix compacts the per-index keys into a dense list when it renders a span, so the gap in indices shows up there as a shorter message list, in the same order.

Response, usage, cost, identity:

| Attribute | When set |
Expand Down
Loading