Skip to content

[bug] pipecat: prompt token count not folded, cache_write and audio details never set #3893

Description

@chrikrah

Describe the bug

LLMTokenMetricsDataExtractor in openinference-instrumentation-pipecat/.../_attributes.py:544 maps pipecat's LLMTokenUsage onto the span with three gaps:

  1. llm.token_count.prompt copies prompt_tokens. For AnthropicLLMService that value excludes the cache, so the span shows prompt below prompt_details.cache_read. spec/semantic_conventions.md:214-218 wants the cache counts added back.
  2. cache_creation_input_tokens is never read, so prompt_details.cache_write is never set.
  3. Both audio details read audio_tokens (:561, :567). LLMTokenUsage has no such field in pipecat 1.3.0 or 1.12.0. Since 1.6.0 it has input_audio_tokens and output_audio_tokens. Neither audio attribute is ever emitted.

To Reproduce

This is repro.py. The first usage is what AnthropicLLMService._report_usage_metrics builds:

from pipecat.metrics.metrics import LLMTokenUsage, LLMUsageMetricsData
from openinference.instrumentation.pipecat._attributes import _llm_usage_metrics_data_extractor as ex
anthropic = LLMTokenUsage(prompt_tokens=100, completion_tokens=40, cache_read_input_tokens=900,
                          cache_creation_input_tokens=50, total_tokens=1090)
audio = LLMTokenUsage(prompt_tokens=300, completion_tokens=120, total_tokens=420,
                      input_audio_tokens=250, output_audio_tokens=100)
for name, u in (("anthropic", anthropic), ("audio", audio)):
    a = ex.extract_from_metrics_data(LLMUsageMetricsData(processor="p", value=u))
    print(name, {k.replace("llm.token_count.", ""): v for k, v in a.items() if "token_count" in k})
$ python repro.py
# pipecat's startup INFO banner omitted
anthropic {'prompt': 100, 'completion': 40, 'total': 1090, 'prompt_details.cache_read': 900}
audio {'prompt': 300, 'completion': 120, 'total': 420}

Expected behavior

For the Anthropic usage, prompt 1050 and prompt_details.cache_write 50. For the audio usage, prompt_details.audio 250 and completion_details.audio 100.

Environment

openinference-instrumentation-pipecat at main 90c1e28, pipecat-ai 1.12.0, Python 3.12.3, Linux.

Additional context

pipecat documents on LLMTokenUsage that Anthropic and Bedrock report prompt_tokens net. OpenAI-compatible services report it gross.

Gaps 2 and 3 need no decision. I will open a pull request for them and extend test_extract_llm_usage_metrics.

@mikeldking, for gap 1, should the extractor fold by service name, or should pipecat expose a gross prompt count?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions