Describe the bug
LLMTokenMetricsDataExtractor in openinference-instrumentation-pipecat/.../_attributes.py:544 maps pipecat's LLMTokenUsage onto the span with three gaps:
llm.token_count.prompt copies prompt_tokens. For AnthropicLLMService that value excludes the cache, so the span shows prompt below prompt_details.cache_read. spec/semantic_conventions.md:214-218 wants the cache counts added back.
cache_creation_input_tokens is never read, so prompt_details.cache_write is never set.
- Both audio details read
audio_tokens (:561, :567). LLMTokenUsage has no such field in pipecat 1.3.0 or 1.12.0. Since 1.6.0 it has input_audio_tokens and output_audio_tokens. Neither audio attribute is ever emitted.
To Reproduce
This is repro.py. The first usage is what AnthropicLLMService._report_usage_metrics builds:
from pipecat.metrics.metrics import LLMTokenUsage, LLMUsageMetricsData
from openinference.instrumentation.pipecat._attributes import _llm_usage_metrics_data_extractor as ex
anthropic = LLMTokenUsage(prompt_tokens=100, completion_tokens=40, cache_read_input_tokens=900,
cache_creation_input_tokens=50, total_tokens=1090)
audio = LLMTokenUsage(prompt_tokens=300, completion_tokens=120, total_tokens=420,
input_audio_tokens=250, output_audio_tokens=100)
for name, u in (("anthropic", anthropic), ("audio", audio)):
a = ex.extract_from_metrics_data(LLMUsageMetricsData(processor="p", value=u))
print(name, {k.replace("llm.token_count.", ""): v for k, v in a.items() if "token_count" in k})
$ python repro.py
# pipecat's startup INFO banner omitted
anthropic {'prompt': 100, 'completion': 40, 'total': 1090, 'prompt_details.cache_read': 900}
audio {'prompt': 300, 'completion': 120, 'total': 420}
Expected behavior
For the Anthropic usage, prompt 1050 and prompt_details.cache_write 50. For the audio usage, prompt_details.audio 250 and completion_details.audio 100.
Environment
openinference-instrumentation-pipecat at main 90c1e28, pipecat-ai 1.12.0, Python 3.12.3, Linux.
Additional context
pipecat documents on LLMTokenUsage that Anthropic and Bedrock report prompt_tokens net. OpenAI-compatible services report it gross.
Gaps 2 and 3 need no decision. I will open a pull request for them and extend test_extract_llm_usage_metrics.
@mikeldking, for gap 1, should the extractor fold by service name, or should pipecat expose a gross prompt count?
Describe the bug
LLMTokenMetricsDataExtractorinopeninference-instrumentation-pipecat/.../_attributes.py:544maps pipecat'sLLMTokenUsageonto the span with three gaps:llm.token_count.promptcopiesprompt_tokens. ForAnthropicLLMServicethat value excludes the cache, so the span showspromptbelowprompt_details.cache_read.spec/semantic_conventions.md:214-218wants the cache counts added back.cache_creation_input_tokensis never read, soprompt_details.cache_writeis never set.audio_tokens(:561,:567).LLMTokenUsagehas no such field in pipecat 1.3.0 or 1.12.0. Since 1.6.0 it hasinput_audio_tokensandoutput_audio_tokens. Neither audio attribute is ever emitted.To Reproduce
This is
repro.py. The first usage is whatAnthropicLLMService._report_usage_metricsbuilds:Expected behavior
For the Anthropic usage,
prompt1050 andprompt_details.cache_write50. For the audio usage,prompt_details.audio250 andcompletion_details.audio100.Environment
openinference-instrumentation-pipecatatmain90c1e28, pipecat-ai 1.12.0, Python 3.12.3, Linux.Additional context
pipecat documents on
LLMTokenUsagethat Anthropic and Bedrock reportprompt_tokensnet. OpenAI-compatible services report it gross.Gaps 2 and 3 need no decision. I will open a pull request for them and extend
test_extract_llm_usage_metrics.@mikeldking, for gap 1, should the extractor fold by service name, or should pipecat expose a gross prompt count?