Is your feature request related to a problem? Please describe.
User reported seeing high CPU utilization in openinference-instrumentation-langchain, specifically in the _update_span -> _convert_io -> _json_dumps workflow.
_update_span is called each time a LangChain run finishes (including every chain, runnable step, parser, tool, LLM call and LangGraph node) -> _convert_io serializes the run's inputs and outputs to JSON -> then _json_dumps for every run.
A couple things make this more expensive than it needs to be:
- The same data is serialized many times. Child runs usually get their parent's data. In LangGraph, the whole graph state (including the message history) is the input and output of every node.
- One bad value causes a second json dump. If any value in the payload isn't serializable,
_json_dumps catches the error and encodes the whole object again with safe_json_dumps.
- Sampling isn't taken into consideration.
_end_trace never checks span.is_recording(), so spans the sampler drops are still fully serialized.
- There may be other ways to make this more performant that I am not capturing here.
Relevant code is in python/instrumentation/openinference-instrumentation-langchain/src/openinference/instrumentation/langchain/_tracer.py
Describe the solution you'd like
A more performant method to minimize how much CPU is utilized during this process.
Describe alternatives you've considered
Upgrading the instrumentor. The serialization path hasn't changed much between the users reported v0.1.61 and the current v0.1.76
Is your feature request related to a problem? Please describe.
User reported seeing high CPU utilization in
openinference-instrumentation-langchain, specifically in the_update_span->_convert_io->_json_dumpsworkflow._update_spanis called each time a LangChain run finishes (including every chain, runnable step, parser, tool, LLM call and LangGraph node) ->_convert_ioserializes the run's inputs and outputs to JSON -> then_json_dumpsfor every run.A couple things make this more expensive than it needs to be:
_json_dumpscatches the error and encodes the whole object again withsafe_json_dumps._end_tracenever checksspan.is_recording(), so spans the sampler drops are still fully serialized.Relevant code is in
python/instrumentation/openinference-instrumentation-langchain/src/openinference/instrumentation/langchain/_tracer.pyDescribe the solution you'd like
A more performant method to minimize how much CPU is utilized during this process.
Describe alternatives you've considered
Upgrading the instrumentor. The serialization path hasn't changed much between the users reported v0.1.61 and the current v0.1.76