Agent Observability

Every agent-native app gets observability out of the box. Traces, automated evals, user feedback, and A/B experiments work with zero configuration — all data lives in the app's own SQL database.

This page covers agent quality metrics: traces, cost, evals, and feedback stored in your database. For product analytics (your app's events flowing to PostHog/Mixpanel/Amplitude), see Tracking.

Three things called "evals"/"observability" — which do I want?

These three pages are easy to confuse. Pick by the question you're asking:

Page The question it answers When it runs Concern
Observability evals (this page, the Evals tab) "How did my real production runs do?" Passive, after every run (LLM-judge sampled) Quality
CI Eval Gate (*.eval.ts) "Does the agent do the right thing on this fixed input?" Active, deterministic, a CI/deploy gate Quality
Observational Memory "Is this long thread staying cheap and inside the window?" Background compaction on long threads Cost / context

Observability and the CI Eval Gate both score quality but from opposite ends — passive post-hoc scoring of real traffic vs. active pass/fail checks on fixed inputs. Promote a completed production trace into a *.eval.ts case (dashboard Promote to eval, promote-trace-eval, or agent-native eval promote <runId> --write evals/from-trace.eval.ts) so the miss becomes the next CI gate. See From a production trace. Observational Memory is unrelated to quality; it's about token cost and context-window pressure.

What's captured automatically

When a user sends a message, the framework automatically records:

  • Token usage — input, output, cache read, cache write
  • Cost — computed from token counts and model pricing
  • Latency — total duration and time per tool call
  • Tool calls — which actions were invoked, success/error status, duration
  • Automated evals — 5 quality scores computed after every run

No code changes needed. The instrumentation hooks into production-agent.ts transparently.

Every run feeds the loop
Agent run
production-agent.ts
Captured automaticallytokens · cost · latency · tool calls
Traces & spans
Evals (5 scorers + LLM judge)
Feedback & frustration index
Dashboard
scoped to the signed-in user

One agent run produces a trace, automated scores, and a feedback hook — all stored in the app's own SQL and surfaced on the dashboard. Experiments split traffic across config variants.

The dashboard

Add the dashboard to any template with a single route:

app/routes/observability.tsx
import { ObservabilityDashboard } from "@agent-native/toolkit/app/observability";
export default function ObservabilityPage() {
  return (
    <div className="min-h-screen bg-background p-6">
      <ObservabilityDashboard />
    </div>
  );
}

All data is scoped to the signed-in user; there is no cross-user admin view today.

The dashboard has 6 tabs:

Tab What it shows
Overview Key metrics — runs, cost, latency, tool success rate, satisfaction, eval score
Conversations Trace list with drill-down to individual spans (agent_run, llm_call, tool_call)
Evals Automated eval scores by criteria, trends over time
Experiments A/B test list with status badges, variant results with confidence intervals
Feedback Thumbs up/down stream, category breakdown, frustration scores
Human review Review the ask and answer, record feedback, and draft instruction updates

Conversations

Select a run to inspect its spans. Expand any span to view captured tool inputs, outputs, errors, and metadata. Use Open full conversation to jump to the associated chat thread.

Tool arguments, results, and full error bodies are off by default. Enable captureToolArgs to store arguments and captureToolResults to store results and full error bodies in the observability config. Sensitive fields are redacted before storage. A failed tool always keeps a bounded first line of its error, with credentials, emails, and long opaque ids redacted, so failures stay diagnosable even when captureToolResults is off.

Prompt capture is off by default. Set capturePrompts: true in the observability config, or set AGENT_NATIVE_OBSERVABILITY_CAPTURE_PROMPTS=true, to store each llm_call span's input and output for this detail view. The system prompt is excluded, structured sensitive fields are redacted, and each side is capped at 128 KiB. Existing spans are not backfilled.

Human review

The Human review tab shows each run’s ask and answer, lets you record feedback, and lets you draft an instruction update for a human to review. Saving the draft does not change instructions automatically.

User feedback

Explicit feedback

Thumbs up/down buttons render inline on every agent message in the chat UI. Thumbs down opens a category popover (Inaccurate, Not helpful, Wrong tool, Too slow). This is wired into AssistantChat.tsx automatically.

Implicit feedback (frustration index)

The framework computes a Frustration Index (0-100) from conversation signals:

Signal Weight What it detects
Rephrasing 30% User repeats similar messages
Retry patterns 20% "Try again", "no that's wrong"
Abandonment 20% Session ends shortly after response
Sentiment 15% Negative language patterns
Length trend 15% Declining message lengths

Score interpretation: 0-20 = healthy, 20-40 = friction, 40-60 = dissatisfied, 60+ = broken session.

Automated evals

Five deterministic scorers run after every agent run:

Criteria What it measures Score range
tool_success_rate % of tool calls without errors 0-1
step_efficiency Penalizes excessive LLM iterations for tool-using runs 0-1
latency_score Normalized against 10s/tool baseline 0-1
cost_efficiency Normalized against cost baseline 0-1
error_recovery Did the agent recover from tool errors? 0 or 1

LLM-as-judge (optional)

Enable sampled LLM-based evaluation by setting evalSampleRate:

// server/plugins/config.ts
import { defineAppConfig } from "@agent-native/core/server";

export default defineAppConfig({
  observability: {
    enabled: true,
    evalSampleRate: 0.05, // 5%
  },
});

Custom criteria use natural language rubrics:

const criteria = {
  name: "helpfulness",
  description: "Was the response helpful and complete?",
  rubric: "0.0 = unhelpful, 0.5 = partially helpful, 1.0 = fully resolved",
};

A/B experiments

Test different models, temperatures, or agent configurations:

// Create via API
POST /_agent-native/observability/experiments
{
  "name": "model-a-vs-b",
  "variants": [
    { "id": "control", "weight": 50, "config": { "model": "<your-model-id>" } },
    { "id": "treatment", "weight": 50, "config": { "model": "<other-model-id>" } }
  ],
  "metrics": ["cost", "latency", "satisfaction"]
}

// Start the experiment
PUT /_agent-native/observability/experiments/:id
{ "status": "running" }

Use the real model identifiers your engine accepts in place of <your-model-id> / <other-model-id> (model names change often — check your provider/engine for the current ids). The agent loop automatically resolves the user's variant and applies the config override. Assignment uses consistent hashing — same user always gets the same variant.

Consistent-hash variant assignment
User id
consistent hash
control · 50%config override A
treatment · 50%config override B
Results per variant
cost · latency · satisfaction

Each user hashes to a stable variant, the loop applies that variant's config override, and results roll up per variant with confidence intervals.

Configuration

Set from a server plugin with defineAppConfig(). Every field also has a deployment environment variable alias — see the declared configuration reference.

// server/plugins/config.ts
import { defineAppConfig } from "@agent-native/core/server";

export default defineAppConfig({
  observability: {
    enabled: true, // Master switch
    capturePrompts: false, // Store prompt content in traces
    captureToolArgs: false, // Store action input arguments
    captureToolResults: false, // Store action results
    captureLlmSpans: true, // Emit one $ai_span per tool call to PostHog
    evalSampleRate: 0, // 0-1, fraction of runs to LLM-judge
  },
});

Content is redacted by default: token counts, costs, and timing are stored, plus, for a failed tool, a bounded first line of its error with credentials, emails, and long opaque ids redacted. Prompts, tool inputs, tool results, and full error bodies stay off unless capturePrompts, captureToolArgs, and captureToolResults are on. Turn them on only when you need prompt or argument content for debugging.

API endpoints

All auto-mounted at /_agent-native/observability/:

Method Path Purpose
GET / Overview stats
GET /traces List trace summaries
GET /traces/:runId Trace detail (summary + spans)
GET /traces/:runId/evals Evals for a run
POST /traces/:runId/promote Promote a completed run to a CI eval
POST /feedback Submit feedback
GET /feedback List feedback
GET /feedback/stats Feedback aggregation
GET /satisfaction Satisfaction scores
GET /evals/stats Eval statistics
POST /experiments Create experiment
GET /experiments List experiments
GET /experiments/:id Get experiment detail
PUT /experiments/:id Update experiment
POST /experiments/:id/results Compute results
GET /experiments/:id/results Get results

All endpoints support ?since=N (ms timestamp) and ?limit=N query params.

PostHog LLM analytics

When POSTHOG_API_KEY is set, every instrumented run is mirrored to PostHog's LLM analytics as a full trace tree:

Event Emitted per Carries
$ai_trace agent run Run name, totals, terminal state, run-level failure
$ai_generation model round-trip Prompt, answer, model, tokens, cost, latency, stop reason
$ai_span tool call Tool name, duration, failure class

Nodes share the run id as $ai_trace_id, so PostHog renders one tree per run, with each tool span under the generation that requested it. A run's tool definitions are never sent: they are the same catalogue on every call, and a call is identified by its name in $ai_output_choices and by its own span.

$ai_is_error means a different thing at each level, and never travels without $ai_error and $ai_error_type. On a generation it is the model call itself failing — a provider error, a dropped stream. On a span it is a tool that crashed or returned an error. On the trace it is the run: a step budget, a timeout, a no-progress cut-off, or a tool that stopped the run. A run that failed after the model answered leaves its generations green, because they did not fail. $ai_session_id is the conversation thread; the browser session is sent separately as $session_id so a trace links to its session replay.

The browser session id travels on the X-Agent-Native-Session-Id header. Both the agent chat and the action client (useActionQuery / useActionMutation / callAction) send it, and the server exposes it to actions and agent runs alike as RequestContext.browserSessionId — so an action the UI calls and an action the agent calls during the same visit correlate. The id rotates after 30 idle minutes by default; to correlate a workflow under an id of your own, pin it once at startup:

import { setAnalyticsSessionId } from "@agent-native/core/client/analytics";

// Pinned ids opt out of idle rotation. Call clearAnalyticsSessionId() to
// return to the rotating default.
setAnalyticsSessionId(myWorkflowId);

Tool calls always appear (PostHog derives its tool tags from them), but their arguments require captureToolArgs and message content requires capturePrompts. When those are off the fields are omitted entirely rather than sent empty, so an unrecorded prompt is never mistaken for an empty one.

Run failure messages are omitted from Analytics and Monitoring, even when content capture is enabled. $ai_error carries a terminal code, named cause, retryability, and fixed message derived from the code. Tracking omits error_message, agent_run_terminal.error_detail, and failed tool result text. The terminal event retains error_code and error_cause. Local agent_runs.error_detail and agent_trace_spans.error_message remain available for owner-scoped debugging in the app's own database.

Run and Builder gateway exceptions retain their type, code, and stack frames, with a fixed Internal Server Error message and no original message header or nested cause text. Monitoring groups these issues by app, type, frame, error code, and failure class, rather than names in the original message. Exceptions retain $ai_trace_id so an issue links to the run's trace. Optional host-owned OpenTelemetry run exports also omit failure text and retain agent.error_code and agent.error_cause. Other exception capture keeps its existing behavior.

Feedback in PostHog

Thumbs, category, and free-text feedback all emit $ai_feedback. PostHog's LLM analytics feedback view reads a survey sent event instead, so set POSTHOG_AI_FEEDBACK_SURVEY_ID to the survey PostHog creates from a trace's Feedback tab. That one id is the whole configuration; without it no survey event is sent.

A thumbs vote answers the survey's first question as PostHog's choice index (1 up, 2 down) and the free text after a thumbs-down answers the follow-up question, so the text never lands where the rating belongs. Both carry one submission id per rated message, which is how PostHog joins them into a single response, and a thumbs-down stays marked incomplete until its follow-up arrives.

Export to external platforms

Send traces to Langfuse, Datadog, Grafana, or any OTel-compatible backend by registering an OpenTelemetry provider in your own app — see OpenTelemetry spans below. The framework emits gen_ai.* semantic convention spans compatible with the OpenTelemetry GenAI spec, and your provider owns the exporter and its credentials.

Core deliberately ships no exporter and holds no export endpoint or token, so a backend credential lives in your provider wiring (and the vault), never in framework configuration.

OpenTelemetry spans

The agent loop emits live OpenTelemetry spans for every run, model call, and tool call — so a host that already runs an OTel collector sees agent activity alongside the rest of its distributed traces.

This layer is optional and no-op by default:

  • @opentelemetry/api is an optional dependency. If it isn't installed, the helpers degrade to silent no-ops — nothing here ever throws into the agent loop.
  • Even when the api package is present, it ships a default no-op tracer. Spans only become real once the host registers a TracerProvider (via @opentelemetry/sdk-node or similar). The framework deliberately does not depend on the heavy SDK/exporter packages or register a provider itself — instrumentation is opt-in by the embedding app.

So the cost when you haven't wired OTel is a couple of cached property reads per call. To turn it on, install the api package plus your SDK and register a provider at server startup the same way you would for any other Node service.

The agent loop emits three span kinds:

Span When Attributes
invoke_agent once per agent run gen_ai.provider.name (the engine), gen_ai.conversation.id (the thread), gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, agent.run_id
chat {model} per model call gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, gen_ai.response.finish_reasons, gen_ai.usage.*
execute_tool {tool} once per action invocation gen_ai.tool.name, gen_ai.tool.call.id, plus success/error status

Spans are finished with OK/ERROR status and record a fixed code-derived message on failure. Zero/sentinel attribute values are pruned so spans aren't cluttered with noise. This OTel layer is purely additive to the in-house agent_trace_spans / agent_trace_summaries tables that power the dashboard above — both are produced from the same run events.

Metrics

Core also records metrics once the host registers a meter provider with registerObservabilityProvider() from @agent-native/core/server. Metrics are recorded for every request and run, with no sampling:

Instrument Attributes
http.server.request.duration http.request.method, http.response.status_code, http.route, error.type
agent_native.http.server.handoff.duration http.request.method, http.response.status_code, http.route, error.type
gen_ai.client.operation.duration gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, error.type
gen_ai.client.token.usage gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.token.type
agent_native.agent.runs status, terminal_reason, gen_ai.provider.name, gen_ai.request.model
agent_native.tool.calls gen_ai.tool.name, error.type
agent_native.telemetry.flush_failures error.type, agent_native.telemetry.signal

agent_native.http.server.handoff.duration has the same attributes and buckets as http.server.request.duration, but it ends when the response hook hands the response back to the runtime instead of when the hook starts. When its percentiles pull away from http.server.request.duration's, finished responses are being held, for example by an inline telemetry flush. The gap is a signal, not any one request's wait, since the two percentiles can come from different requests. Without waitUntil, core records the sample after that flush, so it reaches the back end with the next one.

On metrics, gen_ai.tool.name is a framework action's name or other, and gen_ai.request.model is the model id only when the engine lists it as supported, otherwise _OTHER. App-defined tools and caller-supplied model ids therefore cannot create unbounded series; spans keep the exact values.

http.route is set on every request and comes from a closed set: a framework endpoint's route template (/_agent-native/auth/session, /_agent-native/agent-chat/runs/:runId/events), an action's declared template, the app's own Nitro file route (/api/clips/:id), or a bucket. The buckets are /_agent-native/*, /mcp/*, and /.well-known/* for an unlisted path in those namespaces, /api/* for any other API path, static for file requests, page for any other GET or HEAD, and other. A raw request path is never recorded. The http.server span carries the same value.

Connecting an OpenTelemetry back end

Core ships no OpenTelemetry SDK or exporter. An app builds its own SDK providers and hands them to core with registerObservabilityProvider() from @agent-native/core/server. Core then emits the spans and metrics above through them, and the app's exporter sends them to any OTLP back end.

registerObservabilityProvider()

registerObservabilityProvider() takes one object with two optional fields:

  • tracerProvider: an object with getTracer(name, version?). Core takes its tracer from here. Without one, spans use the global @opentelemetry/api tracer.
  • meterProvider: an object with getMeter(name, version?). Metrics are recorded only when a meter provider is registered. They have no global fallback, so without one every metric is a no-op.

Either provider may also implement forceFlush(), which core calls after each response (see Flushing on serverless). The SDK's BasicTracerProvider and MeterProvider already match these shapes.

function registerObservabilityProvider(
  provider: ObservabilityProvider,
): () => void;

interface ObservabilityProvider {
  tracerProvider?: ObservabilityTracerProvider;
  meterProvider?: ObservabilityMeterProvider;
}

interface ObservabilityTracerProvider {
  getTracer(name: string, version?: string): unknown;
  forceFlush?(): Promise<void>;
}

interface ObservabilityMeterProvider {
  getMeter(name: string, version?: string): unknown;
  forceFlush?(): Promise<void>;
}

A later call replaces the earlier provider. The returned function unregisters the provider if it is still the registered one. Spans also need @opentelemetry/api installed in the app. Without it, core turns tracing off and logs one warning.

Flushing on serverless

Serverless functions freeze between invocations, so an SDK's export timer may never fire. When a registered provider implements forceFlush(), core calls it from the response hook after every request and waits for it. Each flush races one shared 2-second timeout, so telemetry never holds a response longer than that. When the platform provides waitUntil (Cloudflare, or the Netlify invocation context), core hands the flush to it and sends the response first instead of waiting.

Core calls a meter provider's forceFlush() at most once every 10 seconds per process and skips it on the requests in between. Each metric export re-sends every series the process holds, so this keeps upload volume from growing with request rate. Skipped points are not lost: the next export carries them, unless the instance never serves another request. A tracer provider is flushed on every request.

When a flush fails or times out, core stops waiting for it and adds 1 to agent_native.telemetry.flush_failures, which the next flush exports. A timed-out export keeps running and can still arrive, unless the runtime freezes the process first. The failure's error.type is timeout, suspended when the runtime froze the process mid-flush, or the error's name. Because a collector that keeps timing out never receives that counter, core also writes one agent-native.telemetry_flush_failed line to the function log per signal and error type in each process.

On a long-running server, a flush after every response means one export per request. Register providers without forceFlush() there and let the SDK's own timers export. isServerlessRuntime() from @agent-native/core/server tells the two apart. It is true on Netlify, Vercel, AWS Lambda, and Cloudflare.

Example: OTLP over HTTP

This server plugin wires the OpenTelemetry JS SDK to core. Install the SDK packages first:

pnpm add @opentelemetry/api @opentelemetry/context-async-hooks @opentelemetry/exporter-metrics-otlp-proto @opentelemetry/exporter-trace-otlp-proto @opentelemetry/resources @opentelemetry/sdk-metrics @opentelemetry/sdk-trace-base
server/plugins/otel.ts
import { randomUUID } from "node:crypto";

import {
  defineNitroPlugin,
  isServerlessRuntime,
  registerObservabilityProvider,
} from "@agent-native/core/server";
import { context } from "@opentelemetry/api";
import { AsyncLocalStorageContextManager } from "@opentelemetry/context-async-hooks";
import { OTLPMetricExporter } from "@opentelemetry/exporter-metrics-otlp-proto";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-proto";
import {
  defaultResource,
  detectResources,
  envDetector,
  resourceFromAttributes,
} from "@opentelemetry/resources";
import {
  MeterProvider,
  PeriodicExportingMetricReader,
} from "@opentelemetry/sdk-metrics";
import {
  BasicTracerProvider,
  BatchSpanProcessor,
} from "@opentelemetry/sdk-trace-base";

export default defineNitroPlugin(() => {
  if (!process.env.OTEL_EXPORTER_OTLP_ENDPOINT) return;

  // Metric counters are cumulative and restart with each process, so each
  // process needs its own series.
  const resource = defaultResource()
    .merge(resourceFromAttributes({ "service.instance.id": randomUUID() }))
    .merge(detectResources({ detectors: [envDetector] }));

  const meterProvider = new MeterProvider({
    resource,
    readers: [
      new PeriodicExportingMetricReader({
        exporter: new OTLPMetricExporter(),
        exportIntervalMillis: 60_000,
      }),
    ],
  });

  context.setGlobalContextManager(
    new AsyncLocalStorageContextManager().enable(),
  );
  const tracerProvider = new BasicTracerProvider({
    resource,
    spanProcessors: [new BatchSpanProcessor(new OTLPTraceExporter())],
  });

  if (isServerlessRuntime()) {
    registerObservabilityProvider({ meterProvider, tracerProvider });
    return;
  }
  // Without forceFlush(), core leaves exporting to the SDK's own timers.
  registerObservabilityProvider({
    meterProvider: {
      getMeter: (name, version) => meterProvider.getMeter(name, version),
    },
    tracerProvider: {
      getTracer: (name, version) => tracerProvider.getTracer(name, version),
    },
  });
});

The plugin does nothing until OTEL_EXPORTER_OTLP_ENDPOINT is set. The SDK reads the rest of its configuration from the standard OpenTelemetry environment variables:

Variable Effect
OTEL_EXPORTER_OTLP_ENDPOINT Base URL of the OTLP receiver. The exporters append /v1/traces and /v1/metrics.
OTEL_EXPORTER_OTLP_HEADERS Headers sent with every export, e.g. Authorization=Bearer%20<token>
OTEL_SERVICE_NAME service.name on every span and metric
OTEL_RESOURCE_ATTRIBUTES Extra resource attributes, e.g. deployment.environment.name=production
OTEL_TRACES_SAMPLER / OTEL_TRACES_SAMPLER_ARG Trace sampling, e.g. parentbased_traceidratio and 0.1. Metrics are never sampled.

The -proto exporters send OTLP over HTTP with protobuf, so point the endpoint at the receiver's HTTP port (4318 by default), not its gRPC port. Spans carry no prompt or completion text. The capture* settings under Configuration apply to the SQL traces and PostHog, not to OTLP.

The repository's packages/otel holds a fuller version of this plugin with per-signal endpoints and exporter validation. It is not published to npm, so start from the plugin above instead of installing it.

Error reporting (Sentry or PostHog)

Server-side errors that escape Nitro route handlers are reported to Sentry when a DSN is configured. Without it the SDK silently no-ops, so it's safe to leave the env vars unset in dev. Browser and server events can go to the same Sentry project; split them into separate projects only when you want operational separation for ownership, volume, quotas, or alert routing.

Sentry is not required. Error reporting fans out to every configured backend, so setting only POSTHOG_API_KEY gives full server-side error tracking with no Sentry project at all — see Tracking. Configure both and each error goes to both. The noise-filtering rules below are shared by every backend rather than being Sentry-specific.

Surface SDK Env var Notes
Browser / SPA @sentry/browser VITE_SENTRY_CLIENT_DSN, SENTRY_CLIENT_DSN, or SENTRY_DSN Captures unhandled errors and route-change breadcrumbs in the client.
Nitro server @sentry/node SENTRY_SERVER_DSN or SENTRY_DSN Captures 5xx responses and Nitro lifecycle errors. Per-request user.
agent-native CLI @sentry/node hardcoded Crash reports from the published CLI binary; not user-configurable.

Server-side configuration

Set SENTRY_SERVER_DSN or the shared SENTRY_DSN in the deploy environment (Netlify dashboard, Cloudflare secrets, etc.). The framework auto-mounts a Nitro plugin that:

  1. Calls Sentry.init once at startup (idempotent — safe to call from multiple plugins).
  2. Resolves the user via getSession(event) on every API/framework request and attaches id / email / username plus an orgId tag to Sentry's per-request isolation scope. Static-asset paths are skipped to avoid extra DB hits.
  3. Captures every framework-route 5xx with searchable route, method, and userAgent tags.

Optional knobs:

  • SENTRY_SERVER_TRACES_SAMPLE_RATE (float 0–1) — opt in to performance tracing. Defaults to 0 (errors only). Invalid values clamp to 0.
  • AGENT_NATIVE_RELEASE — overrides the release tag. Defaults to agent-native-server@<core-version>.

Templates

Every template inherits this automatically — there's nothing to import. For SSR apps, the server injects a tiny browser config script when SENTRY_CLIENT_DSN, VITE_SENTRY_CLIENT_DSN, or shared SENTRY_DSN is available at runtime, so browser capture is not limited to Vite build-time env. Templates that want custom behavior (extra tags, different DSN per template, hard-disable Sentry) can override by exporting their own plugin from server/plugins/sentry.ts:

server/plugins/sentry.ts
import { createSentryPlugin } from "@agent-native/core/server";
export default createSentryPlugin();

The CLI's hardcoded DSN is intentional — the published binary needs to phone home crashes regardless of which environment runs it. The server module never hardcodes a DSN because it runs inside customer environments where operators decide whether errors should reach Sentry at all.

Privacy & PII

Both server and CLI initialize with sendDefaultPii: false and a beforeSend hook that strips:

  • request.headers.authorization, cookie, set-cookie, proxy-authorization
  • request.cookies
  • user.ip_address (auto-collected without consent)
  • contexts.runtime_env (process env snapshot)
  • Any event whose top-level exception type is ValidationError (treated as expected user-input rejection, not a bug).

Identity fields explicitly set via setUser({ id, email, username }) are preserved.

What's next

  • Tracking — product analytics (PostHog, Mixpanel, Amplitude) for your app's own events
  • Actions — the operations that appear as tool calls in traces
  • Security — data scoping and credential handling
  • Toolkit piece: Observability Kit — the dashboard and feedback UI built on this data model