Agent Observability
Every agent-native app gets observability out of the box. Traces, automated evals, user feedback, and A/B experiments work with zero configuration — all data lives in the app's own SQL database.
This page covers agent quality metrics: traces, cost, evals, and feedback stored in your database. For product analytics (your app's events flowing to PostHog/Mixpanel/Amplitude), see Tracking.
Three things called "evals"/"observability" — which do I want?
These three pages are easy to confuse. Pick by the question you're asking:
| Page | The question it answers | When it runs | Concern |
|---|---|---|---|
| Observability evals (this page, the Evals tab) | "How did my real production runs do?" | Passive, after every run (LLM-judge sampled) | Quality |
CI Eval Gate (*.eval.ts) |
"Does the agent do the right thing on this fixed input?" | Active, deterministic, a CI/deploy gate | Quality |
| Observational Memory | "Is this long thread staying cheap and inside the window?" | Background compaction on long threads | Cost / context |
Observability and the CI Eval Gate both score quality but from opposite ends — passive post-hoc scoring of real traffic vs. active pass/fail checks on fixed inputs. Promote a completed production trace into a *.eval.ts case (dashboard Promote to eval, promote-trace-eval, or agent-native eval promote <runId> --write evals/from-trace.eval.ts) so the miss becomes the next CI gate. See From a production trace. Observational Memory is unrelated to quality; it's about token cost and context-window pressure.
What's captured automatically
When a user sends a message, the framework automatically records:
- Token usage — input, output, cache read, cache write
- Cost — computed from token counts and model pricing
- Latency — total duration and time per tool call
- Tool calls — which actions were invoked, success/error status, duration
- Automated evals — 5 quality scores computed after every run
No code changes needed. The instrumentation hooks into production-agent.ts transparently.
production-agent.ts
scoped to the signed-in user
One agent run produces a trace, automated scores, and a feedback hook — all stored in the app's own SQL and surfaced on the dashboard. Experiments split traffic across config variants.
The dashboard
Add the dashboard to any template with a single route:
import { ObservabilityDashboard } from "@agent-native/toolkit/app/observability";
export default function ObservabilityPage() {
return (
<div className="min-h-screen bg-background p-6">
<ObservabilityDashboard />
</div>
);
}All data is scoped to the signed-in user; there is no cross-user admin view today.
The dashboard has 6 tabs:
| Tab | What it shows |
|---|---|
| Overview | Key metrics — runs, cost, latency, tool success rate, satisfaction, eval score |
| Conversations | Trace list with drill-down to individual spans (agent_run, llm_call, tool_call) |
| Evals | Automated eval scores by criteria, trends over time |
| Experiments | A/B test list with status badges, variant results with confidence intervals |
| Feedback | Thumbs up/down stream, category breakdown, frustration scores |
| Human review | Review the ask and answer, record feedback, and draft instruction updates |
Conversations
Select a run to inspect its spans. Expand any span to view captured tool inputs, outputs, errors, and metadata. Use Open full conversation to jump to the associated chat thread.
Tool arguments, results, and full error bodies are off by default. Enable captureToolArgs to store arguments and captureToolResults to store results and full error bodies in the observability config. Sensitive fields are redacted before storage. A failed tool always keeps a bounded first line of its error, with credentials, emails, and long opaque ids redacted, so failures stay diagnosable even when captureToolResults is off.
Prompt capture is off by default. Set capturePrompts: true in the observability config, or set AGENT_NATIVE_OBSERVABILITY_CAPTURE_PROMPTS=true, to store each llm_call span's input and output for this detail view. The system prompt is excluded, structured sensitive fields are redacted, and each side is capped at 128 KiB. Existing spans are not backfilled.
Human review
The Human review tab shows each run’s ask and answer, lets you record feedback, and lets you draft an instruction update for a human to review. Saving the draft does not change instructions automatically.
User feedback
Explicit feedback
Thumbs up/down buttons render inline on every agent message in the chat UI. Thumbs down opens a category popover (Inaccurate, Not helpful, Wrong tool, Too slow). This is wired into AssistantChat.tsx automatically.
Implicit feedback (frustration index)
The framework computes a Frustration Index (0-100) from conversation signals:
| Signal | Weight | What it detects |
|---|---|---|
| Rephrasing | 30% | User repeats similar messages |
| Retry patterns | 20% | "Try again", "no that's wrong" |
| Abandonment | 20% | Session ends shortly after response |
| Sentiment | 15% | Negative language patterns |
| Length trend | 15% | Declining message lengths |
Score interpretation: 0-20 = healthy, 20-40 = friction, 40-60 = dissatisfied, 60+ = broken session.
Automated evals
Five deterministic scorers run after every agent run:
| Criteria | What it measures | Score range |
|---|---|---|
tool_success_rate |
% of tool calls without errors | 0-1 |
step_efficiency |
Penalizes excessive LLM iterations for tool-using runs | 0-1 |
latency_score |
Normalized against 10s/tool baseline | 0-1 |
cost_efficiency |
Normalized against cost baseline | 0-1 |
error_recovery |
Did the agent recover from tool errors? | 0 or 1 |
LLM-as-judge (optional)
Enable sampled LLM-based evaluation by setting evalSampleRate:
// server/plugins/config.ts
import { defineAppConfig } from "@agent-native/core/server";
export default defineAppConfig({
observability: {
enabled: true,
evalSampleRate: 0.05, // 5%
},
});Custom criteria use natural language rubrics:
const criteria = {
name: "helpfulness",
description: "Was the response helpful and complete?",
rubric: "0.0 = unhelpful, 0.5 = partially helpful, 1.0 = fully resolved",
};A/B experiments
Test different models, temperatures, or agent configurations:
// Create via API
POST /_agent-native/observability/experiments
{
"name": "model-a-vs-b",
"variants": [
{ "id": "control", "weight": 50, "config": { "model": "<your-model-id>" } },
{ "id": "treatment", "weight": 50, "config": { "model": "<other-model-id>" } }
],
"metrics": ["cost", "latency", "satisfaction"]
}
// Start the experiment
PUT /_agent-native/observability/experiments/:id
{ "status": "running" }Use the real model identifiers your engine accepts in place of <your-model-id> / <other-model-id> (model names change often — check your provider/engine for the current ids). The agent loop automatically resolves the user's variant and applies the config override. Assignment uses consistent hashing — same user always gets the same variant.
consistent hash
cost · latency · satisfaction
Each user hashes to a stable variant, the loop applies that variant's config override, and results roll up per variant with confidence intervals.
Configuration
Set from a server plugin with defineAppConfig(). Every field also has a deployment environment variable alias — see the declared configuration reference.
// server/plugins/config.ts
import { defineAppConfig } from "@agent-native/core/server";
export default defineAppConfig({
observability: {
enabled: true, // Master switch
capturePrompts: false, // Store prompt content in traces
captureToolArgs: false, // Store action input arguments
captureToolResults: false, // Store action results
captureLlmSpans: true, // Emit one $ai_span per tool call to PostHog
evalSampleRate: 0, // 0-1, fraction of runs to LLM-judge
},
});Content is redacted by default: token counts, costs, and timing are stored, plus, for a failed tool, a bounded first line of its error with credentials, emails, and long opaque ids redacted. Prompts, tool inputs, tool results, and full error bodies stay off unless capturePrompts, captureToolArgs, and captureToolResults are on. Turn them on only when you need prompt or argument content for debugging.
API endpoints
All auto-mounted at /_agent-native/observability/:
| Method | Path | Purpose |
|---|---|---|
| GET | / |
Overview stats |
| GET | /traces |
List trace summaries |
| GET | /traces/:runId |
Trace detail (summary + spans) |
| GET | /traces/:runId/evals |
Evals for a run |
| POST | /traces/:runId/promote |
Promote a completed run to a CI eval |
| POST | /feedback |
Submit feedback |
| GET | /feedback |
List feedback |
| GET | /feedback/stats |
Feedback aggregation |
| GET | /satisfaction |
Satisfaction scores |
| GET | /evals/stats |
Eval statistics |
| POST | /experiments |
Create experiment |
| GET | /experiments |
List experiments |
| GET | /experiments/:id |
Get experiment detail |
| PUT | /experiments/:id |
Update experiment |
| POST | /experiments/:id/results |
Compute results |
| GET | /experiments/:id/results |
Get results |
All endpoints support ?since=N (ms timestamp) and ?limit=N query params.
PostHog LLM analytics
When POSTHOG_API_KEY is set, every instrumented run is mirrored to PostHog's
LLM analytics as a full trace tree:
| Event | Emitted per | Carries |
|---|---|---|
$ai_trace |
agent run | Run name, totals, terminal state, run-level failure |
$ai_generation |
model round-trip | Prompt, answer, model, tokens, cost, latency, stop reason |
$ai_span |
tool call | Tool name, duration, failure class |
Nodes share the run id as $ai_trace_id, so PostHog renders one tree per run,
with each tool span under the generation that requested it. A run's tool
definitions are never sent: they are the same catalogue on every call, and a
call is identified by its name in $ai_output_choices and by its own span.
$ai_is_error means a different thing at each level, and never travels without
$ai_error and $ai_error_type. On a generation it is the model call itself
failing — a provider error, a dropped stream. On a span it is a tool that
crashed or returned an error. On the trace it is the run: a step budget, a
timeout, a no-progress cut-off, or a tool that stopped the run. A run that
failed after the model answered leaves its generations green, because they did
not fail.
$ai_session_id is the conversation thread; the browser session is sent
separately as $session_id so a trace links to its session replay.
The browser session id travels on the X-Agent-Native-Session-Id header. Both
the agent chat and the action client (useActionQuery / useActionMutation /
callAction) send it, and the server exposes it to actions and agent runs alike
as RequestContext.browserSessionId — so an action the UI calls and an action
the agent calls during the same visit correlate. The id rotates after 30 idle
minutes by default; to correlate a workflow under an id of your own, pin it once
at startup:
import { setAnalyticsSessionId } from "@agent-native/core/client/analytics";
// Pinned ids opt out of idle rotation. Call clearAnalyticsSessionId() to
// return to the rotating default.
setAnalyticsSessionId(myWorkflowId);Tool calls always appear (PostHog derives its tool tags from them), but their
arguments require captureToolArgs and message content requires
capturePrompts. When those are off the fields are omitted entirely rather than
sent empty, so an unrecorded prompt is never mistaken for an empty one.
Run failure messages are omitted from Analytics and Monitoring, even when
content capture is enabled. $ai_error carries a terminal code, named cause,
retryability, and fixed message derived from the code. Tracking omits
error_message, agent_run_terminal.error_detail, and failed tool result text.
The terminal event retains error_code and error_cause. Local
agent_runs.error_detail and agent_trace_spans.error_message remain available
for owner-scoped debugging in the app's own database.
Run and Builder gateway exceptions retain their type, code, and stack frames,
with a fixed Internal Server Error message and no original message header or nested
cause text. Monitoring groups these issues by app, type, frame, error code, and
failure class, rather than names in the original message. Exceptions retain
$ai_trace_id so an issue links to the run's trace. Optional host-owned
OpenTelemetry run exports also omit failure text and retain agent.error_code
and agent.error_cause. Other exception capture keeps its existing behavior.
Feedback in PostHog
Thumbs, category, and free-text feedback all emit $ai_feedback. PostHog's LLM
analytics feedback view reads a survey sent event instead, so set
POSTHOG_AI_FEEDBACK_SURVEY_ID to the survey PostHog creates from a trace's
Feedback tab. That one id is the whole configuration; without it no survey
event is sent.
A thumbs vote answers the survey's first question as PostHog's choice index
(1 up, 2 down) and the free text after a thumbs-down answers the follow-up
question, so the text never lands where the rating belongs. Both carry one
submission id per rated message, which is how PostHog joins them into a single
response, and a thumbs-down stays marked incomplete until its follow-up arrives.
Export to external platforms
Send traces to Langfuse, Datadog, Grafana, or any OTel-compatible backend by registering an OpenTelemetry provider in your own app — see OpenTelemetry spans below. The framework emits gen_ai.* semantic convention spans compatible with the OpenTelemetry GenAI spec, and your provider owns the exporter and its credentials.
Core deliberately ships no exporter and holds no export endpoint or token, so a backend credential lives in your provider wiring (and the vault), never in framework configuration.
OpenTelemetry spans
The agent loop emits live OpenTelemetry spans for every run, model call, and tool call — so a host that already runs an OTel collector sees agent activity alongside the rest of its distributed traces.
This layer is optional and no-op by default:
@opentelemetry/apiis an optional dependency. If it isn't installed, the helpers degrade to silent no-ops — nothing here ever throws into the agent loop.- Even when the api package is present, it ships a default no-op tracer. Spans only become real once the host registers a
TracerProvider(via@opentelemetry/sdk-nodeor similar). The framework deliberately does not depend on the heavy SDK/exporter packages or register a provider itself — instrumentation is opt-in by the embedding app.
So the cost when you haven't wired OTel is a couple of cached property reads per call. To turn it on, install the api package plus your SDK and register a provider at server startup the same way you would for any other Node service.
The agent loop emits three span kinds:
| Span | When | Attributes |
|---|---|---|
invoke_agent |
once per agent run | gen_ai.provider.name (the engine), gen_ai.conversation.id (the thread), gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, agent.run_id |
chat {model} |
per model call | gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, gen_ai.response.finish_reasons, gen_ai.usage.* |
execute_tool {tool} |
once per action invocation | gen_ai.tool.name, gen_ai.tool.call.id, plus success/error status |
Spans are finished with OK/ERROR status and record a fixed code-derived message on failure. Zero/sentinel attribute values are pruned so spans aren't cluttered with noise. This OTel layer is purely additive to the in-house agent_trace_spans / agent_trace_summaries tables that power the dashboard above — both are produced from the same run events.
Metrics
Core also records metrics once the host registers a meter provider with registerObservabilityProvider() from @agent-native/core/server. Metrics are recorded for every request and run, with no sampling:
| Instrument | Attributes |
|---|---|
http.server.request.duration |
http.request.method, http.response.status_code, http.route, error.type |
agent_native.http.server.handoff.duration |
http.request.method, http.response.status_code, http.route, error.type |
gen_ai.client.operation.duration |
gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, error.type |
gen_ai.client.token.usage |
gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.token.type |
agent_native.agent.runs |
status, terminal_reason, gen_ai.provider.name, gen_ai.request.model |
agent_native.tool.calls |
gen_ai.tool.name, error.type |
agent_native.telemetry.flush_failures |
error.type, agent_native.telemetry.signal |
agent_native.http.server.handoff.duration has the same attributes and buckets as http.server.request.duration, but it ends when the response hook hands the response back to the runtime instead of when the hook starts. When its percentiles pull away from http.server.request.duration's, finished responses are being held, for example by an inline telemetry flush. The gap is a signal, not any one request's wait, since the two percentiles can come from different requests. Without waitUntil, core records the sample after that flush, so it reaches the back end with the next one.
On metrics, gen_ai.tool.name is a framework action's name or other, and gen_ai.request.model is the model id only when the engine lists it as supported, otherwise _OTHER. App-defined tools and caller-supplied model ids therefore cannot create unbounded series; spans keep the exact values.
http.route is set on every request and comes from a closed set: a framework endpoint's route template (/_agent-native/auth/session, /_agent-native/agent-chat/runs/:runId/events), an action's declared template, the app's own Nitro file route (/api/clips/:id), or a bucket. The buckets are /_agent-native/*, /mcp/*, and /.well-known/* for an unlisted path in those namespaces, /api/* for any other API path, static for file requests, page for any other GET or HEAD, and other. A raw request path is never recorded. The http.server span carries the same value.
Connecting an OpenTelemetry back end
Core ships no OpenTelemetry SDK or exporter. An app builds its own SDK providers and hands them to core with registerObservabilityProvider() from @agent-native/core/server. Core then emits the spans and metrics above through them, and the app's exporter sends them to any OTLP back end.
registerObservabilityProvider()
registerObservabilityProvider() takes one object with two optional fields:
- tracerProvider: an object with
getTracer(name, version?). Core takes its tracer from here. Without one, spans use the global@opentelemetry/apitracer. - meterProvider: an object with
getMeter(name, version?). Metrics are recorded only when a meter provider is registered. They have no global fallback, so without one every metric is a no-op.
Either provider may also implement forceFlush(), which core calls after each response (see Flushing on serverless). The SDK's BasicTracerProvider and MeterProvider already match these shapes.
function registerObservabilityProvider(
provider: ObservabilityProvider,
): () => void;
interface ObservabilityProvider {
tracerProvider?: ObservabilityTracerProvider;
meterProvider?: ObservabilityMeterProvider;
}
interface ObservabilityTracerProvider {
getTracer(name: string, version?: string): unknown;
forceFlush?(): Promise<void>;
}
interface ObservabilityMeterProvider {
getMeter(name: string, version?: string): unknown;
forceFlush?(): Promise<void>;
}A later call replaces the earlier provider. The returned function unregisters the provider if it is still the registered one. Spans also need @opentelemetry/api installed in the app. Without it, core turns tracing off and logs one warning.
Flushing on serverless
Serverless functions freeze between invocations, so an SDK's export timer may never fire. When a registered provider implements forceFlush(), core calls it from the response hook after every request and waits for it. Each flush races one shared 2-second timeout, so telemetry never holds a response longer than that. When the platform provides waitUntil (Cloudflare, or the Netlify invocation context), core hands the flush to it and sends the response first instead of waiting.
Core calls a meter provider's forceFlush() at most once every 10 seconds per process and skips it on the requests in between. Each metric export re-sends every series the process holds, so this keeps upload volume from growing with request rate. Skipped points are not lost: the next export carries them, unless the instance never serves another request. A tracer provider is flushed on every request.
When a flush fails or times out, core stops waiting for it and adds 1 to agent_native.telemetry.flush_failures, which the next flush exports. A timed-out export keeps running and can still arrive, unless the runtime freezes the process first. The failure's error.type is timeout, suspended when the runtime froze the process mid-flush, or the error's name. Because a collector that keeps timing out never receives that counter, core also writes one agent-native.telemetry_flush_failed line to the function log per signal and error type in each process.
On a long-running server, a flush after every response means one export per request. Register providers without forceFlush() there and let the SDK's own timers export. isServerlessRuntime() from @agent-native/core/server tells the two apart. It is true on Netlify, Vercel, AWS Lambda, and Cloudflare.
Example: OTLP over HTTP
This server plugin wires the OpenTelemetry JS SDK to core. Install the SDK packages first:
pnpm add @opentelemetry/api @opentelemetry/context-async-hooks @opentelemetry/exporter-metrics-otlp-proto @opentelemetry/exporter-trace-otlp-proto @opentelemetry/resources @opentelemetry/sdk-metrics @opentelemetry/sdk-trace-baseimport { randomUUID } from "node:crypto";
import {
defineNitroPlugin,
isServerlessRuntime,
registerObservabilityProvider,
} from "@agent-native/core/server";
import { context } from "@opentelemetry/api";
import { AsyncLocalStorageContextManager } from "@opentelemetry/context-async-hooks";
import { OTLPMetricExporter } from "@opentelemetry/exporter-metrics-otlp-proto";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-proto";
import {
defaultResource,
detectResources,
envDetector,
resourceFromAttributes,
} from "@opentelemetry/resources";
import {
MeterProvider,
PeriodicExportingMetricReader,
} from "@opentelemetry/sdk-metrics";
import {
BasicTracerProvider,
BatchSpanProcessor,
} from "@opentelemetry/sdk-trace-base";
export default defineNitroPlugin(() => {
if (!process.env.OTEL_EXPORTER_OTLP_ENDPOINT) return;
// Metric counters are cumulative and restart with each process, so each
// process needs its own series.
const resource = defaultResource()
.merge(resourceFromAttributes({ "service.instance.id": randomUUID() }))
.merge(detectResources({ detectors: [envDetector] }));
const meterProvider = new MeterProvider({
resource,
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter(),
exportIntervalMillis: 60_000,
}),
],
});
context.setGlobalContextManager(
new AsyncLocalStorageContextManager().enable(),
);
const tracerProvider = new BasicTracerProvider({
resource,
spanProcessors: [new BatchSpanProcessor(new OTLPTraceExporter())],
});
if (isServerlessRuntime()) {
registerObservabilityProvider({ meterProvider, tracerProvider });
return;
}
// Without forceFlush(), core leaves exporting to the SDK's own timers.
registerObservabilityProvider({
meterProvider: {
getMeter: (name, version) => meterProvider.getMeter(name, version),
},
tracerProvider: {
getTracer: (name, version) => tracerProvider.getTracer(name, version),
},
});
});The plugin does nothing until OTEL_EXPORTER_OTLP_ENDPOINT is set. The SDK reads the rest of its configuration from the standard OpenTelemetry environment variables:
| Variable | Effect |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
Base URL of the OTLP receiver. The exporters append /v1/traces and /v1/metrics. |
OTEL_EXPORTER_OTLP_HEADERS |
Headers sent with every export, e.g. Authorization=Bearer%20<token> |
OTEL_SERVICE_NAME |
service.name on every span and metric |
OTEL_RESOURCE_ATTRIBUTES |
Extra resource attributes, e.g. deployment.environment.name=production |
OTEL_TRACES_SAMPLER / OTEL_TRACES_SAMPLER_ARG |
Trace sampling, e.g. parentbased_traceidratio and 0.1. Metrics are never sampled. |
The -proto exporters send OTLP over HTTP with protobuf, so point the endpoint at the receiver's HTTP port (4318 by default), not its gRPC port. Spans carry no prompt or completion text. The capture* settings under Configuration apply to the SQL traces and PostHog, not to OTLP.
The repository's packages/otel holds a fuller version of this plugin with per-signal endpoints and exporter validation. It is not published to npm, so start from the plugin above instead of installing it.
Error reporting (Sentry or PostHog)
Server-side errors that escape Nitro route handlers are reported to Sentry when a DSN is configured. Without it the SDK silently no-ops, so it's safe to leave the env vars unset in dev. Browser and server events can go to the same Sentry project; split them into separate projects only when you want operational separation for ownership, volume, quotas, or alert routing.
Sentry is not required. Error reporting fans out to every configured backend, so
setting only POSTHOG_API_KEY gives full server-side error tracking with no
Sentry project at all — see Tracking.
Configure both and each error goes to both. The noise-filtering rules below are
shared by every backend rather than being Sentry-specific.
| Surface | SDK | Env var | Notes |
|---|---|---|---|
| Browser / SPA | @sentry/browser |
VITE_SENTRY_CLIENT_DSN, SENTRY_CLIENT_DSN, or SENTRY_DSN |
Captures unhandled errors and route-change breadcrumbs in the client. |
| Nitro server | @sentry/node |
SENTRY_SERVER_DSN or SENTRY_DSN |
Captures 5xx responses and Nitro lifecycle errors. Per-request user. |
agent-native CLI |
@sentry/node |
hardcoded | Crash reports from the published CLI binary; not user-configurable. |
Server-side configuration
Set SENTRY_SERVER_DSN or the shared SENTRY_DSN in the deploy environment (Netlify dashboard, Cloudflare secrets, etc.). The framework auto-mounts a Nitro plugin that:
- Calls
Sentry.initonce at startup (idempotent — safe to call from multiple plugins). - Resolves the user via
getSession(event)on every API/framework request and attachesid/email/usernameplus anorgIdtag to Sentry's per-request isolation scope. Static-asset paths are skipped to avoid extra DB hits. - Captures every framework-route 5xx with searchable
route,method, anduserAgenttags.
Optional knobs:
SENTRY_SERVER_TRACES_SAMPLE_RATE(float0–1) — opt in to performance tracing. Defaults to0(errors only). Invalid values clamp to0.AGENT_NATIVE_RELEASE— overrides thereleasetag. Defaults toagent-native-server@<core-version>.
Templates
Every template inherits this automatically — there's nothing to import. For SSR apps, the server injects a tiny browser config script when SENTRY_CLIENT_DSN, VITE_SENTRY_CLIENT_DSN, or shared SENTRY_DSN is available at runtime, so browser capture is not limited to Vite build-time env. Templates that want custom behavior (extra tags, different DSN per template, hard-disable Sentry) can override by exporting their own plugin from server/plugins/sentry.ts:
import { createSentryPlugin } from "@agent-native/core/server";
export default createSentryPlugin();The CLI's hardcoded DSN is intentional — the published binary needs to phone home crashes regardless of which environment runs it. The server module never hardcodes a DSN because it runs inside customer environments where operators decide whether errors should reach Sentry at all.
Privacy & PII
Both server and CLI initialize with sendDefaultPii: false and a beforeSend hook that strips:
request.headers.authorization,cookie,set-cookie,proxy-authorizationrequest.cookiesuser.ip_address(auto-collected without consent)contexts.runtime_env(process env snapshot)- Any event whose top-level exception type is
ValidationError(treated as expected user-input rejection, not a bug).
Identity fields explicitly set via setUser({ id, email, username }) are preserved.
What's next
- Tracking — product analytics (PostHog, Mixpanel, Amplitude) for your app's own events
- Actions — the operations that appear as tool calls in traces
- Security — data scoping and credential handling
- Toolkit piece: Observability Kit — the dashboard and feedback UI built on this data model