See what happens in production
Full visibility into every AI interaction. Every request traced, every response scored, every regression caught before your customers notice.
What is agent observability?
Traditional monitoring tracks uptime and latency. Agent observability tracks whether LLM responses are good.
Trace every step of every LLM call: the prompt, the response, tool calls, retrieved context, and the final output. Score that traffic against quality criteria automatically, and get alerted when something drifts.
Two paths to full visibility
Go from zero observability to full tracing with minimal effort. Use the gateway as your default capture layer, and the SDK for deeper context on your agent’s tool calls, planning steps, or loops.
SDK tracing
Enable auto-instrumentation once and every call to a supported provider is logged with inputs, outputs, latency, tokens, and cost. Wrap your own functions for deeper context on tool calls, planning steps, and agent loops. Logs are buffered and flushed asynchronously to keep overhead low.
Trace LLM callsGateway
Change your base URL to the gateway and pass your Braintrust API key. No other code changes needed. You get a unified API across OpenAI, Anthropic, Google, AWS, and other providers, with automatic caching and observability on every request.
Use the gatewayhttps://gateway.braintrust.dev.Everything you need to observe agents in production. From the first span to the final score.
Traces
Each request becomes a tree of nested spans: LLM calls, tool invocations, and retrieval steps. See the full chain from input to output, with latency, token counts, and cost attached to every step.
Log your first traceDashboards
Track latency, cost, quality, and volume in real time. Break down LLM spend by model, feature, or team, and build views for engineering, product, or leadership.
Explore monitoringAlerts
Set thresholds on any metric, like score drops, latency spikes, or error rate changes. Get notified before your customers do.
Set up alertsOnline scoring
The same scorers you use in offline experiments run automatically on production traffic. Every conversation is scored as it happens.
Score production trafficHow the best AI teams run in production

Paul Klein IV, Founder & CEO
“Braintrust shows us the model's input and output and what actions it took.”

Malte Ubl, CTO
“We didn't realize we needed deep observability until Braintrust.”

Allen Kleiner, AI Engineering Lead
“Loop helps us understand trace details that would be impossible to scan manually.”




