LLM benchmarks, agent evaluations, and experiments.
-
ResearchPrompt caching benchmark: high cache reuse doesn’t always mean lower cost
We benchmarked prompt caching across DeepSeek, GLM, GPT, and Claude using Harbor evals and Phoenix traces, so we could compare cache reuse, estimated cost, and latency on the… Nancy Chauhan October 2, 2026 7 min read -
ResearchAre agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment. Laurie Voss September 29, 2026 9 min read -
ResearchAnthropic says it fixed Claude’s writing. I ran the evals to check.
Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed. Jim Bennett September 29, 2026 12 min read -
ResearchWhat Jev’s probabilities reveal that repeated LLM judgments miss
I tested Jev and five LLM judges across ten Arize Phoenix evaluators to measure how often their answers changed alongside accuracy, cost, and latency. Elizabeth Hutton September 28, 2026 11 min read -
ResearchJev vs. LLM-as-a-Judge: Accuracy and cost benchmarks
We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. Laurie Voss September 23, 2026 11 min read -
ResearchReal-time LLM guardrails with Jev: comparing latency and cost
Compare Jev and GPT-5.4 nano for real-time LLM guardrails, with demo results on latency, cost, and checks on agent inputs, replies, and tool calls. Jim Bennett September 23, 2026 13 min read -
ResearchWhere agent evals are going: Agent-as-a-Judge
Agents changed what failure looks like, and the evaluation layer has to change with them. Why agent-as-a-judge is moving from research paper to production eval stack. Laurie Voss August 19, 2026 8 min read -
ResearchHow to write effective AI agent skills: 6 data-backed practices
Three recent studies show what actually makes an AI agent skill effective: human expertise, compact procedures, tight routing, harness-specific testing, and eval-gated changes—not longer Markdown. Laurie Voss July 24, 2026 11 min read -
ResearchCost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models
Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. Laurie Voss July 23, 2026 16 min read -
ResearchHow to measure human-LLM judge alignment
No single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision,… Elizabeth Hutton July 22, 2026 16 min read -
ResearchWhy AI token costs don’t tell you if your AI is working
Token spend does not prove AI is creating value. Teams need cost-per-outcome metrics that connect AI usage to resolved tickets, accepted code, shipped features, and other business results. Laurie Voss June 19, 2026 9 min read -
ResearchAgent harness vs. agent framework: why harnesses are replacing frameworks
Agent harnesses are replacing frameworks as the real product surface for reliable AI agents, shifting the work from prompt tuning to loops, tools, traces, evals, and operational metrics. Laurie Voss June 18, 2026 8 min read -
ResearchPostgresFS vs. SQL skills: should AI agents fake a filesystem?
Can an AI agent use a database as if it were a filesystem? Arize compared a Postgres-backed filesystem abstraction with a SQL skill and found that locality, accuracy,… Aparna Dhinakaran Sufjan Fana June 11, 2026 12 min read -
ResearchWhat we learned testing 7 models under the same agent harness
Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on… Nancy Chauhan May 20, 2026 10 min read -
ResearchModels got an order of magnitude better at following instructions in one year
A year ago, frontier models started losing track of instructions somewhere around 200–300 simultaneous constraints. With 2026 models, that ceiling is closer to 2,000 — an order-of-magnitude jump.… Laurie Voss May 12, 2026 11 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.