Pricing Docs Blog Resources Changelog Roadmap Community GitHub Slack X / Twitter LinkedIn
Book a demo Get started
All blogs

Launch Week #2 Day 2: Online Evaluation

Today we're launching online evaluation. With online evaluation each request is automatically evaluated. This allows you to monitor things in production.

Nov 11, 2025 5 min read
Launch Week #2 Day 2: Online Evaluation

Did your customer support agent ever mention a competitor?

That would be awful.

And if you build AI apps for health or finance, it could be worse. One wrong answer can change someone’s life.

Pre-production evals can’t catch everything. You will never know exactly how users will interact with your app. You will never see all the edge cases of real-world data until you’re in production.

Today we’re launching Online Evaluation. This feature closes the LLMOps feedback loop and solves this problem.

With Online Evaluation, every AI request gets evaluated in real time. You can spot hallucinations, off-brand answers, and subtle regressions as they happen.

With Online Evaluation, you get:

  • A live view of the reliability of your system in production
  • Confidence that your outputs meet your quality standards
  • A way to find edge cases and add them to your test cases to improve your AI system
  • Clear insight into how prompt changes behave in production

How it works:

  • Pick an evaluator (use an LLM-as-a-judge or write your own evaluator logic in Python)
  • Provide filters to target the right spans and set the sampling rate to control your cost and coverage
  • Measure changes against live traffic, spot regressions, and add them to your test set

You can set up online evaluation in a couple of minutes: Add one line to instrument your application. Then set up online evaluation with a few clicks.

Check out our docs to get started.

More from the blog

The latest updates and insights from Agenta

View all blogs
Introducing Agenta 2.0Article

Introducing Agenta 2.0

Jul 22, 2026
Building the Data Flywheel: How to Use Production Data to Improve Your LLM ApplicationArticle

Building the Data Flywheel: How to Use Production Data to Improve Your LLM Application

Dec 19, 2025
Launch Week #2 Day 5: Jinja2 Prompt TemplatesArticle

Launch Week #2 Day 5: Jinja2 Prompt Templates

Nov 14, 2025
Commercial Open Source Is Hard: Our JourneyArticle

Commercial Open Source Is Hard: Our Journey

Nov 13, 2025

Ship agents that actually work

Build with skills and tools, run on any harness in any environment, and improve with real feedback. All open source.

Start building Book a demo