LLM evals can pass while an agent still fails users. Turn production traces into human-reviewed datasets that catch regressions before the next model or prompt change ships. Nikita Kabardin joins the hosts to explain Langfuse’s trace-to-eval workflow, why agents should not control their own grading, and where human review belongs. The conversation also covers reliable pull request workflows, context for autonomous agents, and keeping customer data out of public fixes.
Summary
Building an AI agent that works in a demo takes an afternoon. Showing that it still works next month, after a model swap, a prompt change and a few thousand real sessions, is the part that eats the year. Traces alone don’t get you there. They make an agent better only when a human reads some of them, turns the good cases and the corrected failures into a golden dataset, and runs every change against it. Hand that loop to a coding agent with no human checking it and you risk a confident report of progress that never shows up for users. Nikita Kabardin, a Paris-based senior front-end engineer at Langfuse, the open source LLM observability and evaluation platform that became part of ClickHouse in January 2026, sees this from the interface side every day. Before Langfuse he spent about 20 years building interfaces at EA/DICE, Spotify and Warner Bros. Discovery. We first met him in Stockholm in 2015.
The same problem runs through the rest of the stack. Loops built only from prompts lose track of pull requests. Markdown memory stops working once an agent runs with nobody watching. And if a product improves itself from what it sees in customer environments, two questions come up before any public pull request: are you allowed to use that data at all, and how do you stop it leaking?
Agents ship fast, so the bottleneck is follow-up
Can a chain of skills manage your pull requests from issue to merge?
It can get you most of the way. Some state should live in real code, because a loop made only of prompts eventually forgets a pull request or stalls on one it doesn’t expect.
When agents do most of the implementation, the week changes shape. Nikita’s starts with planning. Then comes roughly one day of shipping several features in parallel, and it ends with an internal demo where each person gets two minutes, which usually isn’t enough. The rest goes on deciding what to build, explaining decisions, telling customers what’s done and syncing with colleagues. His tooling went from a homemade “meta harness”, a private repo called Langfuse Pilotage that commits everything automatically and opens other terminals for tasks, back to the ChatGPT and Claude desktop apps plus Cursor cloud agents. The hard part is no longer producing code. It’s remembering to follow up so each pull request gets through review and merged.
AI was supposed to shorten the day. So far it has made it more intense, because you make far more decisions, and plenty of people are burning out. Nikita expects organisations to never be satisfied with the productivity gain. The day becomes near-continuous syncing with colleagues, customers and the market, maybe 15 minutes of prompting, and agents working overnight.
Our own follow-up loop is made of skills. A wrap-up skill, first described in episode #20, cleans up worktrees, updates main, merges open pull requests and files every loose end from the session as an issue. Later, issue-triage reads the open issues, sends a Cursor agent to check whether each is done or still makes sense, and shows us the picture. We pick three, run grill-me, implement and wrap up again.
Nikita’s point is that this whole loop is LLM calls. He’s experimenting with real workflow code that moves a task from stage to stage. The failures he expects are the dull ones. A human leaves a proper code review and the agent doesn’t know what to do with it. CI is flaky and the agent starts fixing CI for everyone instead of finishing its own change. A pull request drops out of view.
Skills are a fine place to start. Once a loop runs more often than you can supervise, move transitions such as “opened”, “in review” and “ready to merge” into code that can’t forget them, and let agents do the work inside each stage.
From traces to a golden dataset
Watching everything is the job Langfuse was built for, so it’s worth seeing how the pieces connect.
How do traces become evals that tell you a change made the agent worse?
You instrument the agent, score what it does, promote real cases into a dataset with the output you expect, and run every proposed change as an experiment against that dataset.
Your application sends traces to Langfuse over an OpenTelemetry-based protocol, which Langfuse accepts on an OTLP endpoint. If you’ve used OpenTelemetry or AWS X-Ray, a trace looks familiar: a tree of observations for each agent run, stored in ClickHouse and queryable.
“Nobody will actually instrument it by hand anymore, so we just don’t even put this.”Nikita Kabardin
The quick start hands you a prompt and a repo of skills for your coding agent, which does the instrumentation. Three years ago it was curl | bash. Now it’s a prompt, and running it feels just as scary.
Scores go on top. An LLM-as-a-judge evaluator can run on everything matching a filter. A score works like a tag with a value: accuracy 0.8, or a category such as “billing issue”. Code evaluators are pure Python or TypeScript functions. They receive one observation’s input, output, metadata and tool calls, plus the dataset item’s expected output and metadata when they run inside an experiment, and they attach the score to that observation. To check a whole agent run, target the root observation with the “Is Root Observation” filter. Code evaluators suit deterministic checks such as JSON validity or which tools were called. Langfuse also supports decision-model evaluators for narrow, typed verdicts, the classification job we covered in episode #21.
The human comes in while reading traces. In Langfuse a dataset is a golden set of inputs and the outputs you expect, the known good state of your system. Add the cases where the agent got it right. Also add the ones where it didn’t, with an expert writing the output it should have produced. Langfuse documents exactly that workflow for creating items from production data. Without the failures, a change can’t show you it fixed anything.
Then a new, cheaper model comes out and you want to know it’s no worse. For a simple setup, test an unpublished version of a managed prompt against the dataset inside the Langfuse UI. For a multi-step agent, run the experiment in your own code and send the results to Langfuse. You compare runs side by side in a table. Each dataset-item execution links to its trace, so when one result looks odd you click through to see what the agent actually did.
Don’t let the agent mark its own homework
Linking experiment results to traces matters because the obvious shortcut fails in a predictable way.
Why can’t a coding agent just improve your agent from the eval results?
It can do much of the work, but without a human checking the cases and the grading it may optimise for the eval passing rather than the agent improving.
Langfuse has an MCP server, a CLI and an API, so it’s tempting to hand the whole job over.
“Go to Langfuse and improve my agent. Your Claude will say, yeah, good. I will do that.”Nikita Kabardin
It runs experiments and does a lot of work, and it looks like progress. The agent may still be weak in some cases that matter. We’ve seen the mechanism ourselves. An agent running evals decides a test is failing and changes the test so it passes next time. Or it “improves” the prompt of the LLM judge that’s grading it. From where it sits, both count as success.
The work that actually helps isn’t glamorous. Nikita estimates designing a new multi-agent system is about 1% of AI engineering. The rest is looking at the real system and making it a little better each day, because you can’t improve it radically and fast. Many people drop out here, because it means reading what the customer actually asked and why the agent wasn’t helpful. Working out the why alone can take half a day.
That’s the real cost of evals, and no tool removes it. Agents can propose dataset cases and draft judges. A human still has to validate which cases go in, what the expected outputs say and whether the judge scores what you care about, by checking it now and then against cases they’ve read themselves.
The UI is for observing, the agent is for configuring
If humans have to read traces, the interface they read them in matters. That’s Nikita’s job at Langfuse, and his split is clear.
Do observability tools still need a UI when agents can use the MCP server?
Yes, for observing. Nikita says many people now work through agents without opening the UI, so the team puts less effort into things you do in the UI and more into things you look at there.
Setting up an experiment, configuring scores or annotation queues: ask your agent. Reading real sessions end to end, or spotting a pattern your agent didn’t flag: that needs a good desktop interface. Nikita’s example is cost. You open the dashboard and costs look fine, except for a spike. Why did you use so many tokens yesterday at 1 a.m.? You zoom into the spike, then into the trace behind it. His recent work points the same way: the new filter search bar, and a timeline that stays usable on large traces.
“Agents will not detect everything and we don’t want them honestly to detect absolutely everything.”Nikita Kabardin
Automated monitoring catches a lot. The UI is where a human spots the patterns it misses, then hands the investigation to an agent with a clear question. Annotation is the other human job: colleagues reading traces, scoring them by hand and leaving comments such as “this is weird, we need to look into it”. In practice that’s project management for evals. Reports such as a monthly cost review still belong there too.
Markdown files work until nobody is watching
The most common objection to a separate eval or context product is the one we hear about our own: “I already have Markdown files.”
When does a folder of Markdown files stop being enough for your agent?
When the agent runs on its own. A local agent can ask you when the file is wrong. An autonomous one has nobody to ask.
Locally, Markdown gets you a long way. Say you’re running OpenCode against an AWS organisation with 100 accounts. The thing it needs isn’t in your notes, so it asks, you answer, done. Now picture a couple of services that must stay up. After every production deploy, an autonomous agent watches production for 15 to 20 minutes, works through a checklist and looks for anomalies. That agent needs four things your laptop doesn’t give it:
- Context. Where things are, without a human to fill the gaps.
- Safe tools. Access to logs, metrics and the rest of production that can’t be used to do damage.
- Tracing. It runs in a box with no UI and nobody reads its log output, so something like Langfuse has to record what it did.
- A sandbox. A locked-down environment, so it can’t take your temporary AWS credentials and post them on Twitter.
Even locally, Markdown hits limits. It isn’t built to be queried, and it eats your context window. Many people who run memory add-ons built for Claude say the memory fills up within months and becomes unusable, with a large share of each session’s context spent before work starts. We covered that failure in episode #7.
The context item is what we build B.O.R.I.S for. ClickHouse built something similar internally for on-call engineers who wake up at 3 a.m. and need to know whether an incident is real and whether it’s hitting users.
Learning from customer data without leaking it
Autonomous agents watching production raise a harder question when the product itself is open source.
How can an agent turn what it sees in customer environments into a public fix without leaking customer data?
First decide whether you may use that data at all. If you may, make sure the agent that sees it can’t publish anything: it drafts an abstract issue, a human cleans it up, and only then does it enter the public pipeline, where the agents see nothing but that issue and public code.
Imagine your product is fully open source and you want it to improve itself by watching how it behaves in customer environments. The first constraint isn’t disclosure. Many products promise customers their data won’t be used to improve the product, and that promise rules out the loop before anything nears a pull request. Solve permission before disclosure.
A narrower case: a customer’s service control policy (SCP) blocks an action our scraper expects to succeed. The Lambda function fails, errors come in, and we want a pull request that handles it gracefully, without the account IDs from those alerts.
Layered review, several agents cleaning the change in turn, sounds like the answer. It fails for a simple reason.
“If you put it as a pull request, it’s already leaked.”Nikita Kabardin
Review in public CI is too late for the same reason. A “don’t disclose internals” line in a system prompt prevents some careless mistakes. It isn’t a control.
Nikita’s working answer is about access. The internal agent gets no publishing rights. It drafts a GitHub issue in abstract terms. A human reads it, catches the customer name it slipped in again, has it rewritten and files it. An internal software factory then takes the issue to a reviewed pull request. The agents in that factory get only public access: the approved issue and the public repository, nothing from customer environments or private traces.
Why not let the bot merge? Because nobody can yet say who is to blame when an agent-merged change breaks something. Agent identities are appearing, and Cursor runs Slack-launched agents under your GitHub account, but neither settles it. Keep a person with authority on the merge button, and log what the agents did so they can judge it.
Make your agent write less
Every handoff in these pipelines is text a human has to read.
How do I stop Claude writing walls of text in pull requests and docs?
Give it explicit writing rules and switch on a shorter output style.
Three options, heaviest first:
- ASD-STE100 Simplified Technical English, the aerospace and defence standard for controlled technical English. It’s free, but you fill in a form to download it, and each person gets their own copy, which makes it awkward to package as a shared skill.
- agent-style by Yue Zhao: 21 writing rules for coding agents, 12 from classic style guides and 9 from mistakes seen in LLM output.
agent-style enable claude-codewrites the rules into your project and references them fromCLAUDE.md. Our pull request descriptions and docs read much better since we adopted it. - Claude Code’s Concise output style: on v2.1.269 or later, run
/output-style concise. On earlier versions from v2.1.237, when Concise arrived, pick it through/config. Responses lead with the result and drop the preamble, narration and closing recap. Error reports, failing tests, security warnings and confirmations for destructive actions still come through in full.
Even in Concise mode you’ll sometimes get large amounts of slop in pull requests. Skimming is becoming a core skill, and it’s harder with AI text: speed-reading means knowing which chunks to skip, and in dense generated prose you don’t know where the one important detail is.
How hard you read should depend on what can break. On the UI side, Nikita accepts he’ll break things and fix them quickly, because usually nothing really bad happens. A migration, or anything that might touch ClickHouse performance, becomes a real project, with real humans reviewing every line.
Open source as a go-to-market strategy
All of this sits inside a market question: how does a tool you can self-host for free pay for itself?
How does an open source company make money if everything is free to self-host?
From companies that would rather pay than host and configure it themselves. The open source version earns the trust and name recognition that bring them in.
Nikita came back from AI Engineer Paris with a reality check. Most of his customer conversations start with a problem, so he’d assumed everyone was suffering. At the booth, most people already knew Langfuse and liked it, and the two or three who brought feedback were big customers pushing the product harder than its own team does. He also noticed that in Europe nearly everyone asks about self-hosting, governance, data privacy and sovereign models, while US conferences lean toward customer experience and revenue.
His view of the model: something free that is the default becomes a household name, and landing pages no longer make a product stand out. The usual mistake is to build an open source following first and launch the cloud later, by which time someone else hosts your project. Langfuse ran its cloud from day one and deploys straight from the public repo, with no private fork. The other failure mode is the licence change. Redis is the obvious example, and Valkey is the Linux Foundation fork that followed.
On air, Nikita said nothing is gated any more. The documentation mostly agrees. Langfuse moved its remaining product features to MIT in June 2025, including LLM-as-a-judge, annotation queues and prompt experiments. Enterprise modules such as SCIM, audit logs and data retention policies still need a commercial licence key when you self-host. Check that list against your compliance needs before an air-gapped rollout.
Where B.O.R.I.S fits
B.O.R.I.S is the context layer for AI agents, and it went the other way from Langfuse: it started with a few paying customers so we could check the product fits. The goal is that one tool call returns what an agent needs about the thing it’s working on: the repository, where it runs, in which environments and with which security groups. That leaves the agent’s context for the actual work. The hard parts are joining all of that up in a graph and pulling a coherent answer back out cheaply. If an LLM walks the graph on every query, you wait three minutes and the token cost just moves from your agent to ours. It isn’t open source today. A free tier is coming, and the waitlist is on getboris.ai/pricing.
What this means for teams
The pattern across every section is the same: agents do the volume, and a human decides what counts as good, what goes public and what merges.
- Turn ten traces into a dataset. Read real sessions, keep a few the agent got right and a few it got wrong with the output it should have given, and run your next model or prompt change against them before you ship.
- Keep the agent away from the eval it’s graded on. If an agent runs your experiments, make test and judge-prompt changes need your approval.
- Put a permission boundary between customer data and public output. Confirm you’re allowed to use the data, then let the agent that sees it draft an abstract issue a human files, never the pull request itself. The agents that write the fix see only that issue and the public repository.
- Try the Concise output style for a day, and add agent-style’s rules to repos where agents write docs and pull request descriptions.
Common questions, answered
How do I build an LLM eval dataset from production traces?
Pick real sessions from your traces, including failures, and save each as an input with the output you expect, written or corrected by someone who knows the right answer. Langfuse supports creating dataset items directly from production traces. Then run every model or prompt change as an experiment against that dataset and compare results before shipping.
How do I stop an AI coding agent from changing tests to make evals pass?
Treat tests, expected outputs and judge prompts as human-owned. Let the agent run experiments and propose fixes, but require your approval for any change to the dataset, the tests or the judge prompt. Check the judge against cases you’ve read yourself from time to time, so its scores still match what you care about.
Should an AI agent be allowed to merge its own pull requests?
Not yet, in our view. Who is accountable when an agent-merged change breaks something is still unresolved. A practical setup lets agents take an issue to a reviewed pull request, logs what they did, and leaves the merge to a person with authority to approve it.
Resources
- Nikita Kabardin, on LinkedIn and GitHub: worth following for the Langfuse interface work, including the Filter Search Bar launch and his first Langfuse talk in French at Station F.
- Langfuse evaluation docs: how traces, LLM-as-a-judge, code evaluators, datasets and experiments connect. Start here to build the trace-to-dataset loop.
- Langfuse experiments data model: settles how a run, its dataset-item executions and their traces link, which is what lets you click from a bad result into what the agent did.
- Langfuse via OpenTelemetry: the OTLP endpoint and headers, useful if you already run an OpenTelemetry collector.
- Open sourcing all Langfuse product features and self-hosting: which features became MIT in June 2025 and which enterprise modules still need a licence key.
- ClickHouse welcomes Langfuse: the January 2026 announcement, including the commitment to keep Langfuse open source and self-hostable.
- agent-style (Yue Zhao), Claude Code output styles and the v2.1.269 release notes: the writing rules you can drop into a repo today, what the Concise style removes and keeps, and the release that added
/output-style, so you know whether to use it or/config. - AI Engineer Paris: the conference behind the self-hosting and open source observations above.