Skip to main content

5 posts tagged with "agent"

View All Tags

Voice agents in production: why did one voice agent sound more cheerful when nobody touched its prompt?

· 20 min read
Vadim Nicolai
Senior Software Engineer

A product owner told me one of three outbound agents "sounds the most joyful".

Not the funniest, not the fastest — the most joyful. When I pulled the three system prompts apart, the story wrote itself. One agent's prompt told it to sound "upbeat, with the friendly energy of someone sharing good news". Another opened with "Good morning" where a third opened with "Hello". A tidy causal chain, entirely inside the one artifact an LLM team can edit in an afternoon.

It was wrong.

The cheerful voice was the text-to-speech model. The morning briefing agent had been moved to a newer one before the other two, and on the same text, the same voice and the same settings, that model speaks roughly four semitones higher, with a wider pitch range and a brighter spectrum. By the time anyone asked the question, all three agents ran the same model and every setting on the dashboard was identical. There was nothing left in the configuration to find.

Perceived cheerfulness is set by the model version, not the prompt. What follows is the evidence, the one place where the prompt does reach the audio path, and the reason I trust the number.

Agent architecture: A2A vs MCP — tools inside your agent vs agents across the org chart, and where the line falls now that MCP went stateless and gained Tasks

· 21 min read
Vadim Nicolai
Senior Software Engineer

MCP connects an agent to tools, data sources and context: vertical, inside one agent. A2A connects agents to other agents: horizontal, across teams, departments and organisations. They are complementary, not competing.

That definition is the easy part. The decision it implies is where teams get it wrong, and the 2026 revisions removed the evidence they used to defend it.

A long-running MCP tool call and an A2A task are now the same HTTP request. Both encode calls as JSON-RPC 2.0 over stateless HTTP — MCP as its base encoding, A2A as one of three bindings (MCP specification, A2A specification). Both return a server-generated ID. Both can sit waiting for input, both offer polling plus a push path, and both run behind an ordinary load balancer with no session affinity.

A2A's task lifecycle: what the async contract fixes, and what it leaves to you

· 10 min read
Vadim Nicolai
Senior Software Engineer

The A2A v1.0 specification says SendMessage "MUST return immediately with either task information or response message" (§3.1.1). A dozen subsections later it says "Operations are blocking by default", and a blocking call "MUST wait until the task reaches a terminal state ... or an interrupted state" (§3.2.2). The default client does not wait for completion. It waits for a pause, which may be a human deciding whether to approve something.

That contradiction is a fair summary of the whole protocol: a precise vocabulary for where work is, and almost no guarantee that it happened.

Self-evolving agents: survivorship bias wrong way in stocks

· 19 min read
Vadim Nicolai
Senior Software Engineer

Survivorship bias is supposed to flatter a backtest. A survivor-only universe deletes the names that died along the way. Every number computed on it should therefore come out looking better than the truth. That is the textbook direction — and for this board, the textbooks had it backwards.

The measurement that broke the assumption came from a 10-minute autonomous research loop. It ran the previous evening and logged the result as a measurement only: no lane, constant, module, or gate default was changed.

The loop re-screened its own universe. The survivor-only reference — a single active=true snapshot of Polygon's ticker list — had been used to type every name on all 236 point-in-time dates. That reference produced a benchmark that was too low.

Readmitting every name the gate had silently excluded moved the equal-weighted screened universe from +5.64 to +6.54 bps at k=1, and from +27.34 to +29.07 bps at k=5.

Read that table twice.

equal-weighted screened universesurvivor-onlyall names readmitted
k=1+5.64 bps+6.54 bps
k=5+27.34 bps+29.07 bps

The bias did not flatter the backtest. It censored the names that made the backtest look worse. The reason is structural, not mystical: this panel never observes a delisting as a return. There is no −100% row to be spared.

Removing names did not remove disasters. It removed a type of name — and that type was exactly what the extreme-return lanes were looking for.

Agentic CLEAR: Automating Multi-Level Agent Evaluation — and the Autonomy Gate It Unlocks

· 17 min read
Vadim Nicolai
Senior Software Engineer

Every team running an agent fleet has the same blind spot. Observability platforms—MLflow, Langfuse, home-grown OpenTelemetry—capture execution traces beautifully. They show you what the agent did. They say almost nothing about whether it did it well. So a developer opens the trace viewer, scrolls through a few hundred spans, and tries to eyeball a systemic failure out of thousands of runs. The research alternative is worse: hand-built error taxonomies that take weeks to annotate and go stale the moment the agent changes. What both approaches lack is automated multi-level agent evaluation—judgment of the trajectory itself, not just a record of it.

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents, by Yehudai, Eden, and Shmueli-Scheuer (2026) at IBM Research, attacks exactly this gap. It is an open-source Python package—pip install clear-eval—that reads raw agent traces and produces data-driven evaluation at three levels of granularity, surfaces recurring failure patterns without a predefined taxonomy, and renders the whole thing in an interactive dashboard. It reports up to 0.890 AUC for predicting trajectory success in a fully reference-less setting. This post walks through what the paper actually does, then shows how I wired the same multi-level shape into a 45-graph production fleet as an "autonomy gate"—the component that turns a human approval interrupt into a machine one.