Skip to main content

Voice agents in production: why did one voice agent sound more cheerful when nobody touched its prompt?

· 20 min read
Vadim Nicolai
Senior Software Engineer

A product owner told me one of three outbound agents "sounds the most joyful".

Not the funniest, not the fastest — the most joyful. When I pulled the three system prompts apart, the story wrote itself. One agent's prompt told it to sound "upbeat, with the friendly energy of someone sharing good news". Another opened with "Good morning" where a third opened with "Hello". A tidy causal chain, entirely inside the one artifact an LLM team can edit in an afternoon.

It was wrong.

The cheerful voice was the text-to-speech model. The morning briefing agent had been moved to a newer one before the other two, and on the same text, the same voice and the same settings, that model speaks roughly four semitones higher, with a wider pitch range and a brighter spectrum. By the time anyone asked the question, all three agents ran the same model and every setting on the dashboard was identical. There was nothing left in the configuration to find.

Perceived cheerfulness is set by the model version, not the prompt. What follows is the evidence, the one place where the prompt does reach the audio path, and the reason I trust the number.

Agent architecture: A2A vs MCP — tools inside your agent vs agents across the org chart, and where the line falls now that MCP went stateless and gained Tasks

· 21 min read
Vadim Nicolai
Senior Software Engineer

MCP connects an agent to tools, data sources and context: vertical, inside one agent. A2A connects agents to other agents: horizontal, across teams, departments and organisations. They are complementary, not competing.

That definition is the easy part. The decision it implies is where teams get it wrong, and the 2026 revisions removed the evidence they used to defend it.

A long-running MCP tool call and an A2A task are now the same HTTP request. Both encode calls as JSON-RPC 2.0 over stateless HTTP — MCP as its base encoding, A2A as one of three bindings (MCP specification, A2A specification). Both return a server-generated ID. Both can sit waiting for input, both offer polling plus a push path, and both run behind an ordinary load balancer with no session affinity.

A2A's task lifecycle: what the async contract fixes, and what it leaves to you

· 10 min read
Vadim Nicolai
Senior Software Engineer

The A2A v1.0 specification says SendMessage "MUST return immediately with either task information or response message" (§3.1.1). A dozen subsections later it says "Operations are blocking by default", and a blocking call "MUST wait until the task reaches a terminal state ... or an interrupted state" (§3.2.2). The default client does not wait for completion. It waits for a pause, which may be a human deciding whether to approve something.

That contradiction is a fair summary of the whole protocol: a precise vocabulary for where work is, and almost no guarantee that it happened.

Secure Your AI Agent's VPS: Close Port 22 Without Tailscale

· 15 min read
Vadim Nicolai
Senior Software Engineer

To close port 22 without a VPN: put SSH behind an outbound-only tunnel with an identity check at the edge, prove the new path works, then delete the old port-22 firewall rule over that new path. Nothing on the host listens for inbound connections, and no mesh VPN client has to share the laptop with a corporate one.

An exposed SSH port is not the likeliest way your AI agent's VPS gets owned. It is the likeliest way you lose the ability to fix the machine.

The standard recipe keeps port 22 shut to the world, opens it to your own address, and adds a mesh VPN when you need to get in from anywhere. Both halves tie the security of a machine that holds live API keys to something you do not control: your ISP's address pool, and a laptop routing table that a corporate VPN client already owns. What I would reach for instead is Cloudflare Tunnel with Cloudflare Access in front — an outbound-only connection plus an identity check, which takes both dependencies out of the security model.

An ElevenLabs voice on a Twilio call: four routes, one 8 kHz wire

· 10 min read
Vadim Nicolai
Senior Software Engineer

Here is a fact that should make telephony simple. A Twilio call carries μ-law audio at 8000 Hz, and ElevenLabs' text-to-speech API will produce exactly that if you ask for the ulaw_8000 output format. One vendor's output is the other vendor's wire format. No resampling, no transcoding.

And yet none of the four common ways to put an ElevenLabs voice on a Twilio call just passes those bytes through. Each one converts somewhere, and where it converts tells you who is in charge of the audio.

LiveKit, Pipecat, and the seam where a voice agent actually breaks

· 8 min read
Vadim Nicolai
Senior Software Engineer

Pick a voice stack and you are really picking two things: who owns the audio transport, and what you can see when a turn goes wrong. The vendor list barely matters — LiveKit Agents and Pipecat both put the same speech-to-text, language model and text-to-speech providers behind the same three seams, and both let you swap any of them in a line.

LangSmith shipped tracing for both in July, alongside OpenAI Realtime and Gemini Live. That is a useful forcing function for this comparison, because a trace shows you where a framework's abstraction actually sits — and the two answers are genuinely different.

Twilio behind LiveKit, behind Pipecat, or on its own: three places to put the phone line

· 9 min read
Vadim Nicolai
Senior Software Engineer

"Add Twilio" sounds like one integration. It is three, and they disagree about the most basic question in a voice agent: who holds the audio. LiveKit takes the call as a SIP trunk and turns the caller into a participant in a room. Pipecat takes it as a WebSocket carrying raw audio frames. And Twilio's own ConversationRelay keeps the audio entirely, and hands your server text.

The previous post compared LiveKit and Pipecat on the seam where a voice agent breaks: a caller talking over the bot. Put a phone line in front of either one and that seam moves, because now there is a third system buffering audio that neither framework controls.

LangSmith's voice tracing against the rest: same spans, different audio

· 15 min read
Vadim Nicolai
Senior Software Engineer

In July, LangSmith launched voice tracing for four frameworks: Pipecat, LiveKit, OpenAI Realtime and Gemini Live. The launch post lists what a trace captures, including full conversation audio overlaid on the trace, speech-to-text and text-to-speech latency, voice activity detection events, and interruptions and overlapping speech.

It is not the only way to trace a voice agent, and the alternatives are more alike than their marketing suggests. Most of them read the same spans from the same place. The real differences are further down: which audio each one records, how much of each is actually open source, and what happens when there are no spans to read.

RemoteGraph, and what every other framework does instead

· 13 min read
Vadim Nicolai
Senior Software Engineer

LangGraph ships a class that turns an agent running on another machine into a node in your graph. You construct a RemoteGraph, pass it to builder.add_node(...), and the call site reads exactly like a local subgraph. It subclasses PregelProtocol, so as far as your parent graph is concerned there is no network there at all.

I wanted to know what that costs, and what the other frameworks offer in its place. Two of them can be measured: the same parent-and-child pair runs against a real LangGraph server and against a LlamaIndex workflow server, with the same five contract mismatches pushed through each boundary. CrewAI, the OpenAI Agents SDK and A2A are read from their docs and specs, because two of the three turn out to have no remote-agent primitive to measure.

The headline result is an inversion. RemoteGraph looks typed and validates nothing at the boundary: I renamed one field in the child, redeployed it, left the parent untouched, and the parent returned its own input with no error and no warning — a result indistinguishable from the child never running. LlamaIndex has no RemoteGraph at all; its server-and-client pair looks like a raw HTTP call, publishes a JSON Schema, and rejects a malformed request with a 400 and per-field errors. The same rename throws a KeyError on the line I wrote.

A fifth of my trading universe was delisted stocks, understating every signal by half

· 19 min read
Vadim Nicolai
Senior Software Engineer

The number I did not expect was 20.5%. That is the share of my screen-eligible universe that had already delisted and fallen out of the research panel before I scored a single signal on it. The direction was the bigger surprise. Dropping those names did not flatter my backtest; it suppressed it, understating every lane's spread by a mean 19.3 bps (se 6.5) against a headline lane spread of 37.99 bps. Just over half of a result, traceable to an instrument nobody had audited.