Skip to main content

Voice agents in production: why did one voice agent sound more cheerful when nobody touched its prompt?

· 20 min read
Vadim Nicolai
Senior Software Engineer

A product owner told me one of three outbound agents "sounds the most joyful".

Not the funniest, not the fastest — the most joyful. When I pulled the three system prompts apart, the story wrote itself. One agent's prompt told it to sound "upbeat, with the friendly energy of someone sharing good news". Another opened with "Good morning" where a third opened with "Hello". A tidy causal chain, entirely inside the one artifact an LLM team can edit in an afternoon.

It was wrong.

The cheerful voice was the text-to-speech model. The morning briefing agent had been moved to a newer one before the other two, and on the same text, the same voice and the same settings, that model speaks roughly four semitones higher, with a wider pitch range and a brighter spectrum. By the time anyone asked the question, all three agents ran the same model and every setting on the dashboard was identical. There was nothing left in the configuration to find.

Perceived cheerfulness is set by the model version, not the prompt. What follows is the evidence, the one place where the prompt does reach the audio path, and the reason I trust the number.

The prompt never reached the acoustics — the TTS model did​

Loading diagram…

The clean experiment is the one where nothing moves except the model. Two of the agents' voicemail texts, same voice, stability 0.55, similarity 0.8, speed 1.0, four renders per text per model, through the text-to-speech API, with every model name taken from the models page.

Flash v2 rendered the text at a median 202 Hz, 95% CI [198, 205]. Eleven v4 Turbo rendered the same text at 250 Hz, CI [248, 253]. That is +3.9 semitones on identical input — nearly a major third, more than the gap between two adjacent keys on a piano. (The 3.9 headline is the pooled 251 Hz versus 200 Hz estimate across all eight renders on each side; the two medians quoted above, 250 and 202, come to 3.7.) This is not a subtle difference in timbre. It is the same sentence sung a third higher.

Total separation is what convinced me. Not one of the eight Eleven v4 Turbo renders came in below any of the eight Flash v2 renders. Every interval here is a bootstrap over per-clip medians, 3,000–4,000 resamples, per Efron (1979).

The stakes were concrete. The request was to make the other two agents sound like the cheerful one. Rewriting the prompt was the cheap, plausible, one-afternoon fix — and it would have shipped, been called done, and moved the voice by less than the run-to-run noise of the thing they were chasing. A ticket closed with no measurement is worse than an open one, because nobody looks at it again.

This is measured, not read: four renders per text per model against the live API, with bootstrap intervals, not a claim from documentation.

"Cheerful" is a pitch level, a pitch range and a spectral centroid​

The vocal-emotion literature was already waiting for me. Banse and Scherer's (1996) paper is literally titled "Acoustic profiles in vocal emotion expression", and it maps listener-rated emotion onto measured acoustic parameters: pitch level, pitch movement, spectral energy distribution, speech rate. Juslin and Laukka (2003) carry the same dimensions across two channels and title the result "Different channels, same code?". Both are claims about human voices, so the transfer needs an argument rather than an assumption: a listener brings the same perceptual apparatus to a rendered voice as to a recorded one, and if an acoustic profile is what carries the percept in the human case, the same dimensions carry it in the synthetic case.

That argument holds for correlates. It does not make the acoustic delta identical to the percept. Hold that thought.

What it does mean is that "joyful" decomposes into quantities I can measure. Median F0 over voiced frames, tracked with pYIN, Mauch and Dixon (2014), 150–450 Hz search band — high enough to drop most of a male callee bleeding into a mono recording. Pitch range as the 10th-to-90th percentile spread in semitones. Spectral centroid over the louder 60% of frames, as the brightness proxy.

Semitones, not hertz, for every comparison: 12·log2(f1/f2). The three agents' voiced pitches differed by 38 Hz in the first experiment, which sounds like a lot until you notice it is 2.8 semitones, and that the semitone figure is the one that tracks what an ear hears.

There is a second reason to insist on the acoustic framing. "It's in a better mood today" is not a hypothesis. Mood language hides the dials.

Every knob your team actually reaches for is flat​

This is the part that ought to change how people work, and it is almost entirely negative results.

On Eleven v4 Turbo, crossing stability 0.3 / 0.55 / 0.8 against speed 0.9 / 1.0 / 1.1 gave nine cells spanning 246–255 Hz. Flat. Stability, similarity and speed are real parameters — they are the fields the request payload carries, which is exactly why the pinning test later in this piece reads the payload rather than the console. They do other things. They did not move pitch level by anything resembling the effect I was chasing.

Then the wording. "Good morning" versus "Hello" in the same opening line: under one semitone. Added exclamation marks: under one semitone. Swapping one agent's entire voicemail text for another agent's: under one semitone.

Read that list again. Those are the four levers a team pulls when a stakeholder says "make it sound friendlier". Together they sum to less than the difference between two adjacent renders of the same text.

There was one prompt-side effect, and it deserves its own paragraph because it is the strongest thing anyone can say against me. The "upbeat … sharing good news" line did change what the model emitted: in expressive mode, the LLM sometimes prepended an audio tag such as [happy] to its reply — on the two agents whose prompt carried that line, and never on the third. That is the prompt reaching the synthesis path. What I did not do is measure what the tag did to the audio. So there is evidence of a text-level change and no evidence of an acoustic one, which is not the same as evidence of no acoustic change.

Version history answers what the config dump cannot​

Here is where my expectation lost.

I assumed the three live agents would differ today, and that a config dump would show me where. So I sent each live agent the same "Hello?" over a text websocket session — no phone call placed — five times, and measured the spoken reply. All three: 247–253 Hz. No gap at all.

The current configuration could not explain the owner's ear. Neither could any version of it I could read.

The archive could. Pulling the first agent turn of each call over the previous 30 days from the vendor's stored call audio (agents platform overview) gave the morning briefing agent a median 266 Hz across 25 openers, against 221 Hz across 8 for the job-offer agent and 228 Hz across 25 for the re-engagement agent. That is +2.8 semitones, Cohen's d = 1.32, and a random morning-briefing opener came in higher than a random other opener 81% of the time.

So the instrument I kept reaching for was pointed the wrong way in time. A config dump is a snapshot of the present tense. The question was historical. Had I trusted it, I would have reported "no difference exists", delivered with full confidence, having measured the one thing that had already converged.

Grouping those recorded openers by date turned a distribution into a staircase. The morning briefing agent sat at 180–192 Hz before its switch and 258–304 Hz after. The job-offer agent went from 195–226 Hz to 290 Hz. The re-engagement agent moved from roughly 200–250 Hz to 275–285 Hz. Each step lined up with a new agent version in the vendor's history, diffed through the agents platform voice settings.

Then the useful part. The one change common to all three agents was the TTS model setting, moving to Eleven v4 Turbo — from Flash v2 on two agents, from Eleven v3 Conversational on the third. For the job-offer agent it was the only change. Stability, speed, similarity, expressive mode, the LLM and the prompt's tone line were identical on both sides.

The morning briefing agent switched first. That is the stretch in which it alone "sounded joyful", and I am deliberately not saying how long the stretch was, because the honest answer is a small number of days and the useful answer is the mechanism.

Pooled by the model each version ran, the picture tightens: Eleven v4 Turbo openers at 276 Hz (n=25), Flash v2 at 211 Hz (n=20), Eleven v3 Conversational at 231 Hz (n=13). v4 Turbo minus the older models is +57 Hz, 95% CI [+46, +68], with pitch range +1.2 semitones [+0.3, +2.0] and spectral centroid +117 Hz [+67, +167]. Pitch level, pitch movement and brightness — the three dimensions the emotion literature points at.

The morning briefing agent alone, after minus before its switch, is +73 Hz with CI [+45, +95] across 22 calls against 3. Three. The one agent a human actually complained about carries the weakest before-side evidence in this piece, and I am not going to let it carry the argument when the pooled comparison does the job properly.

The obvious instrument was measuring the telephone line​

If "joyful" is the question, run an emotion classifier. That was the plan, and it was the wrong instrument.

I ran the SUPERB emotion model — wav2vec2, four classes, via the SUPERB benchmark from Yang et al. (2021), trained on IEMOCAP, Busso et al. (2008) — over every clip. Two results killed it.

First, a control. Passing the same clean clip through an 8 kHz mu-law round trip raised P(happy) by +0.34 on average, on both models. Pitch moved under 1 Hz. The clip did not get happier. It got compressed. The codec is the one ITU-T calls "Pulse code modulation (PCM) of voice frequencies" — the standard's own title for the 8-bit, 8 kHz companded PCM the telephone network runs on (ITU-T G.711, retrieved for this article).

Second, the real calls. v4 Turbo minus the older models in P(happy): −0.04, 95% CI [−0.17, +0.10]. Nothing. On audio where pitch, pitch range and spectral brightness had all moved clearly and consistently.

The likely reason is in the corpus it was trained on. IEMOCAP is studio-recorded, acted, wideband speech. A classifier fit to that is reading the channel as much as the voice — and a phone line is a channel. Had I installed that model as an alert, I would have a detector that reacts to codec changes and sleeps through the thing I care about, and it would have looked rigorous the whole time, because it has a confidence interval and a model card.

A 4-semitone claim is evidence only after the tracker has been tested​

Before putting a 3.9-semitone finding in front of anyone, I tried to break the thing that produced it.

I wrote a plain-Python YIN, de Cheveigné and Kawahara (2002) — a paper whose title states its scope exactly, "YIN, a fundamental frequency estimator for speech and music" — alongside pYIN, so the tracker itself could be unit tested where numpy is not installed. Then I ran it against synthetic voices of known pitch.

A 160–400 Hz grid: worst error 0.025 semitones. A missing fundamental — harmonics 2–6 only, which is what a phone line tends to leave you — still read the fundamental. White noise at 30, 20 and 10 dB SNR: within 0.1 semitones. At 5 dB and 0 dB SNR it reported no pitch rather than a wrong one, which is the behaviour that matters most; a tracker that guesses at the noise floor will manufacture effects. Vibrato of ±1 semitone measured 1.83 semitones of range against a theoretical 1.90, SD 0.68 against 0.71. A 4-semitone glide measured 3.07 against 3.2, and ended high. A male voice at 110 Hz was never reported — below the floor, and its 220 Hz harmonic was not mistaken for the agent.

An 8 kHz mu-law round trip, again via G.711, moved measured pitch by 0.000 semitones. The channel does not move pitch. That is why the classifier's channel response in the previous section was not a pitch artifact, and why the pitch numbers in this one survive the same codec.

On the real audio the two trackers agreed within 0.24–0.7 semitones on average, at worst 1.6 — well inside a 4-semitone effect. A naive tracker without a voicing decision disagreed by 1.28 semitones on average by reading unvoiced frames as pitch. That is the argument for pYIN's voicing probabilities: Mauch and Dixon's title is "pYIN: A fundamental frequency estimator using probabilistic threshold distributions", and it is the distributions, not the threshold, that stop a silent frame being reported as a note.

I also planted defects to check the tests were awake. Removing the tracker's parabolic interpolation fails 9 tests. Making it skip the first dip under the threshold fails 19.

The strongest case that I measured the wrong thing​

Let me state the opposing view as its holder would.

You measured the wrong thing, and the prompt does reach the voice. Your own Experiment 5 shows the "upbeat, sharing good news" line changing model behaviour — it prepended [happy] on exactly the two agents whose prompt carried the line and never on the third. You concede pitch, range and brightness are correlates of perceived cheerfulness, that you ran no listening test, that the agent a human actually heard as cheerful has three "before" recordings, and that all your clean numbers come from voicemail-text renders rather than the calls the owner was complaining about. You have a strong result about synthetic renders and an anecdote about the incident.

The advocate is right about three things, and I concede them without hedging. The prompt can reach the audio path — the [happy] tag is real, it appeared only where the tone line was, and I did not measure its acoustic effect. That is a gap in my work, not a gap in the argument, but it is a gap. The perception claim is genuinely weaker than the acoustic claim: the owner's ear found the problem, and my measurements explain a correlate of what that ear heard — I cannot claim the 3.9 semitones are the cheerfulness. And the morning-briefing before/after is thin on the before side, three calls against 22.

The thesis survives where it matters. The 3.9-semitone effect was measured with text, voice and settings held constant, with total separation between the two groups of eight renders. Every knob a team would actually turn — stability, speed, greeting wording, exclamation marks, even a full text swap — landed under one semitone. And the live agents are indistinguishable at 247–253 Hz today, so nothing in the current configuration explains the owner's ear either.

The prompt may have moved a tag. It did not move the voice.

Practical takeaways: pin the model, probe the agents, re-derive the gap​

Three things came out of this, and none of them is "be more careful".

A test that holds every voice agent to the same TTS model, voice and settings, read from the payload actually sent to the vendor rather than from the config file. Reading the config would have caught nothing here, because the config was correct and identical. I paired it with a test that pulls a live agent whose model was downgraded in the console and shows the pin catching it — a self-run demonstration, not an independent reproduction, and worth labelling as such.

A measuring script that sends the same input to every live agent and flags any agent more than 1.5 semitones from the reference. Run on demand, no phone call placed. That threshold is chosen against a measured noise floor — two trackers agreeing within 0.24–0.7 semitones on real audio — against a 3.9-semitone effect. A zero threshold would have fired on tracker disagreement and closed the investigation on a rounding error.

And the tracker's own tests, plus a fixture of real pitch traces — numbers only, no audio — on which the model gap and its bootstrap interval are re-derived in CI. If a future model release moves the voice, I want the fixture to fail before a stakeholder's ear does.

Now the alternatives, with what each actually costs.

Steering tone in the prompt, by wording or by expressive-mode audio tags, is cheap and honest about one thing: it is the only lever that lets you shift expression per turn. It costs you reach — under one semitone of pitch movement — plus a tag that appeared non-deterministically on two agents and never on the third, whose acoustic effect I never measured. Choose it when you want an intentional expressive beat in a specific turn and you are willing to test the render. Do not choose it to explain or prevent drift.

An emotion classifier in the eval loop is fast to stand up and reads the channel: +0.34 on a codec round trip against −0.04 on the real calls. Choose it when your audio path is wideband and identical across everything you compare, so the classifier is not reading the codec instead of the voice.

A human listening panel is the right primary instrument when the question is the brand percept — "is this the voice we want" — rather than "did the acoustics move". It costs time, it needs a rubric and enough raters to be stable, and there is no validated scale for vocal cheerfulness I could borrow from the material I had.

Upgrading the whole fleet together, so no agent is ever the odd one out, buys cross-agent consistency at the price of a stable reference: you cannot separate a model change from a content or context change, you re-baseline on every provider release, and you never learn what the previous sound was.

The decision framework is short. If the question is "why did it change", diff the recorded audio against the version history — do not read the config, because the config has converged. If the question is "has it moved", run the pinned-model test and the 1.5-semitone probe. If the question is "is this the voice we want", put humans on it, alongside the acoustic measures rather than instead of them.

Two candidate causes that belong on any list of this kind are not tested here: conversation history and memory contamination, and sampling non-determinism in the LLM that writes the text. Neither is in this measurement set, and I am not going to imply otherwise. They would surface as text-side changes first, so the pin-and-probe pair above would catch them only indirectly — the honest statement is that I did not look.

The limits, stated plainly. Three calls on the before side for the morning briefing agent. Mono recordings with a 150 Hz floor that removes most of a male callee and not a female one, which is exactly why the opener turn was used, since the callee is usually silent through it. One voice; another voice on the same models may move differently. Pitch, range and brightness are correlates of perceived cheerfulness, not a listening test, and the owner's ear was the original observation. And these numbers are for the model versions served on one day; the vendor's models change.

Short answers to the questions this raises​

Can a voice agent's tone change without a prompt edit? Yes. The spoken output is produced by the LLM, the TTS model, the voice settings, the injected context and the audio path. In the case above, the prompt was untouched and the TTS model had moved.

Why does my agent sound different after a provider update? Because managed providers update models behind stable-looking names. The mitigating move is to read the model and voice fields out of the request payload you actually send, assert on them, and alert when they change.

Does conversation history affect tone? It can, by changing the words the LLM emits. It is not what happened here, and I did not measure it.

How do I test for drift? Send identical input to every live agent on demand, compare pitch against a reference, and re-derive the effect size from stored traces in CI. Do not rely on subjective listening alone, and do not diff transcripts — a transcript diff is blind to F0 by construction.

The second author​

The prompt is the only part of a voice agent you can see, edit and redeploy in an afternoon. That is precisely why it is the first suspect and the last explanation. The cheerful agent had no change anywhere a human had reached. Someone else moved a model, the audio moved nearly a major third, and every artifact written in text — the prompt, the config, the transcript — recorded nothing.

That generalises past voice agents. Any behaviour you produce through someone else's versioned model has a second author. They ship on their schedule, their changelog is not your regression suite, and the only evidence they leave you is the output itself.

So the question is not whether your prompt is well written. It is whether you have anything that would notice if the voice changed tonight.