Voice agents in production: why did one voice agent sound more cheerful when nobody touched its prompt?
A product owner told me one of three outbound agents "sounds the most joyful".
Not the funniest, not the fastest — the most joyful. When I pulled the three system prompts apart, the story wrote itself. One agent's prompt told it to sound "upbeat, with the friendly energy of someone sharing good news". Another opened with "Good morning" where a third opened with "Hello". A tidy causal chain, entirely inside the one artifact an LLM team can edit in an afternoon.
It was wrong.
The cheerful voice was the text-to-speech model. The morning briefing agent had been moved to a newer one before the other two, and on the same text, the same voice and the same settings, that model speaks roughly four semitones higher, with a wider pitch range and a brighter spectrum. By the time anyone asked the question, all three agents ran the same model and every setting on the dashboard was identical. There was nothing left in the configuration to find.
Perceived cheerfulness is set by the model version, not the prompt. What follows is the evidence, the one place where the prompt does reach the audio path, and the reason I trust the number.
