Playbooks

A test plan for your voice agent before it takes real calls

8 September 2026 Β· 5 min read Β· by the NordTell team

An agent that sounded great in a demo is not launched. It has passed one call, with a friendly person, in a quiet room, asking the question it was built for. Real callers mumble, interrupt, switch language, ask if it is a robot and hang up mid-sentence. This is a plan for finding out what your agent does then, before your customers do. It takes an afternoon.

Round one: the browser

Start with the microphone test in the browser, before any number is attached. It is free, it is fast, and it lets you change the instructions and try again within a minute.

Write down the five to ten things callers actually want, in their words, not yours: "can I get a time on Thursday", "how much is a check-up", "I need to speak to someone about my bill". Run each one. Listen for three things: did the greeting say it is an AI, did it read names and dates back correctly, and did it use the right tool at the right moment? A booking that appears in the calendar before the agent says "booked" is a pass. A "booked" with no row in the calendar is a fail, and the most important one to catch.

Fix the instructions after each failure, one line at a time, and rerun the whole list. Keep the list. It becomes your regression test for every later change.

Round two: a real phone

The phone is a different animal. The audio is narrower, there is delay, and people speak differently to a phone than to a laptop. Call the agent from a mobile, from a landline, from a car with the window down, and on speakerphone in a kitchen. Ask a colleague with a strong accent to call. Ask someone who has never heard of the project to call and book something without instructions.

Test the transfer for real: does the hold music play, does the colleague get a briefing, and what happens when the colleague does not pick up? Point the transfer at a voicemail on purpose; the post on how a warm transfer works describes what should happen. Then check the transcript and summary of each call in the dashboard. The summary is what your team will read every morning, so it needs to be right.

The edge cases

These are the calls that separate a demo from a product. Decide what you want to happen for each, then test whether it does.

  • Silence: say nothing after the greeting. It should ask once whether you are there, wait, and end politely, not talk to itself for a minute.
  • Interruptions: talk over it mid-sentence. It should stop and listen, then pick up from what you said.
  • Wrong language: answer in a language other than the agent's. It should switch, or say plainly which languages it speaks.
  • "Are you a robot?": it should say yes, simply, and carry on. No jokes, no dodging.
  • Asking for a human: it should offer a transfer or a callback the first time, without a sales pitch first.
  • Hanging up mid-sentence: the call should end cleanly and the summary should still be written.
  • Numbers and names: give a phone number quickly, spell an unusual name. It should read them back before using them.
  • The rambler: tell a long story with the question buried in the middle. It should find the question and ask one thing at a time.
  • The angry caller: complain, loudly. It should apologise once, not argue, and fetch a person.
  • Two people talking: have someone chat in the background. It should stay with the caller.

What to write down

Keep a plain test sheet: the line you said, what the agent did, pass or fail, and the fix. Four columns are enough. After each fix, rerun the whole sheet, not just the line that failed. A change that fixes the rambler can break the booking.

The dashboard gives you the raw material: the transcript of each test call, the summary and outcome, and the minutes it cost. Reading ten transcripts in a row shows patterns a single call hides, such as an agent that always asks for the phone number twice.

Launch small, then widen

Do not switch the main line over on a Monday morning. Route the after-hours calls first, or the overflow when the desk is busy, and keep the main line where it is. Listen in live on the first day; you can hear a call as it happens without joining it. Read every summary for a week. Then widen.

Before the first real call, set the things that are hard to undo: recording (off, on, or on after consent), how long transcripts and recordings are kept, and who on the team can see them. And put the test sheet somewhere you will find it, because the next change to the instructions needs the same afternoon.

Where NordTell fits

Every NordTell agent can be tested in the browser with a microphone before a number is attached, and every call, test or real, has a transcript, a summary, an outcome and a minutes cost. You can listen in live, and the agent says it is an AI in its greeting and keeps a way to a person.

Agents in the product Β· Calls and transcripts Β· Languages