Engines

Pipeline, speech-to-speech or hybrid?

September 2026 Β· 7 min read Β· by the NordTell team

Every NordTell agent has an engine setting with three options. People ask which is best. The honest answer is that they are good at different things, and that the setting exists because we could not pick one for everybody. Here is how to choose.

Pipeline: three specialists in a row

The classic design. A speech-to-text model listens and writes down what the caller said. A language model reads it and writes a reply. A text-to-speech model reads the reply aloud. Three models, three companies if you like, each the best at its job.

Good at: control and choice. You pick the listener, the thinker and the voice separately. It has the widest range of voices and the lowest cost per minute. The transcript is exact, because the transcript is literally what the thinker saw.

Less good at: rhythm. Each hop adds a little delay, and the agent cannot hear tone, only words. For a receptionist that books appointments this rarely matters. For a long, emotional conversation it can.

Speech-to-speech: one model that hears and talks

A realtime model takes audio in and gives audio out. There is no transcript in the middle; the model "hears" the caller, hesitations and all, and answers in its own voice with almost no delay. Laughter, interruptions and overlapping speech feel natural.

Good at: feel. Callers relax. It is the closest to a person on the line.

Less good at: discipline. Realtime models are eager. Left alone, one will occasionally say "I have booked that" before the calendar tool has answered, or wander off the script when a caller is persuasive. Fewer voices, and a higher price per minute.

Hybrid: realtime, with a second model watching

This is what we recommend, and what most of our agents run. The caller talks to a realtime model, so the rhythm is natural. But every reply is checked by a second model before it stands: did the agent claim an action that has not happened? Did the caller ask for a person? Is this the moment to transfer, or to end the call? If so, the safety net acts: it corrects, it transfers, it hangs up politely.

The nets are written per language. "Stil mig om til en kollega" and "put me through to someone" are different sentences with the same meaning, and the net knows both. Thirteen languages have them today.

Good at: the combination. Natural feel, and an agent that does not bluff.

Less good at: nothing dramatic. It costs a little more per minute than pipeline, because two models are working.

A simple rule

  • Booking, reception, order status, surveys: hybrid. The feel matters and so does the discipline.
  • High volume, tight budget, strict script: pipeline. Cheapest, most controllable.
  • Long conversations where tone is the product: speech-to-speech, and test it hard.

You can change the engine per agent at any time, and the agent keeps its prompt, tools and number. Try two on the same prompt and listen; it takes ten minutes and tells you more than this post.

Engines in the product Β· Try it free