Skip to content

A session, end to end

A Voqalize call touches four things you own: a WebSocket route, an agent record, one HTTP request from your server, and a page running pipecat. Everything else — the audio, the recognizer, the voice, the turn-taking — happens between them.

This page is that path in order. It is the map; each step names the page that goes deeper.

flowchart LR
  A["Write a brain<br>one WebSocket route"] --> B["Give it an address<br>a public URL, or Cortex"]
  B --> C["Create an agent<br>name + brain_url"]
  C --> D["Keep the sk_ key<br>shown once, at creation"]

Voqalize dials one URL, once per session:

{brain_url}?session_id={session_id}

The connection opens when the call starts and closes when it ends. Frames are protobuf and the contract is published — the wire is 497 lines of it. In Python that route is a Brain subclass and a call to run_session; the FastAPI example is a working one in about sixty lines.

Your route needs to be reachable from us. When it cannot be — a laptop, a VPC with no ingress, a serverless function — Cortex has your brain dial out instead, and the brain code is unchanged.

2. Behind the socket, whatever you already run

Section titled “2. Behind the socket, whatever you already run”

The brain receives finalized text and returns the words to speak. What produces those words is yours: a model call, an agent framework, a state machine, a lookup table. GeminiBrain ships in the SDK for the case where you want a model loop already wired, and the eleven demo brains are readable source.

Your model, your prompts, your tools and your retrieval stay in the process you deploy today.

3. Drive it before there is a call to make

Section titled “3. Drive it before there is a call to make”

The conformance harness opens a real socket, speaks real protobuf, and drives your brain through the scenarios Voqalize produces — barge-in, idle, an abandoned turn — with no microphone and no account. This is the loop to be in while you are writing the brain; steps 4 onward are for when it answers correctly. See testing a brain.

Two fields matter:

create_agent(tenant, name, brain_url="https://…/voice")
→ { agent, session_key } # sk_… , shown once

Over the MCP server from your editor, or the console. Keep the agent_id and the sk_ key: every credential names one agent, and the raw key is stored only as a hash.

The agent record says where the brain lives and whether recording defaults to on. It carries no voice, language, recognizer or idle settings — those depend on this caller. Our own lead-qualification brain reads a state from the enquiry form and answers a caller in Tamil Nadu in Tamil, which is a fact that does not exist until the call starts. Step 6 holds that call’s configuration.

sequenceDiagram
  participant P as Your page
  participant S as Your server
  participant V as Voqalize
  participant B as Your brain

  P->>S: start a call (your own auth)
  S->>V: POST /sessions.connect — agent_id, config, init
  V-->>S: webrtc_request_params, session_id
  S-->>P: that body, verbatim
  P->>V: client.connect(params)
  Note over P,V: WebRTC, straight to the machine running the call
  V->>B: WebSocket, one per session
  V->>B: SessionStart, carrying init
  B-->>V: a greeting, as a string
  V-->>P: the greeting, spoken
  loop every turn
    P->>V: the caller speaks
    V->>B: finalized text
    B-->>V: speech units, and actions
    V-->>P: audio, and RTVI on the data channel
    P->>V: a click, a form, a client message
    V->>B: the same message, verbatim
  end
  B-->>V: session.end()
  V->>B: Finalize, with what was heard
  V->>B: End

The browser asks your backend for a call. How that request is authenticated is entirely yours — a session cookie, a bearer token, whatever your app already does. We never see it.

There is a second path where a publishable pk_ key sits in the page and the browser calls us directly, which suits a demo or a public page. On that path the browser may send the same tts, stt and idle configuration shown below; record: false is accepted and record: true is refused. The handshake covers both. The rest of this page follows the server path, because it is the one where you decide who may start a call.

One request, holding two named things:

POST https://app.voqalize.com/api/v1/sessions.connect
Authorization: Bearer sk_live_…
Content-Type: application/json
{
"agent_id": "agt_…",
"config": {
"tts": { "voice": "VOICE_OMNIVOICE_GAURI", "language": "LANGUAGE_TA" },
"stt": { "language": "LANGUAGE_TA" },
"idle": { "timeout_ms": 8000 },
"record": false
},
"init": { "order_id": "A-1183", "tier": "gold" }
}

idle.timeout_ms defaults to 0, which is offon_user_idle never fires until something sets a timeout, because a nudge nobody asked for talks over a caller who was thinking.

config is how this call sounds and listens. It is the same Config the brain sends mid-call, parsed as proto3 JSON — enum members are their names, and a name we do not serve is refused at mint with the field pointed at. Send only what you want moved; both legs of a language change travel together, because moving one leaves the call listening in a language it is answering out of. See the catalog for what is on the roster and why there is no provider slot for the question underneath it.

record rides beside the three sections and stays out of the wire Config, because its lifetime is different: tts, stt and idle move any time, and recording is decided once, here. A pk_ key may turn it off and may not turn it on — recordings says why.

init is what your brain gets and nobody else reads. It arrives at session.init under that exact name, uninterpreted by everything in between: the account this caller is signed into, the order they are asking about, the plan they are on. It is stored on the session record, so send identifiers rather than personal data.

Both are optional. A request with an agent_id and nothing else starts a call in English on both legs.

Two more keys exist, and both are labels. display_name is what the console shows instead of an id, and metadata is a flat string→string map — at most 10 keys, values at most 256 characters — for the handful of things you want to recognise a call by later: a support ticket, a build, an experiment arm. It is not queryable and it is not storage; that is your own database’s job, and the cap is where the line is drawn. Anything the brain needs to act on goes in init.

The answer is what a pipecat transport connects with:

{
"webrtc_request_params": {
"endpoint": "https://…/webrtc",
"headers": { "Authorization": "Bearer <session token>" }
},
"session_id": ""
}

Return it to the browser unchanged and hand it to connect. A pipecat page forwards this body rather than reading it — startBotAndConnect is literally connect(await startBot(params)) — which is why the response is these two keys and no session record.

Two things to know before the first call: the endpoint is one machine, chosen when the session is minted, so it cannot be a constant in your page; and headers has to be a real Headers object today, one line in your page, for a reason the handshake writes down.

The media is direct UDP from the browser to that machine. Nothing of ours proxies the audio.

Voqalize builds the pipeline, dials your brain, and sends SessionStart with init on it. Your brain returns a greeting.

The caller is connected and hearing silence while that call returns, which is why a greeting is a string — a fixed line, or a template over what init carried. A model call here is a second and a half of nothing, at the one moment a caller has no idea whether the call is working.

on_session_start runs alongside it, and it is where a brain sets the voice and language for this caller with session.configure(...).

The caller speaks. Voqalize decides when they have finished and hands your brain the finalized text.

Your brain yields speech in units. Speaking starts on the first unit, so the reply begins before it has finished being generated, and each unit is one thing the caller can be interrupted out of. When they do interrupt, the audio stops and the in-flight turn is cancelled.

At the end of each unit your brain is told what the caller actually heard. A reply that generated three sentences and was cut after one is remembered as one — that reconciliation is the brain’s job, and the SDK keeps no history for you. Interruption and heard truth is the long version.

The same RTVI data channel carries the transcript, so both of these land while audio is still flowing.

Brain to page: session.dispatch(SomeAction(...)) sends a typed message your page receives as an event — a form to open, a row to highlight, a total to update. A number the caller has to hold in their head is a number that belongs on the screen.

Page to brain: pipecat’s own client methods send back what the caller clicked or typed, and it arrives at on_rtvi while the floor stays where it was. That callback cannot speak, so an agent cannot talk over the person who just clicked.

Ten message types cross, five each way. The RTVI plane is the list, and says which of them are pipecat’s rather than ours.

Either side ends it: session.end() from the brain, or the caller hangs up. on_session_end is where you write your own record of what happened.

Ours is readable back through the MCP server or the API — the session log with its wire frames, the recording if you asked for one, and usage.

Three levels, and the later one wins:

  1. Voqalize’s defaults — English on both legs.
  2. The session creator’s config, at connect. This may be your server with an sk_, or a browser holding a publishable pk_.
  3. The brain, with session.configure(Config(...)) — at session start, or any time during the call.

The brain always has the last word, because on_session_start runs after connect. Levels 2 and 3 are the same message seen from two sides: the session creator sets the call up immediately before it starts, and the brain moves it knowing how the conversation is going.

None of this is needed for a first call, and each has a page:

  • Cortex — your brain dials out, when it cannot accept inbound connections.
  • Recording — per agent as a default, per session as a decision.
  • The avatar — a talking head in the page, driven by the same session. The processor is already in every pipeline, so this is a browser-side change.
  • Idle detectionidle.timeout_ms hands the brain the floor after silence, and 0 turns it off.
  • Voice and language — two personas, English and 22 Indic languages.
  • The MCP server — agents, keys and sessions from inside your editor.