Web and apps
Talking to the AI from a browser or a mobile app: creating a session, client_token and WebRTC, audio versus text output, and pairing voicast with avacast.
Browsers and mobile apps connect to the voicast voice server over WebRTC. The API key stays on your server; the browser or app gets only a single-use client_token.
Your server ── POST /v1/sessions with the API key ──→ voicast
│ passes only client_token, webrtc_url and ice_servers
Browser / app ── WebRTC offer with client_token ──→ voicast voice server1. Create a session (server)#
curl https://api.voicast.jp/v1/sessions \
-H "Authorization: Bearer $VOICAST_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "tenant_id": "ten_…", "output": "audio" }'{
"id": "sess_…",
"tenant_id": "ten_…",
"agent_id": "agt_…",
"output": "audio",
"client_token": "vct_…",
"expires_at": 1790000120000,
"webrtc_url": "https://voice.voicast.jp/v1/webrtc/offer",
"ice_servers": [{ "urls": "stun:stun.cloudflare.com:3478" }]
}| Input | Required | Description |
|---|---|---|
tenant_id |
yes | Tenant ID |
agent_id |
Agent ID. Defaults to the tenant's agent | |
output |
audio (the AI speaks; default) or text (reply text only) |
client_tokenexpires after 2 minutes and works once. To reconnect, create a new session.- The agent's latest version at connect time is used.
- Pass
ice_serverstoRTCPeerConnection'siceServersas is. Besides STUN, it includes TURN with short-lived credentials (username,credential) when voicast provides TURN, which helps on networks that block UDP. - When new calls are refused for billing reasons, you get 402 (Errors and limits). A disabled tenant returns 409
tenant_disabled; if no agent can be determined you get 422no_agent.
2. Connect (browser or app)#
With the SDK#
voicast/web in the SDK handles the microphone, the WebRTC connection and the events.
import { VoiceSession } from 'voicast/web';
// Get client_token, webrtc_url and ice_servers from your own server
const session = new VoiceSession({ clientToken, webrtcUrl, iceServers });
session.on('user-transcript', ({ text, final }) => showUser(text, final));
session.on('bot-output', ({ text }) => showBot(text));
session.on('closed', ({ reason }) => console.log('closed', reason));
await session.start(); // asks for the microphone, then connects
hangupButton.onclick = () => session.stop();Browsers may only play audio in response to a user action, so call start() from a click.
With a Pipecat client#
The voice server speaks the same protocol as Pipecat's SmallWebRTC. With a Pipecat client (JavaScript, iOS, Android) and its SmallWebRTC transport, use webrtc_url as the connection URL and pass { token: client_token } as requestData.
With your own WebRTC code#
Any WebRTC stack can connect like this:
- Create an
RTCPeerConnectionwith your microphone track and a data channel (any name, e.g.chat) - Create an offer, wait for ICE gathering to finish, and POST
{ sdp, type, token }towebrtc_url - Set the returned
{ sdp, type, pc_id }as the answer - Send ICE candidates found later with
PATCH webrtc_urland{ pc_id, candidates: [{ candidate, sdp_mid, sdp_mline_index }] } - When the data channel opens, send an RTVI
client-readymessage, then send apingstring every second (the voice server treats a few seconds without pings as a disconnect)
const pc = new RTCPeerConnection({ iceServers }); // the session's ice_servers
const mic = await navigator.mediaDevices.getUserMedia({ audio: { echoCancellation: true, noiseSuppression: true } });
pc.addTransceiver(mic.getAudioTracks()[0], { direction: 'sendrecv' });
pc.ontrack = (e) => { audioEl.srcObject = e.streams[0]; audioEl.play(); };
const dc = pc.createDataChannel('chat', { ordered: true });
dc.onopen = () => {
dc.send(JSON.stringify({ label: 'rtvi-ai', type: 'client-ready', id: crypto.randomUUID(), data: { version: '1.0.0' } }));
setInterval(() => dc.readyState === 'open' && dc.send(`ping: ${Date.now()}`), 1000);
};
dc.onmessage = (e) => {
const msg = JSON.parse(e.data);
if (msg.label === 'rtvi-ai') onRtvi(msg); // see "What you receive"
};
await pc.setLocalDescription(await pc.createOffer());
await iceGatheringComplete(pc);
const r = await fetch(webrtcUrl, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ sdp: pc.localDescription.sdp, type: pc.localDescription.type, token: clientToken }),
});
if (!r.ok) throw new Error(`cannot connect (${r.status})`); // 403: token expired or already used
const answer = await r.json();
await pc.setRemoteDescription({ type: answer.type, sdp: answer.sdp });If the voice server asks to renegotiate, the data channel receives { type: 'signalling', message: { type: 'renegotiate' } }. POST a new offer with pc_id and restart_pc: false to the same webrtc_url.
What you receive#
With output: 'audio', the AI's voice arrives as an audio track. The data channel carries RTVI messages (label: 'rtvi-ai').
type |
Contents |
|---|---|
bot-ready |
The AI is ready |
user-transcription |
What the user said (data.text, data.final) |
bot-output |
A sentence the AI spoke (with output: 'audio') |
bot-started-speaking, bot-stopped-speaking |
The AI started or stopped speaking |
server-message |
voicast messages (below) |
Text output#
A session with output: 'text' does not speak. Each sentence of the reply arrives as a server-message. Use it to drive an avatar (such as avacast) or your own speech synthesis, or to show the reply as text.
type VoicastMessage =
| { type: 'voicast.reply'; text: string } // one sentence of the reply (greeting and closing too)
| { type: 'voicast.interrupted' } // the user started talking: stop the current reply
| { type: 'voicast.end'; outcome: string }; // the conversation is over (the connection closes next)The user's voice is always sent as the microphone track, whatever the output.
Pairing with avacast#
avacast animates an avatar made from a photo in real time. Let voicast listen and write the replies, and avacast give them a face and a voice, and you have an on-screen receptionist.
import { VoiceSession } from 'voicast/web';
import { AvacastSession } from 'https://avacast.jp/sdk/v1.js';
// Create both tokens on your server (the voicast session uses output: 'text')
const { voicast, avacast } = await fetch('/api/start-reception', { method: 'POST' }).then((r) => r.json());
const avatar = new AvacastSession(avacast.client_token, {
signalingUrl: avacast.webrtc.signaling_url,
iceServers: avacast.webrtc.ice_servers,
});
avatar.attach(document.getElementById('avatar'));
await avatar.start();
const talk = new VoiceSession({
clientToken: voicast.client_token,
webrtcUrl: voicast.webrtc_url,
iceServers: voicast.ice_servers,
});
talk.on('reply', ({ text }) => avatar.speak(text)); // the avatar says voicast's reply
talk.on('interrupted', () => avatar.interrupt()); // stop when the user starts talking
talk.on('end', ({ outcome }) => showResult(outcome));
await talk.start();On your server, create the voicast session (POST /v1/sessions with output: 'text') and the avacast session side by side, and return only the two tokens to the page.
Results#
Web and app calls produce the same call records and webhooks as phone calls. They have channel: 'web', and the session id (sess_…) is the call's provider_call_id.
curl https://api.voicast.jp/v1/calls/by-provider/sess_… \
-H "Authorization: Bearer $VOICAST_API_KEY"Notes#
- If the AI's voice leaks back into the user's microphone, the AI may react to itself. Keep the browser's echo cancellation on (
echoCancellation: true); headphones help even more. - Networks that block UDP (some corporate networks) may prevent the connection. Always pass the session's
ice_servers(with TURN in it, the connection goes through TURN). - A refusal at the offer stage returns 403 (expired or used token, billing reasons and so on). Billing refusals will later return 402, so treat both 403 and 402 as "cannot connect".