Live voice
A real-time voice conversation, like a phone call: your app streams the microphone to viora, and viora answers out loud while it's still writing the answer.
viora does the hard parts on the server:
- Turn taking. It notices when the person starts talking and when they've finished. There's no button to hold.
- Speech to text in English, Swahili and about 100 other languages.
- Grounded answers from your information, in short spoken-style sentences.
- Speech. Each sentence is spoken as soon as it's written, by the natural voice for its language: English (Aura-2) or Swahili (Microsoft neural, female or male). Swahili speech arrives as MP3 for clients that list
"mp3"inhello.formats(the JavaScript SDK in browsers and the Python client do this for you); other clients get those sentences as text for the device's own voice. - Interruptions. When the person talks over viora, it stops and listens.
- Wake word (optional). With
"hey viora", viora ignores everything else it hears. Follow-up questions within 20 seconds don't need the wake word.
In a browser
The JavaScript SDK opens the microphone, plays viora's voice and handles interruptions:
import { Viora } from "https://viora-cloud.depriver-tech.workers.dev/sdk.js";
const viora = new Viora({ api: "https://viora-cloud.depriver-tech.workers.dev", site: "your-site" });
const live = viora.live({ wake: "hey viora" });
live.on("state", (e) => console.log(e.state)); // listening, hearing, thinking, speaking
live.on("transcript", (e) => console.log("you:", e.text));
live.on("token", (e) => console.log(e.text));
button.onclick = () => live.start(); // must start from a click or tap
Other methods: live.sendText("…") for a typed question in the same conversation, live.interrupt(), live.endTurn() for push-to-talk, and live.close(). The level event (0 to 1) tells you how loud the microphone is, for an animation.
In Python
import asyncio
from viora.live import LiveSession
async def main():
async with LiveSession("https://viora-cloud.depriver-tech.workers.dev", "your-site") as live:
asyncio.create_task(stream_microphone(live)) # await live.send_audio(frame), 16 kHz PCM
async for event in live.events():
if event["type"] == "transcript":
print("you:", event["text"])
elif event["type"] == "token":
print(event["text"], end="", flush=True)
elif event["type"] == "audio":
play(event["pcm"]) # 24 kHz PCM
asyncio.run(main())
For a ready-made assistant with microphone and speaker, run viora talk --site your-site.
Protocol
Open a WebSocket to wss://…/v1/sites/{site}/live. Audio frames are binary messages; everything else is JSON text.
| Audio | Format |
|---|---|
| Microphone, you send | 16-bit little-endian PCM, mono, 16 kHz. Frames of 20 to 100 ms. Keep sending, including silence. |
| Speech, viora sends | 16-bit little-endian PCM, mono, 24 kHz. Play the frames in order as they arrive. |
You send
| Message | Meaning |
|---|---|
{"type": "hello", …} | Optional, first. speak: "server" (viora's voice, the default), "client" (your device speaks the speak events) or "none". Also voice (English voice: a name, female or male), voice_sw (Swahili: rehema, daudi, zuri, rafiki, female or male), wake, language ("en" or "sw" to always answer in one language), history and formats (["pcm", "mp3"] to receive Swahili speech). |
| binary | Microphone audio. |
{"type": "text", "text": "…"} | A typed question. |
{"type": "end_turn"} | Answer now, without waiting for a pause (push-to-talk). |
{"type": "interrupt"} | Stop the current answer. |
{"type": "bye"} | End the conversation. |
viora sends
| Message | Meaning |
|---|---|
ready | Connected. Includes the audio formats and the assistant's name. |
speech_started, speech_ended | The person started or stopped talking. On speech_started, stop playing viora's voice. |
transcript | What the person said: text, language. ignored: true means the wake word was missing. |
wake | The wake word was said on its own. viora is listening for the question. |
token | Answer text as it's written. |
speak | A sentence about to be spoken: text, language. audio: true means viora's speech follows as binary frames; otherwise speak it with the device's voice. |
| binary | viora's speech. |
done | The answer is complete: text, answered, language, sources, handoff (a WhatsApp link when viora couldn't answer). |
interrupted, audio_clear | The answer was stopped. Drop any speech you haven't played yet. |
error | message to show the person. code: "quota" means the daily limit was reached. |
bye | The session ended. reason: idle after 2 minutes without audio, time_limit after 20 minutes, or client. |
Echo and interruptions
When viora's voice plays through a speaker, the microphone hears it too. Browsers remove most of that echo, and viora needs clearly louder speech before it treats a sound as an interruption. Without echo cancellation, for example a Python program on a laptop, use headphones, or stop sending microphone audio while viora speaks. viora talk does this unless you pass --barge-in.
Limits
- Each spoken or typed question counts toward your plan's daily total.
- 30 live sessions per visitor per hour, up to 20 minutes each.
- viora's natural English voice is included on paid plans: 300 answers a month on Starter and 1,000 on Business. On the Free plan, and once the month's allowance is used,
speakevents arrive withaudio: falseand the SDK and widget speak them with the device's own voice. Thereadyevent'sspeech_languageslists the languages viora can voice ("sw"on every plan). Aspeakevent withaudio: truealso says theformatof the binary that follows:"pcm"(16-bit, 24 kHz, in chunks) or"mp3"(one file per sentence). - Self-hosted: live voice needs a speech-to-text provider (
VIORA_VOICE). Spoken answers need a text-to-speech provider (VIORA_SPEECH:cloudflareoropenai). Without one, viora sendsspeakevents for the client to voice.