viora by INCPRITECH

Live voice

A real-time voice conversation, like a phone call: your app streams the microphone to viora, and viora answers out loud while it's still writing the answer.

viora does the hard parts on the server:

  • Turn taking. It notices when the person starts talking and when they've finished. There's no button to hold.
  • Speech to text in English, Swahili and about 100 other languages.
  • Grounded answers from your information, in short spoken-style sentences.
  • Speech. Each sentence is spoken as soon as it's written, by the natural voice for its language: English (Aura-2) or Swahili (Microsoft neural, female or male). Swahili speech arrives as MP3 for clients that list "mp3" in hello.formats (the JavaScript SDK in browsers and the Python client do this for you); other clients get those sentences as text for the device's own voice.
  • Interruptions. When the person talks over viora, it stops and listens.
  • Wake word (optional). With "hey viora", viora ignores everything else it hears. Follow-up questions within 20 seconds don't need the wake word.

In a browser

The JavaScript SDK opens the microphone, plays viora's voice and handles interruptions:

import { Viora } from "https://viora-cloud.depriver-tech.workers.dev/sdk.js";

const viora = new Viora({ api: "https://viora-cloud.depriver-tech.workers.dev", site: "your-site" });
const live = viora.live({ wake: "hey viora" });

live.on("state", (e) => console.log(e.state));          // listening, hearing, thinking, speaking
live.on("transcript", (e) => console.log("you:", e.text));
live.on("token", (e) => console.log(e.text));

button.onclick = () => live.start();                      // must start from a click or tap

Other methods: live.sendText("…") for a typed question in the same conversation, live.interrupt(), live.endTurn() for push-to-talk, and live.close(). The level event (0 to 1) tells you how loud the microphone is, for an animation.

In Python

import asyncio
from viora.live import LiveSession

async def main():
    async with LiveSession("https://viora-cloud.depriver-tech.workers.dev", "your-site") as live:
        asyncio.create_task(stream_microphone(live))     # await live.send_audio(frame), 16 kHz PCM
        async for event in live.events():
            if event["type"] == "transcript":
                print("you:", event["text"])
            elif event["type"] == "token":
                print(event["text"], end="", flush=True)
            elif event["type"] == "audio":
                play(event["pcm"])                         # 24 kHz PCM

asyncio.run(main())

For a ready-made assistant with microphone and speaker, run viora talk --site your-site.

Protocol

Open a WebSocket to wss://…/v1/sites/{site}/live. Audio frames are binary messages; everything else is JSON text.

AudioFormat
Microphone, you send16-bit little-endian PCM, mono, 16 kHz. Frames of 20 to 100 ms. Keep sending, including silence.
Speech, viora sends16-bit little-endian PCM, mono, 24 kHz. Play the frames in order as they arrive.

You send

MessageMeaning
{"type": "hello", …}Optional, first. speak: "server" (viora's voice, the default), "client" (your device speaks the speak events) or "none". Also voice (English voice: a name, female or male), voice_sw (Swahili: rehema, daudi, zuri, rafiki, female or male), wake, language ("en" or "sw" to always answer in one language), history and formats (["pcm", "mp3"] to receive Swahili speech).
binaryMicrophone audio.
{"type": "text", "text": "…"}A typed question.
{"type": "end_turn"}Answer now, without waiting for a pause (push-to-talk).
{"type": "interrupt"}Stop the current answer.
{"type": "bye"}End the conversation.

viora sends

MessageMeaning
readyConnected. Includes the audio formats and the assistant's name.
speech_started, speech_endedThe person started or stopped talking. On speech_started, stop playing viora's voice.
transcriptWhat the person said: text, language. ignored: true means the wake word was missing.
wakeThe wake word was said on its own. viora is listening for the question.
tokenAnswer text as it's written.
speakA sentence about to be spoken: text, language. audio: true means viora's speech follows as binary frames; otherwise speak it with the device's voice.
binaryviora's speech.
doneThe answer is complete: text, answered, language, sources, handoff (a WhatsApp link when viora couldn't answer).
interrupted, audio_clearThe answer was stopped. Drop any speech you haven't played yet.
errormessage to show the person. code: "quota" means the daily limit was reached.
byeThe session ended. reason: idle after 2 minutes without audio, time_limit after 20 minutes, or client.

Echo and interruptions

When viora's voice plays through a speaker, the microphone hears it too. Browsers remove most of that echo, and viora needs clearly louder speech before it treats a sound as an interruption. Without echo cancellation, for example a Python program on a laptop, use headphones, or stop sending microphone audio while viora speaks. viora talk does this unless you pass --barge-in.

Limits

  • Each spoken or typed question counts toward your plan's daily total.
  • 30 live sessions per visitor per hour, up to 20 minutes each.
  • viora's natural English voice is included on paid plans: 300 answers a month on Starter and 1,000 on Business. On the Free plan, and once the month's allowance is used, speak events arrive with audio: false and the SDK and widget speak them with the device's own voice. The ready event's speech_languages lists the languages viora can voice ("sw" on every plan). A speak event with audio: true also says the format of the binary that follows: "pcm" (16-bit, 24 kHz, in chunks) or "mp3" (one file per sentence).
  • Self-hosted: live voice needs a speech-to-text provider (VIORA_VOICE). Spoken answers need a text-to-speech provider (VIORA_SPEECH: cloudflare or openai). Without one, viora sends speak events for the client to voice.