Skip to main content
Voice turns a session into a natural, full-duplex spoken conversation. GPT-Live handles listening, speaking, and interruptions while the session handles questions and work with its selected model, tools, context, and durable transcript. You can move between voice and text without switching to a separate voice-only agent.

Enabling voice

Voice uses OpenAI GPT-Live-1 and needs an OpenAI API key from a project with GPT-Live access. A deployment admin can enter one from Integrations > Voice, or a self-hosted operator can set R_VOICE_OPENAI_API_KEY; the environment variable is used when both exist, and an admin can still turn voice off for the deployment from the same card. Voice is opt-in: the deployment’s general OPENAI_API_KEY is not used, so enabling OpenAI for task inference does not turn voice on. See the Voice integration page for details. The voice integration settings include the GPT-Live voice selection and an AI-generated audio preview. Existing selections are preserved. New and unset connections default to Marin, one of OpenAI’s recommended voices. When the voice key is not configured the voice button does not appear. The key stays on the control plane. The browser sends its WebRTC connection offer to Roomote and receives only the negotiated session answer; it never receives the API key. On Roomote Cloud, each user must accept the experimental voice data flow before their first call. The confirmation explains that microphone audio, voice transcripts, and workspace context such as repository and integration names are sent to OpenAI. Roomote saves acceptance for that user across browsers and sessions. Canceling or dismissing the confirmation does not start microphone capture or contact OpenAI. Self-hosted deployments keep their existing voice flow and do not show this Cloud consent step.

Using voice

  1. Select the voice button in a composer: in an open session, or on the home page and the New Session dialog. From the home page or dialog a new session is created and the call starts inside it. When Private is selected, the new session and its transcript remain owner-only under the Private Sessions boundary; the voice data described above is still processed by OpenAI.
  2. Grant microphone access when the browser asks. A short rising tone confirms the call is open; a falling tone marks the end. A Call started marker appears in the session.
  3. Talk to Roomote the way you would on a phone call. It acknowledges each request in a few words, hands the work to the session, and reports the result out loud when it lands. Greetings, thanks, and small talk are answered directly without starting session work.
  4. Speak at any time to interrupt. Roomote keeps listening while it speaks, and follow-ups go back through the same session. You can also type in the composer during the call.
  5. Use the in-call controls to mute your microphone, silence Roomote’s audio without muting yourself, or end the call. The button stays highlighted while the call is active, and a Call ended marker records its length.
Voice input requires a browser with microphone and WebRTC support, which includes current Chrome, Edge, Safari, and Firefox.

How the transcript works

A voice call is transcribed into the session as the record of what was said. Your speech appears as your messages, and what Roomote said out loud appears as its replies. The session’s work, such as tool calls, launched tasks, and reports, appears between those turns exactly as it does in a typed session, so the timeline shows both the conversation and the work behind it. During a call the session returns its results to the voice rather than writing them as chat replies; Roomote then reports them in its own words, keeping numbers, names, paths, and link labels exact. The exact result stays in the transcript as a collapsed Reported result to voice row, including when the call drops or speech is interrupted. Typed messages sent during a call are answered in writing as usual. Each spoken request is cleaned up (filler words, false starts, and misheard terms) by the deployment’s helper model before it reaches the session. GPT-Live is told which repositories, environments, and integrations the session can reach, so it recognises their names, and the same names guide the cleanup so a misheard repository name is corrected to the real one. Ending the call stops the microphone; work already started remains visible in the session and follows the normal session lifecycle.