Skip to main content
Models decide how Roomote thinks through a task. The sandbox provider gives a Roomote agent a sandbox to work in. The inference provider gives it access to the model that reads the prompt, reasons about the workspace, writes code, runs tools, reviews output, and explains what changed. Configure models from Settings > Models.

Inference providers

An inference provider is the service that hosts or routes model calls. Roomote supports direct model APIs, multi-provider gateways, coding subscriptions, and self-hosted OpenAI-compatible endpoints. You can connect more than one inference provider in the same deployment. That lets you mix and match models by provider instead of betting the whole deployment on one account, one vendor, or one model family.

Managed Roomote inference

Some hosting deployments provision Roomote inference with a limited, spend-capped credit grant during setup. It is separate from your own provider connections: in particular, you can add an OpenRouter key in Settings > Models even when Roomote inference is active. If the hosting deployment does not offer it, the option is not shown. The Roomote Cloud trial includes $5 of inference, enough to complete several tasks before you connect your own provider. Trial token usage and estimated model cost appear in task details and Cost Analytics, while the remaining credit line reflects the hosted grant’s authoritative limit. Its key is stored with your other provider credentials, so deleting the Roomote inference provider in Settings > Models disables it permanently. For example, a deployment might use:
  • an OpenRouter-routed model for the default coding model
  • a direct Anthropic or OpenAI model for planning or review
  • a lower-cost provider model for helper work
  • a vision-capable model only when visual inspection is needed
Connecting an API-key provider that has no configured models yet automatically adds a short list of recommended models for it — enabled and including the provider’s default model — so you land on a usable model list right away, both in the setup wizard and from Settings > Models. You stay in control: you can disable or remove any of the added models, add more later, and reconnecting a provider never re-adds models you removed. When you connect or update an API-key provider, Roomote first checks the submitted credentials with a small live model request. If the provider rejects the key, nothing is saved: fix the credential and save again. An account with no remaining credits or quota is still saved, with a warning, so you can add credits while you finish setting up. Other failures (a rate limit, a model your account cannot access, or a transient provider outage) never block the save. Self-hosted OpenAI-compatible endpoint connections are checked by listing the endpoint’s models instead. Saving again without changing any credential skips the live check.

API and gateway providers

These connections use metered API billing or a provider-managed gateway:

Subscription providers

Subscription connections consume an included plan allowance instead of a general API balance: Roomote shows plan usage for subscription providers when the provider returns usable quota data. These usage endpoints are not always documented or stable, so a missing usage line does not by itself mean the connection is broken.

Self-hosted and gateway providers

OpenAI-compatible endpoints let you bring any server that speaks the OpenAI /v1 API — including LiteLLM, Ollama, vLLM, and custom proxies. Models are discovered from the configured endpoint instead of Roomote’s recommended catalog. Connect the provider in Settings > Models, then select from the discovered models and choose the defaults and role mappings that fit your deployment. These endpoints must be reachable from the Roomote deployment, not merely from your laptop or an individual task sandbox. See the provider page for the expected URL, network, security, and cost behavior. The recommended set is a single curated list of models that ships with each Roomote release, so it is predictable for a given version. Every catalog-backed provider draws from the same list: a provider offers the subset it serves, under its own model ids, with the same names everywhere. Endpoint providers that discover models dynamically use the model list returned by their endpoint. Available Models always lists the full recommended set for every connected catalog-backed provider: recommendations you have not enabled appear toggled off, and they cannot be deleted while their provider stays connected — turn a model off to stop using it. To go beyond the recommended set, add any model by its slug from the add-model field. Claude Sonnet 5.5 is the current curated Claude recommendation where a connected provider exposes its verified route. New connections can discover it through supported direct and gateway providers, while existing Claude Sonnet 5 selections remain usable until you choose to switch. Claude Haiku 5.5 is recommended through Roomote inference, OpenRouter, Vercel AI Gateway, Requesty, Azure OpenAI, Azure AI Foundry, Anthropic, OpenCode Zen, OpenCode Go, and Amazon Bedrock. Enable it in Available Models, then assign it to a role or select it for a session or task. The Anthropic helper and explore recommendations and Amazon Bedrock’s Recommended mapping preset now use Haiku 5.5, which supports adaptive thinking. Existing saved Haiku 4.5 selections stay unchanged; GitHub Copilot retains its verified Haiku 4.5 recommendation. Model access still depends on the connected provider account, region, or plan. GPT-6 Astra is also recommended through Vercel AI Gateway, GitHub Copilot, and OpenCode Zen. Connect the provider, enable Astra in Available Models, and assign it to a role or select it for a session or task. Adding these routes does not change existing model defaults; access still depends on the connected provider account or plan. GPT-6.1 Sol is available through the same curated provider routes as GPT-6 Sol. OpenAI describes it as near-Astra performance at a lower cost. Existing defaults and saved mappings stay unchanged; differentiated recommendations that previously used GPT-6 Sol now use GPT-6.1 Sol. DeepSeek V4.1 Flash is recommended through DeepSeek, OpenRouter, Vercel AI Gateway, and OpenCode Go. Roomote uses the provider’s matching route and no longer recommends the older V4 Flash 0731 model. Existing defaults do not change automatically; enable V4.1 Flash and assign it to a role when you want to use it. V4.1 Flash has three distinct reasoning levels: Low, High, and Max. Its provider maps a saved Medium or X-High request to High. The model picker shows that effective level consistently in settings and chat. Viewing the picker leaves saved defaults and inherited session settings intact; choosing a different stop saves that level as an explicit selection. GPT 5.6 Luna is the default coding model for Roomote’s default OpenRouter setup and for new Vercel AI Gateway, Requesty, OpenAI, Azure OpenAI, Azure AI Foundry, OpenCode Zen, OpenCode Go, Amazon Bedrock, GitHub Copilot, and ChatGPT subscription connections. Where a provider offers a Default mapping preset, it uses that provider’s GPT 5.6 Luna route for every role except orchestration, which follows the coding model, and uses medium reasoning for coding. Existing saved defaults and role mappings do not change automatically. Providers can also offer differentiated presets for the model roles below — for example a strong model for planning and code review and a fast, low-cost model for helper and explore work. The Recommended OpenAI and ChatGPT subscription presets use GPT-6.1 Sol for coding, while the GitHub Copilot preset uses GPT-6 Luna. Connecting one of these providers in the setup wizard applies its Default preset; choose a differentiated preset later when you want to split roles across models. From Settings > Models, use Use a mapping preset on the Model mapping card to preview and apply any preset offered by a connected provider. Confirm the preset to set the role selections and enable any missing models. It is an apply-once action, not a lock: you can adjust every role afterward. Roles a provider has no specific recommendation for are set to Same as coding model, and roles managed by environment variables are left untouched. You can also choose Add your own to save a private preset with a model and reasoning level for each role when that model supports reasoning. Your custom presets are available only to your account. They remain listed if a model later becomes unavailable, but cannot be applied until every unavailable role has a current model selection; deleting a preset does not change the active model mapping.

Env-based setup

Most deployments should configure providers from Settings > Models. Use environment variables when provider credentials are managed by your hosting platform, secret manager, or local development shell. At minimum, set a default coding model and the matching provider key:
The provider is the first segment of the model ID. Direct-provider access uses the provider’s normal key:
You can also split model roles with env vars:
Roomote automatically makes common provider configuration available to the model runtime, including OpenRouter, Requesty, Vercel AI Gateway, OpenAI, Azure OpenAI, Azure AI Foundry, Anthropic, Google Gemini, Moonshot, Kimi for Coding, MiniMax, Z.AI, OpenCode, Amazon Bedrock, xAI, and GitHub Copilot. For gateway-supported providers, credentials stay on the control plane and model requests are proxied instead of forwarding the credentials into task workers. Use R_MODEL_ENV_KEYS when a provider key uses a custom env var name and must be forwarded:
See Environment Variables for the full supported key list.

What Models settings controls

Settings > Models has two layers:
  • Inference Providers stores the provider credentials Roomote can use.
  • Models controls the provider/model pairs that are available, the default model, and specialized model roles.
When you connect an inference provider in setup, Roomote tests the candidate credentials and selected model before saving them. Authentication, quota, model-access, endpoint, and timeout failures are reported during setup so you can correct the connection before starting a task. Providers configured through environment variables are validated by the deployment at startup instead. For subscription-based providers (ChatGPT, GitHub Copilot, Kimi for Coding, xAI Grok / SuperGrok, and Z.AI Coding Plan), the provider row also shows current plan usage when the provider reports it, such as remaining legacy Copilot premium-request quota, the percent used of the ChatGPT 5-hour and weekly limits, or Grok included-usage windows. These numbers come from unofficial provider endpoints, so the usage line is hidden whenever the provider does not return usable data. Under Automations > Inference Provider Usage Alerts, admins can configure hourly checks for queryable limits from ChatGPT subscription, GitHub Copilot, xAI Grok subscription, OpenRouter, Kimi for Coding, OpenCode Go, Z.AI, and Z.AI Coding Plan. Roomote posts one warning when a credential crosses its configured threshold in a limit window, plus a critical warning at 100%. Alerts can report through Slack or Discord destinations and supported primary conversations in Teams, Telegram, or Discord. The warning names the provider, credential fingerprint or label, limit window, and reported usage. A provider that does not return usable quota data is skipped. xAI can be connected with an API key, a SuperGrok or eligible X Premium+ subscription, or both. When both are configured, the subscription is preferred at runtime for xai/ models. Admins can enable or disable models from the task model list. The default model must stay enabled, because Roomote uses it when a task does not request a specific model. Model settings affect new task starts. Running tasks keep their selected model unless you change it or a configured fallback switches it after a provider failure. Resumed tasks also keep a saved model when it is still registered and its provider is connected, unless a usable task override applies. If the saved model is missing or its provider is disconnected, Roomote resumes with the deployment default instead. If a preserved saved model is reported as unavailable when work resumes, Roomote retries once with the deployment default.

Model roles

Roomote can use different models for different parts of the system. You can leave these roles on the default model at first, then split them when you know where you want more speed, quality, or cost control. You do not need a separate model for every role. Set a specialized role to Same as coding model to inherit the default. Many teams start with one strong default model, then split out a faster orchestration or helper model, or a stronger review model, after they can see real usage. The Vision model handles image attachments when the active model cannot inspect them. The separate Audio and video model handles audio transcription and video descriptions. Each can be configured independently in Settings > Models. Roomote warns when catalog metadata shows that the selected media model does not support audio or video. The audio and video model defaults to Same as coding model, inheriting the coding model when no separate model is selected. The task model switcher also includes an Audio and video role override for media processing owned by that task. Changes apply from the task’s next message. Automatic chat-attachment transcription from Slack, Discord, Telegram, and Microsoft Teams uses the deployment audio and video model instead of a per-task override.

Fallback models

Admins can turn on Fallback config in Settings > Models and choose one enabled fallback model and reasoning level for each role. A fallback must differ from that role’s effective default. Mapping presets change only the Default rows; they do not replace saved fallbacks. Roomote switches immediately when the active provider runs out of credits or quota, rejects credentials with a 401 or 403 response, or reports the model as unavailable with a 404 response. Other retryable provider failures, including rate limits and transient server errors, switch after three Roomote-level attempts. Safety refusals, context-length failures, output-length failures, and user aborts do not trigger a fallback. An HTML gateway or WAF 403 is treated as a transient provider failure rather than a credential rejection. After a switch, the task or session stays on the fallback for the rest of its life, including snapshot resumes. New tasks and sessions start on their defaults. Choosing another model manually takes precedence and re-arms fallback protection for that model. The transcript shows an expandable fallback row, and work started from a communications provider also receives one notice in its originating thread. Fallbacks cover task coding and review models, visual, explore, and advisor subagents, session orchestration, and Roomote’s helper calls for titles, summaries, and routing. OpenCode’s private helper calls inside a task sandbox, such as OpenCode-generated session titles, are not observable to Roomote and cannot trigger fallback. This error fallback is separate from Same as coding model, which only makes a role inherit the coding model before a request starts.

Coding model routing rules

Admins can add natural-language routing rules below Coding model in Settings > Models. Each rule chooses an enabled coding model, an optional reasoning level, and a condition such as Complex migrations and architecture work or Small documentation fixes where speed matters. When a new session delegates coding work, Roomote evaluates every saved rule independently of its list position. It applies only the single strongest clear match. An explicit model or reasoning choice in the request takes precedence; weak, ambiguous, failed, or unmatched decisions keep the deployment defaults. Rules that point to a model that is no longer enabled are ignored. Model routing is separate from environment routing: one chooses the model and reasoning level, while the other chooses the workspace that contains the relevant repositories and services. To verify a rule, start a new session with a request that clearly matches its condition and check the model shown on the delegated task.

Judgment model

The judgment model is optional. Many routing and triage steps are bounded decisions rather than text: whether a channel message should start work, which model a delegated task needs, or whether a thread reply is meant for Roomote. A judgment model answers these as typed questions with probabilities in one fast call. Roomote supports a judgment model it runs itself and TypeSafe’s Jev through these routes: Choose it under Judgment model in Settings > Models, or set R_JUDGMENT_MODEL to roomote, typesafe, openrouter, vercel, or off. A TypeSafe key selects Jev via TypeSafe when Settings has no saved route. Choose the Roomote judgment model, OpenRouter, or Vercel AI Gateway to route decisions through those providers. To compare Jev with the Roomote judgment model, select Jev and set R_JUDGMENT_SHADOW=on; the Roomote model scores the same decisions for calibration while Jev remains the selected route. Many deployments already have OpenRouter or Vercel AI Gateway keys for task inference, and selecting those routes sends decision text to that provider. Enable Judgement repository rules under Settings → Experimental to give coding tasks the Judgement skill and judgement command on their next run. The experiment is off by default and requires a configured Jev route and provider key. You can ask a Roomote agent to prevent a recurring repository mistake, such as “add a rule so we stop introducing this wording.” The skill helps it scope a rule, add labeled examples, and test the cutoff. Judgement uses your configured Jev provider; no additional task API key or skill installation is needed. The command supports check, test, calibrate, capture, and compare. Rules and examples live in the repository’s .judgement/ directory. Calibration makes inference requests; capturing examples and comparing saved reports do not. The skill is unavailable when a supported Jev route is not configured or setup fails. It does not automatically add rules or change which CI checks are required. For deployment-wide conversational preferences, use Agent Guidance. Enable Jevgrep code search under Settings → Experimental to give coding tasks Jevgrep (jg) and its agent skill for finding relevant code from natural-language questions. This experiment is off by default and requires Jev to be selected with its provider key configured. Roomote sets it up on the next task run; no separate login is needed. Searches send selected source excerpts to the selected Jev provider. If setup or evaluation is unavailable, the agent can use ordinary code search. Jevgrep is currently available with the three Jev routes above; selecting Off or the Roomote judgment model disables it for new task runs. Turning off the experiment also stops new evaluations from existing tasks. The agent, including OpenCode’s explore subagent, is instructed to make at most one attempt within a specific subsystem folder during task discovery, stop after 30 seconds, and fall back to ordinary search. Whole app or package source trees are too broad. Exact symbols and known paths use ordinary search directly. Jevgrep evaluations do not use the Roomote shadow evaluator. The Roomote judgment model is an open-weight model served on Roomote-operated infrastructure, where its decision text is processed. It answers the same typed questions as Jev and follows the same confidence rules for decisions in its evaluated policy. Self-hosted deployments can point R_JUDGMENT_UPSTREAM_URL at any server that speaks the same typed decisions request. When a judgment model is on, Roomote asks it first for:
  • Auto-respond to channels launch criteria, including whether a message repeats an incident that already launched
  • whether a new task request asks a question, wants a plan, or wants an implementation
  • whether an unmentioned reply in a Slack, Discord, or Teams thread is meant for Roomote after someone else spoke; it routes the reply only when confident, and otherwise Roomote still needs a mention
  • dropping inbound email that is clearly an automatic reply, such as an out-of-office or bounce notice
  • whether a settled session or task turn holds something worth saving to Memory that the agent did not record itself
Roomote applies a judgment when its confidence meets the configured threshold. When a decision needs more context, Roomote uses the configured Helper model role. The text a decision needs is sent to the selected judgment model, so choosing Jev sends it to TypeSafe or the gateway you picked; keys stay on the control plane and are never sent to task sandboxes. Some decisions run in bulk, such as the Memory check after every session and task turn. Roomote evaluates those high-volume workloads with a configured judgment model (Jev or the Roomote judgment model), using one typed evaluation for each check. To see how the judgment model answers, admins can open Test decisions under Judgment model in Settings > Models. It lists each decision Roomote asks, with the exact questions and a sample state to edit, and shows the answer’s probabilities and latency. When Jev answers and the Roomote judgment model is also configured, it can ask both side by side. Test decisions go to the same judgment model as real ones but are not captured or shadowed. For Repository Judgement, choose a rule and labeled example, select its initial or expanded evidence, and click Load example. Edit the inputs or ask for three repeated runs. Each answer shows the probability of a rule violation and whether it reaches the rule’s cutoff. Recent runs retain their original inputs and labels. Expected labels are never sent to the model. These are individual evidence requests. A partial screen cannot approve a file, and the tester’s 20-second timeout does not measure the commit hook’s three-second budget. Confirm a proposed improvement through the full Judgement checker before changing a rule or its violation cutoff.

Reasoning settings

Some models expose reasoning controls. Roomote lets admins set reasoning levels for the main model roles: Low, Medium, High, Extra high, or Max. Higher reasoning can improve planning, debugging, and review quality, but it can also increase latency and cost. Use it where deeper thinking changes the outcome, not everywhere by default. A practical starting point:
  • use Medium for the default coding model
  • use Low for orchestration, helper, media, and explore work unless you see quality issues
  • use High for code review and advisor work when you want more careful analysis
  • reserve Extra high or Max for models and workflows where the added cost is justified
If a model does not support reasoning controls, Roomote hides or ignores the reasoning selector for that role.

Per-task model switching

Deployment settings define the defaults, but each task can override them from the web task view. The model chip in the message composer shows the task’s current coding model and reasoning level; open it to switch the coding model, or expand All roles to override the planning, code review, explore, helper, or media role for that task only. Changes apply from the next message: a turn that is already running finishes on the old settings, and the next turn (including its sub-agents) uses the new ones. Overrides persist for the life of the task, including snapshot resumes, and Reset to defaults returns every role to the deployment configuration. When Roomote switches after a provider failure, the model chip and session composer show the fallback immediately; a manual selection overrides it. You can also just ask the agent — for example “switch to Fable 5 with max reasoning for the rest of this task”. The agent applies the same change through its task-management tool, subject to the same allowed-models list. Viewing a custom automation’s session does not grant control over its task model. Only the automation’s creator or a deployment administrator can change model selection for those tasks. Conversational sessions expose the same model and reasoning choices in their composer. New sessions show the deployment’s orchestration model by default; direct environment and repository tasks show the coding model. Leaving the session selection untouched keeps deployment-default behavior, while an explicit selection is saved before the next message can be sent, persists across page refreshes, and applies to later session turns without changing a turn that is already running.

Session delegation and consultation

The session sees the exact coding models enabled for delegated tasks. When it launches repository work, it can choose one of those models for that task or omit the choice to use the deployment default. An unavailable model is rejected before launch, so Roomote can correct the selection without consuming its launch attempt. The chosen model applies to the delegated task just as if it had been selected when starting a task from the web app. For a structured pull request review, Roomote can likewise choose an enabled model and a supported reasoning effort for the review task. Ask for those choices when the review needs a particular balance of speed and depth; omitted choices use the deployment’s code review defaults. A single session turn can launch multiple independent tasks, so one request can delegate separate workstreams without waiting for each task to finish first. Roomote still posts a kickoff for every task and keeps repeated identical launch requests from creating duplicate work. Roomote can also consult the advisor and judge roles for focused planning or completion checks without launching a workspace-backed task. These consultations can read deployment integrations and task status for context, but they cannot inspect a repository workspace, post to chat, or launch, message, or cancel tasks. Roomote remains responsible for the user-facing answer and any execution it delegates. OpenCode allows two levels of subagents: the main agent can delegate to a subagent, which can make one further nested delegation or consultation. This does not permit unbounded recursion or expand a role’s permissions. Roomote consultants still cannot inspect a repository workspace or control tasks; Roomote owns any workspace-backed execution.

Choosing models

Start by choosing for reliability, then optimize for cost and speed once tasks are working. For the default coding model, prioritize:
  • strong coding and debugging performance
  • reliable tool use across long multi-step tasks
  • enough context window for your repositories and logs
  • predictable behavior with your preferred inference provider
For helper work, prioritize:
  • fast responses
  • low cost
  • acceptable accuracy on short routing and summarization prompts
For media work, prioritize:
  • support for the image, audio, and video inputs you use
  • layout and screenshot understanding, accurate transcription, and clear video descriptions
For code review, prioritize:
  • careful reasoning over speed
  • good false-positive control
  • attention to tests, regressions, security, and edge cases
For advisor work, prioritize:
  • structured reasoning
  • ability to break work into practical steps
  • consistency with your deployment-wide and environment-specific guidance

Mixing providers

Mixing providers is normal. It can help when:
  • one provider has better pricing for a model you use heavily
  • another provider has better availability or rate limits
  • you want direct-provider access for one model and gateway routing for another
  • you are comparing model families before changing the default
  • a specialized model, such as a vision model, only exists behind one provider
The main tradeoff is operational complexity. Each provider adds credentials, account limits, billing, and possible regional or data-handling requirements. Keep the enabled list focused enough that teammates can understand which models to pick.

Keep model metadata fresh

Model context windows, output limits, supported input types, reasoning support, and pricing can change. Settings > Models can refresh model metadata so the admin UI has current information for enabled and custom models. Refresh metadata after adding custom models, changing providers, or upgrading a deployment. It helps admins compare models without relying on stale defaults.

Common issues

  • No models are available. Connect at least one inference provider and enable at least one model.
  • A model cannot be selected. Confirm its provider is connected and that the model is enabled in Settings > Models.
  • Tasks are expensive or slow. Move helper work to a cheaper model, lower reasoning where quality allows, or choose a faster default model.
  • An image attachment cannot be read. Check the selected Vision model in Settings > Models for image input support.
  • An audio or video attachment is not processed. Choose a model that supports the required input under Settings > Models > Audio and video model. If the role is Same as coding model, check that the coding model supports the attachment, or select a capable audio and video model explicitly. The Vision model handles images separately.
  • A model works from one provider but not another. Check provider-specific credentials, rate limits, model availability, and model ID prefix.
  • A task switched to a fallback model. Check the failed provider’s credits, credentials, and model availability. The current task stays on the fallback; a new task starts on the configured default.
  • Work stops with “You seem to have run out of credits”. The provider reported that the account has no remaining credits or quota, or that a ChatGPT subscription reached its usage limit. If that role has an enabled fallback, Roomote switches immediately; otherwise add credits, wait for the subscription limit to reset, or switch to another provider. Temporary rate limits are retried before fallback.