Docs
Use InferAll with OpenCode
Add InferAll as an OpenCode provider, list your models, and run a small read-only task. The example uses exact routing: a failed request is returned instead of switching models.
Tested with OpenCode 1.18.35. For OpenCode's own options, see its custom provider docs and CLI docs.
1. Install OpenCode
Install the version this guide was tested with.
Install
npm install -g opencode-ai@1.18.35
opencode --version2. Set your key
Create a key at /keys and export it in the terminal you will use. The config refers to the variable by name, keeping the secret out of opencode.json.
Shell
# Use your own key from inferall.ai/keys. Never paste it into a file you commit.
export INFERALL_API_KEY="paste-your-inferall-key-here"3. Create a scratch project
Use a new directory so no existing project config is overwritten. If $HOME/opencode-inferall-test already exists, pick another name in every command.
Directory
mkdir -p "$HOME/opencode-inferall-test"Keep the quotes around 'EOF' when you copy the config, so the shell leaves {env:INFERALL_API_KEY} for OpenCode to resolve.
opencode.json (scratch only)
cat > "$HOME/opencode-inferall-test/opencode.json" <<'EOF'
{
"$schema": "https://opencode.ai/config.json",
"enabled_providers": ["inferall"],
"model": "inferall/nvidia/nemotron-3.5-lightning-30b-a3b",
"small_model": "inferall/nvidia/nemotron-3.5-lightning-30b-a3b",
"provider": {
"inferall": {
"npm": "@ai-sdk/openai-compatible",
"name": "InferAll",
"options": {
"baseURL": "https://api.inferall.ai/v1",
"apiKey": "{env:INFERALL_API_KEY}",
"headers": { "X-Inferall-Routing": "exact" },
"timeout": 90000
},
"models": {
"nvidia/nemotron-3.5-lightning-30b-a3b": {
"name": "Nemotron 3.5 Lightning (free)",
"tool_call": true,
"limit": { "context": 131072, "output": 1024 },
"options": { "reasoningEffort": "none" }
},
"gemini/gemini-2.5-flash": {
"name": "Gemini 2.5 Flash (paid)",
"tool_call": true,
"limit": { "context": 131072, "output": 1024 },
"options": { "reasoningEffort": "none" }
}
}
}
},
"share": "disabled",
"autoupdate": false,
"snapshot": false,
"lsp": false,
"formatter": false,
"mcp": {},
"plugin": [],
"compaction": { "auto": false, "prune": false },
"permission": { "*": "deny", "read": "allow", "external_directory": "deny" },
"agent": {
"readonly": {
"mode": "primary",
"steps": 2,
"permission": { "*": "deny", "read": "allow", "external_directory": "deny" }
},
"title": { "disable": true },
"summary": { "disable": true }
}
}
EOF- Two models are listed: the free Nemotron 3.5 Lightning and the paid Gemini 2.5 Flash. Listing a model does not call it, and does not grant access to it.
- The default model is the free one. Each task still names its model explicitly in step 4.
limit.contextandlimit.outputare the budgets OpenCode uses for this client, not the provider's maximum. Check /models for current model details.reasoningEffortmust be camelCase in OpenCode config. OpenCode sends it asreasoning_effort: "none"; a snake_case key is silently ignored.- The restrictive settings (read-only permissions, no LSP, formatter, MCP or plugins, 2 steps) are for this first check only. They are not a recommended everyday setup.
4. Choose one model
Pick one and keep using the same terminal so the variable stays set. The run command refuses to start if it is empty. Exact routing returns a failed request to you instead of substituting another model.
Free: Nemotron 3.5 Lightning. Available within eligible trial allowances and daily quotas.
Choose the free model
export INFERALL_TEST_MODEL="nvidia/nemotron-3.5-lightning-30b-a3b"Paid: Gemini 2.5 Flash. This spends your included monthly usage or prepaid balance, and needs paid-model access. Trial accounts can call only the free models; see /billing and /models.
Choose the paid model
export INFERALL_TEST_MODEL="gemini/gemini-2.5-flash"Run only one of the two exports.
In our October 7, 2026 checks, Lightning completed the read-and-summary task using the explicit file path below; Gemini also completed the task. Both used an internal project key. An earlier Lightning probe timed out. These checks do not establish trial-account access or ongoing availability.
5. Add a synthetic README and run
The task reads one small made-up file and reports what it says. The prompt supplies the full file path so the model does not need to infer its directory.
Synthetic README
cat > "$HOME/opencode-inferall-test/README.md" <<'EOF'
# Synthetic project notes
Open tasks:
- IFA-217: Build the CSV importer. Owner: Mina.
- IFA-318: Add the request history view. Owner: Jules.
EOFThe command runs in the scratch directory with the read-only readonly agent, writes fresh events to out.jsonl, and only runs the checker if OpenCode succeeded. The checker looks for one completed read of the README, a final answer naming both IDs and owners, a normal stop, and no error events.
Run the check
: "${INFERALL_TEST_MODEL:?Choose a model first}" \
&& : "${INFERALL_API_KEY:?Set your key first}" \
&& cd "$HOME/opencode-inferall-test" \
&& opencode run --pure --format json --agent readonly \
--model "inferall/${INFERALL_TEST_MODEL}" \
--title 'InferAll first task' \
"Read $PWD/README.md and report both task IDs, owners and descriptions. Use the read tool; do not modify any files." \
> out.jsonl \
&& node <<'JS'
const fs = require('node:fs');
const events = fs.readFileSync('out.jsonl', 'utf8')
.trim().split('\n').filter(Boolean).map(line => JSON.parse(line));
const reads = events.filter(event => event.type === 'tool_use'
&& event.part?.tool === 'read');
const errors = events.filter(event => event.type === 'error');
const finish = events.findLast(event => event.type === 'step_finish');
const text = reads.length === 1
? events.slice(events.indexOf(reads[0]) + 1)
.filter(event => event.type === 'text')
.map(event => event.part?.text || '').join('\n')
: '';
const expected = ['IFA-217', 'Mina', 'IFA-318', 'Jules'];
if (errors.length || reads.length !== 1
|| reads[0].part.state?.status !== 'completed'
|| !/(?:^|\/)README\.md$/.test(reads[0].part.state.input?.filePath || '')
|| finish?.part?.reason !== 'stop'
|| !expected.every(value => text.includes(value))) {
console.error(errors[0] ? JSON.stringify(errors[0]) : 'Task incomplete: check out.jsonl.');
process.exit(1);
}
console.log(text);
JSExpect a short summary naming IFA-217 (CSV importer, Mina) and IFA-318 (request history view, Jules). Wording will vary. On failure the command exits nonzero and prints the first error or a short message; the raw events are in out.jsonl.
6. Use it in your own projects
Merge only the provider.inferall block into your own opencode.json, keeping your other settings. Leave out enabled_providers unless you want to hide your other providers. The test's permission lines would block normal coding tools, so do not copy them. Model names are always inferall/ plus the full model ID, slash included.
Start OpenCode with an InferAll model
opencode --model "inferall/${INFERALL_TEST_MODEL:?choose a model first}"Troubleshooting
ProviderModelNotFoundError or an unknown model
The model is not listed under provider.inferall.models, or you passed a name without the inferall/ prefix. Add the model ID exactly as shown on /models.
Invalid key (401)
Check the variable is set in this terminal and the key is active at /keys.
A paid model is refused, or the trial or daily quota is used up
Trial accounts can call only free models. Complete a credit purchase at /billing for paid models. Trial allowances and daily caps are separate limits; the response says which applies. See rate limits.
502 "Service temporarily overloaded"
With exact routing, the failure is returned instead of switching models. Retry later. Choosing the other model is your call, and the paid one spends usage. See routing.
Answer cut off
The config caps output at 1024 tokens per request. Raise limit.output for longer answers; the checker requires a normal stop.