Docs

Use InferAll with Pi

This guide adds InferAll as a custom provider in the Pi coding agent, then runs one finite, read-only task in a scratch directory. It uses the OpenAI-compatible endpoint with X-Inferall-Routing: exact, so a failed request returns the failure instead of switching models.

On October 7, 2026, with Pi 1.0.4 and an internal project key, each model listed below completed one synthetic read-and-summary task using these provider settings with exact routing. The free Ultra route also returned provider-overload (502) errors in earlier tests. These internal tests do not establish trial-account behavior or sustained availability. Try your own workload before relying on it. Pi is a separate project; for its own options see the models docs and CLI docs.

1. Install Pi

Pi needs Node.js 22.19 or newer. Install the version this guide was tested with.

Install

npm install -g @earendil-works/pi-coding-agent@1.0.4
pi --version

2. Set your key

Create a key at /keys and export it in the terminal you will use. The config below refers to the variable by name, so the key itself never goes into a config file. Do not use --api-key, which puts the key in the process arguments.

Shell

# Use your own key from inferall.ai/keys. Never paste it into a file you commit.
export INFERALL_API_KEY="paste-your-inferall-key-here"

3. Create a scratch config

For the first test, use a dedicated Pi config directory so nothing in your everyday ~/.pi/agent is touched. Pi reads this directory from PI_CODING_AGENT_DIR; the test command sets it for that one command only.

Directories

mkdir -p "$HOME/pi-inferall-test/agent" "$HOME/pi-inferall-test/work"

The heredoc delimiter is quoted (<<'EOF') so the shell leaves $INFERALL_API_KEY as text for Pi to resolve. Keep the quotes when you copy it.

models.json (provider block)

cat > "$HOME/pi-inferall-test/agent/models.json" <<'EOF'
{
  "providers": {
    "inferall": {
      "baseUrl": "https://api.inferall.ai/v1",
      "api": "openai-completions",
      "apiKey": "$INFERALL_API_KEY",
      "headers": {
        "X-Inferall-Routing": "exact"
      },
      "models": [
        {
          "id": "nvidia/nemotron-3-ultra-550b-a55b",
          "maxTokens": 1024,
          "samplingParams": { "reasoning_effort": "none" }
        },
        {
          "id": "gemini/gemini-2.5-flash",
          "maxTokens": 1024,
          "samplingParams": { "reasoning_effort": "none" }
        }
      ]
    }
  }
}
EOF

This registers two models: the current free Nemotron 3 Ultra and the paid Gemini 2.5 Flash. Registering a model does not call it; Pi only sends requests to the model you select in the next step.

maxTokens: 1024 is the output cap requested on each call. It is not a limit on the whole run or on spend, and not the model's context maximum.

samplingParams makes Pi send reasoning_effort: "none" explicitly. On these two models, it requests thinking off for this check.

The block does not set a context window or prices; check /models for the model's current details.

This settings file turns off Pi's own retries and install telemetry for the test. If Pi retried on an error, one task could send more requests than you expect.

settings.json (scratch only)

cat > "$HOME/pi-inferall-test/agent/settings.json" <<'EOF'
{
  "retry": {
    "enabled": false,
    "maxRetries": 0,
    "provider": { "maxRetries": 0, "timeoutMs": 90000 }
  },
  "compaction": { "enabled": false },
  "cacheWarming": "off",
  "enableInstallTelemetry": false,
  "quietStartup": true
}
EOF

4. Choose one model

Pick one of the two and keep using the same terminal, so the variable stays set. The run command refuses to start if it is empty, and Pi never switches between them on its own.

Free default: Nemotron 3 Ultra. It is available within eligible trial allowances and daily quotas. We also observed provider-overload errors on this route.

Choose the free model

export INFERALL_TEST_MODEL="nvidia/nemotron-3-ultra-550b-a55b"

Paid alternative: Gemini 2.5 Flash. It runs through a different provider and requires paid-model access plus available monthly included usage or prepaid balance. Trial accounts can call only the free models; complete a credit purchase to activate paid-model access. See /billing, /models and /pricing.

Choose the paid model

export INFERALL_TEST_MODEL="gemini/gemini-2.5-flash"

Run only one of the two exports. A different provider may avoid one provider's overload, but both routes go through the same InferAll gateway, so this does not guarantee availability. Another NVIDIA model would still depend on the same upstream service. Choosing Gemini is a manual switch and uses paid-model usage.

5. Add a synthetic README and run

The task reads one small file and summarizes it. The file is made up.

Synthetic README

cat > "$HOME/pi-inferall-test/work/README.md" <<'EOF'
# Synthetic project notes

Open issues:

- ISSUE-17: Add a retry-limit setting. Owner: Mira.
- ISSUE-23: Document timeout handling. Owner: Leon.
EOF

The command runs from the scratch work directory and exits when the task is done. It loads no extensions, MCP servers, skills, prompt templates, themes or context files, keeps no session, ignores project-local files, and allows only the read tool. Provider, model (from INFERALL_TEST_MODEL) and thinking level are set explicitly. JSON mode is used so you can inspect tool events.

Run the check

: "${INFERALL_TEST_MODEL:?Choose one model in step 4}" \
&& cd "$HOME/pi-inferall-test/work" \
&& env PI_CODING_AGENT_DIR="$HOME/pi-inferall-test/agent" \
  pi --print --mode json --offline --no-approve --no-session \
  --no-extensions --no-mcp --no-skills --no-prompt-templates \
  --no-themes --no-context-files \
  --tools read \
  --provider inferall --model "$INFERALL_TEST_MODEL" --thinking off \
  --system-prompt "You are a read-only assistant for this synthetic task. Use read exactly once on ./README.md, and no other file. Then summarize the two issues and their owners. Do not invent details." \
  "Read ./README.md with the read tool and summarize both issue IDs, titles and owners." \
  > ../out.jsonl \
&& node <<'JS'
const fs = require('node:fs');
const events = fs.readFileSync('../out.jsonl', 'utf8')
  .trim().split('\n').filter(Boolean).map(line => JSON.parse(line));
const end = events.findLast(event => event.type === 'agent_end');
const answer = end?.messages?.findLast(message => message.role === 'assistant');
const reads = events.filter(event => event.type === 'tool_execution_start'
  && event.toolName === 'read' && /(?:^|\/)README\.md$/.test(event.args?.path || ''));
const readFinished = events.some(event => event.type === 'tool_execution_end'
  && event.toolCallId === reads[0]?.toolCallId && event.isError === false);
const text = answer?.content?.filter(part => part.type === 'text')
  .map(part => part.text).join('\n') || '';
if (!answer || answer.stopReason !== 'stop' || reads.length !== 1 || !readFinished
    || !['ISSUE-17', 'ISSUE-23', 'Mira', 'Leon'].every(value => text.includes(value))) {
  console.error(answer?.errorMessage || 'Task incomplete: check the tool events and answer.');
  process.exit(1);
}
console.log(text);
JS

The read tool is not a filesystem sandbox. It can open any path your user account can read, and the model chooses the path. The scratch directory only limits what this test is pointed at. Do not run it from a directory with files you would not want sent to a model provider.

What to expect:

6. Merge into your everyday config

Once the check passes, open your everyday ~/.pi/agent/models.json (or the agent directory you normally set in PI_CODING_AGENT_DIR). Add the inferall entry under providers, keeping your other provider entries. If InferAll is already configured, review that entry instead of adding a duplicate. Keep the environment reference for the key. Set INFERALL_API_KEY again when you open a new terminal. The scratch settings are optional for everyday use. You can add only the model you chose. The command below reads the same INFERALL_TEST_MODEL variable, so export your choice again in each new terminal; it stops if the variable is unset.

Then, with your usual tools and settings

pi --provider inferall --model "${INFERALL_TEST_MODEL:?choose a model first}" --thinking off

Limits, cost and accuracy

Troubleshooting

No API key, or the model is not found

In the same terminal, run test -n "$INFERALL_API_KEY" && echo set; it should print set. An unset variable leaves Pi without a key, and Pi may report a missing key or an unknown model instead of sending a request. Run PI_CODING_AGENT_DIR="$HOME/pi-inferall-test/agent" pi --offline --list-models inferall to check the scratch provider loaded.

Config or model selection errors

Pi reports models.json parse and schema problems at startup. Make sure the file is in the directory PI_CODING_AGENT_DIR points to, and that --provider inferall is paired with --model using the full model ID.

A paid model is refused on a trial key

Trial accounts can call only the free models marked on /models. Use one of those IDs in models.json, or complete a credit purchase at /billing for paid models.

402, trial exhausted, or a daily quota

These are separate limits. A used-up trial allowance and a plan's daily request cap are counted differently; the response body says which applies. See rate limits.

The request fails with exact routing

A 502 with "Service temporarily overloaded" means the selected provider could not serve that request. With exact, the failure is returned and InferAll does not silently switch to another one. Retry later, choose the other model in step 4 yourself (the paid route is not a trial fallback; trial accounts stay on free models), or remove the X-Inferall-Routing header to allow fallback (see routing).