Correction, 2026-09-02 — pricing. An earlier version of this post said premium provider tokens are billed at the provider's rate with zero markup. That was not accurate: InferAll bills premium tokens at its own published per-token rates, which sit above the underlying provider cost. Current rates are on /pricing. The post has been corrected; nothing else about the models or the setup has changed.
If you've been hitting Anthropic's $5/month minimum or just want a sandbox for Claude Code without funding a balance, point it at InferAll instead. Two env vars, one CLI command, and a free trial before any signup nag.
Setup
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=your_inferall_key
export ANTHROPIC_MODEL=nvidia/nemotron-3-ultra-550b-a55b
claude
That's it. The InferAll key is an ifu_... token you create at inferall.ai/keys. No card needed to sign up; you get a free trial against NVIDIA NIM open-source models.
What actually runs
Good free routes to try:
nvidia/nemotron-3-ultra-550b-a55b— 550B parameter open model (55B active); our current default free model.nvidia/nemotron-3-ultra-550b-a55b— the largest free model; use it when quality matters more than speed.openai/gpt-oss-20b— OpenAI's open-weight model; low-latency general and coding turns.
All three are $0 input / $0 output. The trial covers successful calls across all three.
If you want actual Claude
The same ANTHROPIC_BASE_URL config also accepts a model="anthropic/claude-sonnet-4-6" style prefix to force the real Claude provider through InferAll's billing at our published per-token rates, with one API key instead of separate Anthropic + Anthropic-via-bedrock + Anthropic-via-vertex keys. After the free trial, the $5 starter pack unlocks ongoing free NIM use AND becomes spendable balance for premium routing.
What's the catch?
Honestly nothing for typical Claude Code usage. The free NIM models are good enough for most code generation, refactoring, and chat tasks. Where the catch shows up:
- Tool calls / agentic flows: the NIM models support OpenAI-style function calls; Anthropic's tool-use schema is auto-translated. Mostly works, occasional edge cases.
- Long-context (128K+): current NIM model context windows vary and quality at the tail varies. For >50K context, Claude Sonnet via the prefix usually wins.
- Vision: NIM doesn't currently expose multimodal models; vision routes to Anthropic/Gemini directly (charged from balance).
Why this exists
InferAll is an AI gateway — one base URL, one API key, every major provider routed underneath at the provider's published rate at our published rates. The free NIM tier exists because NVIDIA gives us free inference on the open-source catalog and it's cheaper to pass that through than to gatekeep it. The $5 starter pack is the conversion event for users who want premium-on-demand; everyone else stays in the free tier indefinitely and we're fine with that.
If you want to read the docs, they cover the OpenAI SDK route, the raw REST endpoint, and the trial-status headers SDKs can pick up to surface usage inline.
Using the Claude Desktop app rather than the CLI? See Claude Desktop with InferAll.