The current Claude family is available through InferAll with the same key you use for every other model: Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1, plus the Claude 4 models.
| Model id | Anthropic list price per 1M tokens (input / output) |
|---|---|
anthropic/claude-opus-5-5 |
$4 / $20 |
anthropic/claude-opus-5 |
$5 / $25 |
anthropic/claude-sonnet-5 |
$2 / $10 |
anthropic/claude-fable-5-1 |
$10 / $50 |
These are paid models, billed from your InferAll balance. The free trial covers the open NVIDIA models only, so add the $5 starter pack at /billing first.
What we measured (2026-09-26, through the live API)
Two short prompts per model on /v1/chat/completions and one on /v1/messages. Every call answered as the model named.
| Model | Chat completions | Messages |
|---|---|---|
| Opus 5.5 | 1.9s, 1.7s | 1.1s |
| Opus 5 | 1.8s, 1.6s | 1.5s |
| Sonnet 5 | 1.8s, 1.9s | 1.7s |
| Fable 5.1 | 3.6s, 3.6s | 3.0s |
These are short prompts; long outputs take longer.
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Review this function for bugs: ..."}],
max_tokens=2000,
)
print(response.choices[0].message.content)
Python (Anthropic SDK)
import anthropic
client = anthropic.Anthropic(
base_url="https://api.inferall.ai",
api_key="ifu_your_key_here",
)
message = client.messages.create(
model="anthropic/claude-opus-5-5",
max_tokens=2000,
messages=[{"role": "user", "content": "Plan a database migration for ..."}],
)
print(message.content[0].text)
Claude Code
Once your account has credits, pick any Claude model (Claude Code's own default Claude models also work without the last export):
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=anthropic/claude-sonnet-5
claude
On the free trial, set ANTHROPIC_MODEL=nvidia/nemotron-3-super-120b-a12b instead: trial accounts cannot use paid models, and the gateway refuses a Claude id rather than answering with a different model.
Things that behave differently on these models
- Sampling parameters are not accepted by the Claude 5 models. If you send
temperatureortop_pthrough our OpenAI-compatible endpoint, InferAll drops them instead of failing the request. - Forced tool use is not accepted by Fable 5.1 and Opus 5.5. If you send
tool_choice: "required"or a named function through/v1/chat/completions, InferAll sendsautoplus an instruction to call the tool. In our tests both models called the tool; it is a strong instruction rather than a guarantee. /v1/modelslists these ids exactly as you should send them (anthropic/claude-...), so tools that build a model menu from it, such as Claude Desktop in gateway mode, work without edits.
Get started
- Create an account at inferall.ai/keys.
- Add the $5 starter pack at /billing.
- Call any model above with your key. The full list is on the models page.