← Blog

Claude Opus 5.5, Opus 5, Sonnet 5 and Fable 5.1 through one API key

Call the current Claude models through InferAll's OpenAI- and Anthropic-compatible endpoints: measured latency, list prices, working Python and Claude Code setup.

InferAll Team

3 min read
ClaudeClaude Opus 5.5Claude Sonnet 5Claude Fable 5.1Anthropic APIOpenAI APIClaude CodeLLM API

The current Claude family is available through InferAll with the same key you use for every other model: Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1, plus the Claude 4 models.

Model id Anthropic list price per 1M tokens (input / output)
anthropic/claude-opus-5-5 $4 / $20
anthropic/claude-opus-5 $5 / $25
anthropic/claude-sonnet-5 $2 / $10
anthropic/claude-fable-5-1 $10 / $50

These are paid models, billed from your InferAll balance. The free trial covers the open NVIDIA models only, so add the $5 starter pack at /billing first.

What we measured (2026-09-26, through the live API)

Two short prompts per model on /v1/chat/completions and one on /v1/messages. Every call answered as the model named.

Model Chat completions Messages
Opus 5.5 1.9s, 1.7s 1.1s
Opus 5 1.8s, 1.6s 1.5s
Sonnet 5 1.8s, 1.9s 1.7s
Fable 5.1 3.6s, 3.6s 3.0s

These are short prompts; long outputs take longer.


Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key="ifu_your_key_here",  # inferall.ai/keys
)

response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Review this function for bugs: ..."}],
    max_tokens=2000,
)
print(response.choices[0].message.content)

Python (Anthropic SDK)

import anthropic

client = anthropic.Anthropic(
    base_url="https://api.inferall.ai",
    api_key="ifu_your_key_here",
)

message = client.messages.create(
    model="anthropic/claude-opus-5-5",
    max_tokens=2000,
    messages=[{"role": "user", "content": "Plan a database migration for ..."}],
)
print(message.content[0].text)

Claude Code

Once your account has credits, pick any Claude model (Claude Code's own default Claude models also work without the last export):

export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=anthropic/claude-sonnet-5
claude

On the free trial, set ANTHROPIC_MODEL=nvidia/nemotron-3-super-120b-a12b instead: trial accounts cannot use paid models, and the gateway refuses a Claude id rather than answering with a different model.


Things that behave differently on these models

  • Sampling parameters are not accepted by the Claude 5 models. If you send temperature or top_p through our OpenAI-compatible endpoint, InferAll drops them instead of failing the request.
  • Forced tool use is not accepted by Fable 5.1 and Opus 5.5. If you send tool_choice: "required" or a named function through /v1/chat/completions, InferAll sends auto plus an instruction to call the tool. In our tests both models called the tool; it is a strong instruction rather than a guarantee.
  • /v1/models lists these ids exactly as you should send them (anthropic/claude-...), so tools that build a model menu from it, such as Claude Desktop in gateway mode, work without edits.

Get started

  1. Create an account at inferall.ai/keys.
  2. Add the $5 starter pack at /billing.
  3. Call any model above with your key. The full list is on the models page.

Try it with one key: create a free account and your first 25 calls on open models are free, no card needed.

Start building free