← Blog

gpt-oss-20b free API: OpenAI's open-weight model, OpenAI-compatible, $0

Call openai/gpt-oss-20b for free through InferAll's OpenAI- and Anthropic-compatible API. Measured latency, working Python, tool calling, JSON mode, streaming and Claude Code setup.

InferAll Team

2 min read
gpt-ossgpt-oss-20bfree LLM APIOpenAI APIopen-weight modelsNVIDIA NIMfunction calling

gpt-oss-20b is OpenAI's open-weight model. InferAll serves it for free, as openai/gpt-oss-20b, through NVIDIA's inference service: $0 per input and output token on the free trial, no card required.

Everything below was run against the live API on 2026-09-27.

What we measured Result
Short answer, 5 calls median 1.96s (1.33s to 2.89s)
Streaming first text after 0.74s
Tool calling (tool_choice auto) 3 of 3 calls returned a correct call
JSON mode (response_format) valid JSON
/v1/messages (Anthropic SDK) works
/v1/responses (Responses API) works
Claude Code completed a file-reading task, 2 of 2

Quick start (Python)

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key=os.environ["INFERALL_API_KEY"],
)

response = client.chat.completions.create(
    model="openai/gpt-oss-20b",
    messages=[{"role": "user", "content": "In one sentence, what is a hash map?"}],
    max_tokens=300,
)
print(response.choices[0].message.content)

Tool calling

import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.inferall.ai/v1", api_key=os.environ["INFERALL_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

response = client.chat.completions.create(
    model="openai/gpt-oss-20b",
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
    tools=tools,
    max_tokens=400,
)
call = response.choices[0].message.tool_calls[0]
print(call.function.name, json.loads(call.function.arguments))

This printed get_weather {'city': 'Paris'} in our run.

curl

curl https://api.inferall.ai/v1/chat/completions \
  -H "Authorization: Bearer $INFERALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-20b",
    "messages": [{"role": "user", "content": "Say hello in five words."}],
    "max_tokens": 200
  }'

Claude Code

export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=openai/gpt-oss-20b
claude

Things to know

  • Send the full id, openai/gpt-oss-20b. The openai/ part is the model's name on NVIDIA, not a sign that the request goes to OpenAI; InferAll routes this one id to NVIDIA.
  • gpt-oss-120b is not available. NVIDIA retired it; InferAll returns an error for it rather than a different model.
  • It is a 20B model. It is fast and free, and for harder reasoning or long code changes a larger model will do better. The models page lists the others, including free NVIDIA models and paid Claude, Gemini and OpenAI models.

Get started

  1. Create a key at inferall.ai/keys.
  2. Run the Python example above with INFERALL_API_KEY set.

Try it with one key: create a free account and your first 25 calls on open models are free, no card needed.

Start building free