gpt-oss-20b is OpenAI's open-weight model. InferAll serves it for free, as openai/gpt-oss-20b, through NVIDIA's inference service: $0 per input and output token on the free trial, no card required.
Everything below was run against the live API on 2026-09-27.
| What we measured | Result |
|---|---|
| Short answer, 5 calls | median 1.96s (1.33s to 2.89s) |
| Streaming | first text after 0.74s |
| Tool calling (tool_choice auto) | 3 of 3 calls returned a correct call |
JSON mode (response_format) |
valid JSON |
/v1/messages (Anthropic SDK) |
works |
/v1/responses (Responses API) |
works |
| Claude Code | completed a file-reading task, 2 of 2 |
Quick start (Python)
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key=os.environ["INFERALL_API_KEY"],
)
response = client.chat.completions.create(
model="openai/gpt-oss-20b",
messages=[{"role": "user", "content": "In one sentence, what is a hash map?"}],
max_tokens=300,
)
print(response.choices[0].message.content)
Tool calling
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.inferall.ai/v1", api_key=os.environ["INFERALL_API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
response = client.chat.completions.create(
model="openai/gpt-oss-20b",
messages=[{"role": "user", "content": "What's the weather in Paris?"}],
tools=tools,
max_tokens=400,
)
call = response.choices[0].message.tool_calls[0]
print(call.function.name, json.loads(call.function.arguments))
This printed get_weather {'city': 'Paris'} in our run.
curl
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-20b",
"messages": [{"role": "user", "content": "Say hello in five words."}],
"max_tokens": 200
}'
Claude Code
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=openai/gpt-oss-20b
claude
Things to know
- Send the full id,
openai/gpt-oss-20b. Theopenai/part is the model's name on NVIDIA, not a sign that the request goes to OpenAI; InferAll routes this one id to NVIDIA. - gpt-oss-120b is not available. NVIDIA retired it; InferAll returns an error for it rather than a different model.
- It is a 20B model. It is fast and free, and for harder reasoning or long code changes a larger model will do better. The models page lists the others, including free NVIDIA models and paid Claude, Gemini and OpenAI models.
Get started
- Create a key at inferall.ai/keys.
- Run the Python example above with
INFERALL_API_KEYset.