The OpenAI Agents SDK works with InferAll, so one agent codebase can run on Claude, Gemini, OpenAI and the free NVIDIA models. InferAll serves both APIs the SDK uses: the Responses API at /v1/responses and Chat Completions at /v1/chat/completions.
We ran every example on this page against the live API on 2026-09-27 with openai-agents 0.22.3 and openai 3.19.2. The tool-calling agent answered correctly on nvidia/nemotron-3-super-120b-a12b, gemini/gemini-2.5-flash, anthropic/claude-haiku-4-5-20251001 and openai/gpt-4o-mini, with and without streaming.
Install
pip install openai-agents
export INFERALL_API_KEY=ifu_your_key_here # inferall.ai/keys
Pass a model object, not a model string
This is the one thing that differs from the SDK's own examples. If you write model="nvidia/nemotron-3-super-120b-a12b", the SDK reads nvidia/ as its own provider prefix and stops with Unknown prefix: nvidia before any request is sent. openai/... strings have their prefix removed by the SDK. Build the model object yourself and give it an InferAll client:
import asyncio
import os
from openai import AsyncOpenAI
from agents import Agent, OpenAIResponsesModel, Runner, function_tool, set_tracing_disabled
client = AsyncOpenAI(
base_url="https://api.inferall.ai/v1",
api_key=os.environ["INFERALL_API_KEY"],
)
set_tracing_disabled(True) # tracing uploads go to OpenAI, not InferAll
@function_tool
def get_weather(city: str) -> str:
"""Return the current weather for a city."""
return f"{city}: 17C, light rain"
agent = Agent(
name="assistant",
instructions="Answer briefly. Use tools for weather questions.",
model=OpenAIResponsesModel(
model="nvidia/nemotron-3-super-120b-a12b",
openai_client=client,
),
tools=[get_weather],
)
async def main():
result = await Runner.run(agent, "What is the weather in Paris?")
print(result.final_output)
asyncio.run(main())
Output from our run:
The weather in Paris is currently 17 °C with light rain.
OpenAIResponsesModel uses /v1/responses. OpenAIChatCompletionsModel uses /v1/chat/completions. Both work; the rest of the agent code is the same.
Streaming
import asyncio
import os
from openai import AsyncOpenAI
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled
from openai.types.responses import ResponseTextDeltaEvent
client = AsyncOpenAI(base_url="https://api.inferall.ai/v1", api_key=os.environ["INFERALL_API_KEY"])
set_tracing_disabled(True)
agent = Agent(
name="writer",
instructions="Write short, plain answers.",
model=OpenAIChatCompletionsModel(model="nvidia/nemotron-3-super-120b-a12b", openai_client=client),
)
async def main():
result = Runner.run_streamed(agent, "Give me three names for a hiking app.")
async for event in result.stream_events():
if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
print(event.data.delta, end="", flush=True)
print()
asyncio.run(main())
Text streams as it is generated on both /v1/chat/completions and /v1/responses, for NVIDIA, OpenAI, Claude and Gemini models (updated 2026-09-27: earlier versions of this page said some of these delivered the whole answer at the end; that is no longer the case).
Switching models
Change the model= string. The free NVIDIA models work on the free trial. Claude, Gemini and OpenAI models are paid and billed from your balance, so add the $5 starter pack at /billing first. The full list is on the models page; send ids with their prefix, for example anthropic/claude-sonnet-5 or gemini/gemini-2.5-flash.
What is not supported yet
- previous_response_id. InferAll does not store responses, so it returns a 400. The Agents SDK sends the full conversation by default, which works.
- Hosted tools such as
WebSearchToolandFileSearchTool. These run on OpenAI's servers; on/v1/responsesInferAll returns a 400 that names the tool. Function tools work on every provider. - Tracing. The SDK uploads traces to OpenAI, not to InferAll, so turn it off with
set_tracing_disabled(True)or give it a separate OpenAI key.
Get started
- Create a key at inferall.ai/keys.
- Run the example above with a free NVIDIA model.
- Add credits at /billing when you want Claude, Gemini or OpenAI models.