← Blog

Run the OpenAI Agents SDK on Claude, Gemini and free NVIDIA models

Use the OpenAI Agents SDK with InferAll's /v1/responses and /v1/chat/completions endpoints. The model-object setup that works, a tool-calling agent, streaming, and what is not supported yet.

InferAll Team

3 min read
OpenAI Agents SDKResponses APIAI agentsfunction callingClaudeGeminiNVIDIA NIMLLM API

The OpenAI Agents SDK works with InferAll, so one agent codebase can run on Claude, Gemini, OpenAI and the free NVIDIA models. InferAll serves both APIs the SDK uses: the Responses API at /v1/responses and Chat Completions at /v1/chat/completions.

We ran every example on this page against the live API on 2026-09-27 with openai-agents 0.22.3 and openai 3.19.2. The tool-calling agent answered correctly on nvidia/nemotron-3-super-120b-a12b, gemini/gemini-2.5-flash, anthropic/claude-haiku-4-5-20251001 and openai/gpt-4o-mini, with and without streaming.

Install

pip install openai-agents
export INFERALL_API_KEY=ifu_your_key_here   # inferall.ai/keys

Pass a model object, not a model string

This is the one thing that differs from the SDK's own examples. If you write model="nvidia/nemotron-3-super-120b-a12b", the SDK reads nvidia/ as its own provider prefix and stops with Unknown prefix: nvidia before any request is sent. openai/... strings have their prefix removed by the SDK. Build the model object yourself and give it an InferAll client:

import asyncio
import os

from openai import AsyncOpenAI
from agents import Agent, OpenAIResponsesModel, Runner, function_tool, set_tracing_disabled

client = AsyncOpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key=os.environ["INFERALL_API_KEY"],
)
set_tracing_disabled(True)  # tracing uploads go to OpenAI, not InferAll


@function_tool
def get_weather(city: str) -> str:
    """Return the current weather for a city."""
    return f"{city}: 17C, light rain"


agent = Agent(
    name="assistant",
    instructions="Answer briefly. Use tools for weather questions.",
    model=OpenAIResponsesModel(
        model="nvidia/nemotron-3-super-120b-a12b",
        openai_client=client,
    ),
    tools=[get_weather],
)


async def main():
    result = await Runner.run(agent, "What is the weather in Paris?")
    print(result.final_output)


asyncio.run(main())

Output from our run:

The weather in Paris is currently 17 °C with light rain.

OpenAIResponsesModel uses /v1/responses. OpenAIChatCompletionsModel uses /v1/chat/completions. Both work; the rest of the agent code is the same.

Streaming

import asyncio
import os

from openai import AsyncOpenAI
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled
from openai.types.responses import ResponseTextDeltaEvent

client = AsyncOpenAI(base_url="https://api.inferall.ai/v1", api_key=os.environ["INFERALL_API_KEY"])
set_tracing_disabled(True)

agent = Agent(
    name="writer",
    instructions="Write short, plain answers.",
    model=OpenAIChatCompletionsModel(model="nvidia/nemotron-3-super-120b-a12b", openai_client=client),
)


async def main():
    result = Runner.run_streamed(agent, "Give me three names for a hiking app.")
    async for event in result.stream_events():
        if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
            print(event.data.delta, end="", flush=True)
    print()


asyncio.run(main())

Text streams as it is generated on both /v1/chat/completions and /v1/responses, for NVIDIA, OpenAI, Claude and Gemini models (updated 2026-09-27: earlier versions of this page said some of these delivered the whole answer at the end; that is no longer the case).

Switching models

Change the model= string. The free NVIDIA models work on the free trial. Claude, Gemini and OpenAI models are paid and billed from your balance, so add the $5 starter pack at /billing first. The full list is on the models page; send ids with their prefix, for example anthropic/claude-sonnet-5 or gemini/gemini-2.5-flash.

What is not supported yet

  • previous_response_id. InferAll does not store responses, so it returns a 400. The Agents SDK sends the full conversation by default, which works.
  • Hosted tools such as WebSearchTool and FileSearchTool. These run on OpenAI's servers; on /v1/responses InferAll returns a 400 that names the tool. Function tools work on every provider.
  • Tracing. The SDK uploads traces to OpenAI, not to InferAll, so turn it off with set_tracing_disabled(True) or give it a separate OpenAI key.

Get started

  1. Create a key at inferall.ai/keys.
  2. Run the example above with a free NVIDIA model.
  3. Add credits at /billing when you want Claude, Gemini or OpenAI models.

Try it with one key: create a free account and your first 25 calls on open models are free, no card needed.

Start building free