OpenAI's GPT-4.1 family is now available through InferAll — the same OpenAI-compatible endpoint that already routes to Anthropic, Gemini, and 40+ free NVIDIA NIM models.
You get all three tiers with one key:
| Model | Input | Output | Best for |
|---|---|---|---|
gpt-4.1 |
$2.00/M | $8.00/M | Complex reasoning, long context |
gpt-4.1-mini |
$0.40/M | $1.60/M | Most production workloads |
gpt-4.1-nano |
$0.10/M | $0.40/M | High-volume, latency-sensitive |
Prices are OpenAI's published list rates — InferAll passes them through at zero markup.
Drop-in with the OpenAI SDK
Using a new key? A no-card trial key routes to the free NIM models only, so the paid model in these snippets returns a 402 until you add a card. To get a working first call right now, keep everything below and swap the model to
meta/llama-3.1-8b-instruct.
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # get one free at inferall.ai/keys
)
# Full model — complex tasks
response = client.chat.completions.create(
model="openai/gpt-4.1",
messages=[{"role": "user", "content": "Review this code for security issues: ..."}],
max_tokens=1024,
)
# Mini — most workloads, 5× cheaper
response = client.chat.completions.create(
model="openai/gpt-4.1-mini",
messages=[{"role": "user", "content": "Summarize this document in three bullets."}],
max_tokens=256,
)
# Nano — high-volume classification, routing, structured extraction
response = client.chat.completions.create(
model="openai/gpt-4.1-nano",
messages=[{"role": "user", "content": "Classify this support ticket: ..."}],
max_tokens=64,
)
print(response.choices[0].message.content)
The base_url swap is the only change. Your existing OpenAI SDK code, LangChain pipelines, and LlamaIndex retrievers all work unchanged.
TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const response = await client.chat.completions.create({
model: "openai/gpt-4.1-mini",
messages: [{ role: "user", content: "Explain async/await in one paragraph." }],
max_tokens: 200,
});
console.log(response.choices[0].message.content);
Also new: o3 and o4-mini
The same deploy that brought GPT-4.1 also added OpenAI's reasoning models:
# o3 — strong reasoning, slower
response = client.chat.completions.create(
model="openai/o3",
messages=[{"role": "user", "content": "Prove that the square root of 2 is irrational."}],
)
# o4-mini — faster reasoning, lower cost
response = client.chat.completions.create(
model="openai/o4-mini",
messages=[{"role": "user", "content": "Debug this Python traceback: ..."}],
)
Why route through InferAll
One key, every provider. The same ifu_... key routes to GPT-4.1, Claude Sonnet, Gemini Flash, and 40+ free NVIDIA models. You don't manage separate OpenAI, Anthropic, and Google credentials.
Switch models without changing code. Want to compare GPT-4.1-mini vs Claude Sonnet 4.6 on the same prompt? Change one string. The response shape is identical.
Free trial, no card required. New accounts get a free trial on the 40+ open-source NIM models. The one-time $5 activation unlocks ongoing free-NIM use and becomes spendable balance for premium providers.
Get your key at inferall.ai/keys.