Models
Every model, with its real price.
135 models reachable from one base URL and one key. 41 of them are $0 input and $0 output. Premium providers bill at the provider's published rate with zero markup.
Pass an ID below as the model argument exactly as written, except for the 10 marked served by a default model, which we do not route yet and which answer from a Llama model instead. Premium IDs are shown with the provider prefix they need, for example openai/gpt-4o-mini or anthropic/claude-sonnet-4-6. Sent without that prefix an ID routes to the free NVIDIA tier instead of the provider you named. This list is generated from the live catalog and was last refreshed on 2026-08-05.
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_inferall_key",
)
resp = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct", # any id from the free table below
messages=[{"role": "user", "content": "Hello"}],
)Free open models (41)
$0 input and $0 output, served on NVIDIA NIM. A new trial key can call these without a card. If yours returns a 402 here, your account already started checkout at some point and needs it finished before free calls resume.
| Model ID | Input / M | Output / M |
|---|---|---|
baai/bge-m3served by a default model | $0 | $0 |
deepseek-ai/deepseek-v4-flashserved by a default model | $0 | $0 |
deepseek-ai/deepseek-v4-proserved by a default model | $0 | $0 |
google/diffusiongemma-26b-a4b-itserved by a default model | $0 | $0 |
google/gemma-4-31b-itserved by a default model | $0 | $0 |
meta/llama-3.1-70b-instruct | $0 | $0 |
meta/llama-3.1-8b-instruct | $0 | $0 |
meta/llama-3.2-11b-vision-instruct | $0 | $0 |
meta/llama-3.2-1b-instruct | $0 | $0 |
meta/llama-3.2-3b-instruct | $0 | $0 |
meta/llama-3.2-90b-vision-instruct | $0 | $0 |
meta/llama-3.3-70b-instruct | $0 | $0 |
meta/llama-guard-4-12b | $0 | $0 |
minimaxai/minimax-m3served by a default model | $0 | $0 |
mistralai/mistral-medium-3.5-128b | $0 | $0 |
mistralai/mistral-nemotron | $0 | $0 |
nvidia/ai-synthetic-video-detector | $0 | $0 |
nvidia/ising-calibration-1.5-31b | $0 | $0 |
nvidia/llama-3.1-nemoguard-8b-content-safety | $0 | $0 |
nvidia/llama-3.1-nemoguard-8b-topic-control | $0 | $0 |
nvidia/llama-3.1-nemotron-nano-8b-v1 | $0 | $0 |
nvidia/llama-3.1-nemotron-nano-vl-8b-v1 | $0 | $0 |
nvidia/llama-3.1-nemotron-safety-guard-8b-v3 | $0 | $0 |
nvidia/llama-3.3-nemotron-super-49b-v1 | $0 | $0 |
nvidia/llama-3.3-nemotron-super-49b-v1.5 | $0 | $0 |
nvidia/nemoretriever-parse | $0 | $0 |
nvidia/nemotron-3-nano-30b-a3b | $0 | $0 |
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning | $0 | $0 |
nvidia/nemotron-3-super-120b-a12b | $0 | $0 |
nvidia/nemotron-3-ultra-550b-a55b | $0 | $0 |
nvidia/nemotron-3.5-content-safety | $0 | $0 |
nvidia/nemotron-mini-4b-instruct | $0 | $0 |
nvidia/nemotron-nano-12b-v2-vl | $0 | $0 |
nvidia/nemotron-parse | $0 | $0 |
nvidia/nvidia-nemotron-nano-9b-v2 | $0 | $0 |
nvidia/riva-translate-4b-instruct-v1.1 | $0 | $0 |
nvidia/riva-translate-4b-instruct-v2 | $0 | $0 |
poolside/laguna-xs-2.1served by a default model | $0 | $0 |
stepfun-ai/step-3.7-flashserved by a default model | $0 | $0 |
thinkingmachines/inklingserved by a default model | $0 | $0 |
z-ai/glm-5.2served by a default model | $0 | $0 |
Premium text models (55)
Billed from your balance at the provider's published rate, cheapest input first. A trial key cannot call these until you add a card. The IDs below already include the provider prefix they need (openai/, anthropic/, or gemini/). Our catalog stores them bare, because that is the key we price against once the prefix has been read off. Sent that way, as gpt-4o rather than openai/gpt-4o, they route to the free NVIDIA tier. The reply does say so: its model field names the model that actually answered, so compare it against the ID you sent. What you will not get is an error or a warning header, so code that never reads response.model will not notice.
| Model ID | Input / M | Output / M |
|---|---|---|
openai/gpt-4.1-nano | $0.10 | $0.40 |
gemini/gemini-2.5-computer-use-preview-10-2025 | $0.15 | $0.60 |
gemini/gemini-2.5-flash | $0.15 | $0.60 |
gemini/gemini-2.5-flash-image | $0.15 | $0.60 |
gemini/gemini-2.5-flash-lite | $0.15 | $0.60 |
gemini/gemini-3-flash-preview | $0.15 | $0.60 |
gemini/gemini-3-pro-image | $0.15 | $0.60 |
gemini/gemini-3-pro-image-preview | $0.15 | $0.60 |
gemini/gemini-3.1-flash-image | $0.15 | $0.60 |
gemini/gemini-3.1-flash-image-preview | $0.15 | $0.60 |
gemini/gemini-3.1-flash-lite | $0.15 | $0.60 |
gemini/gemini-3.1-flash-lite-image | $0.15 | $0.60 |
gemini/gemini-3.1-flash-lite-preview | $0.15 | $0.60 |
gemini/gemini-3.1-pro-preview | $0.15 | $0.60 |
gemini/gemini-3.1-pro-preview-customtools | $0.15 | $0.60 |
gemini/gemini-3.5-flash | $0.15 | $0.60 |
gemini/gemini-3.5-flash-lite | $0.15 | $0.60 |
gemini/gemini-3.6-flash | $0.15 | $0.60 |
gemini/gemini-embedding-001 | $0.15 | $0.60 |
gemini/gemini-embedding-2 | $0.15 | $0.60 |
gemini/gemini-embedding-2-preview | $0.15 | $0.60 |
gemini/gemini-flash-latest | $0.15 | $0.60 |
gemini/gemini-flash-lite-latest | $0.15 | $0.60 |
gemini/gemini-omni-flash-preview | $0.15 | $0.60 |
gemini/gemini-pro-latest | $0.15 | $0.60 |
gemini/gemini-robotics-er-1.6-preview | $0.15 | $0.60 |
gemini/gemini-robotics-er-2-preview | $0.15 | $0.60 |
openai/gpt-4o-mini | $0.15 | $0.60 |
openai/gpt-5.4-nano | $0.20 | $1.25 |
openai/gpt-4.1-mini | $0.40 | $1.60 |
openai/gpt-3.5-turbo | $0.50 | $1.50 |
openai/gpt-5.4-mini | $0.75 | $4.50 |
anthropic/claude-haiku-4-5-20251001 | $0.80 | $4.00 |
openai/o3-mini | $1.10 | $4.40 |
openai/o4-mini | $1.10 | $4.40 |
gemini/gemini-2.5-pro | $1.25 | $10.00 |
openai/gpt-5 | $1.25 | $10.00 |
openai/gpt-4.1 | $2.00 | $8.00 |
openai/gpt-4o | $2.50 | $10.00 |
openai/gpt-5.4 | $2.50 | $15.00 |
anthropic/claude-sonnet-4-5-20250929 | $3.00 | $15.00 |
anthropic/claude-sonnet-4-6 | $3.00 | $15.00 |
openai/gpt-5.5 | $5.00 | $30.00 |
anthropic/claude-fable-5 | $10.00 | $50.00 |
openai/gpt-4-turbo | $10.00 | $30.00 |
openai/o3 | $10.00 | $40.00 |
anthropic/claude-opus-4-1-20250805 | $15.00 | $75.00 |
anthropic/claude-opus-4-5-20251101 | $15.00 | $75.00 |
anthropic/claude-opus-4-6 | $15.00 | $75.00 |
anthropic/claude-opus-4-7 | $15.00 | $75.00 |
anthropic/claude-opus-4-8 | $15.00 | $75.00 |
openai/o1 | $15.00 | $60.00 |
openai/gpt-4 | $30.00 | $60.00 |
openai/gpt-5.5-pro | $30.00 | $180.00 |
openai/o1-pro | $150.00 | $600.00 |
Image, video, and audio (39)
Billed per image or per second of output rather than per token.
| Model ID | Kind | Price |
|---|---|---|
gpt-image-1 | image | $0.04 / image |
imagen-4.0-fast-generate-001 | image | $0.04 / image |
imagen-4.0-generate-001 | image | $0.04 / image |
imagen-4.0-ultra-generate-001 | image | $0.04 / image |
black-forest-labs/flux-1.1-pro | prediction | $0.04 / second |
black-forest-labs/flux-schnell | prediction | $0.0030 / second |
bytedance/latentsync | prediction | $0.0008 / second |
devxpy/cog-wav2lip | prediction | $0.0001 / second |
eleven_multilingual_v2 | prediction | $0.0034 / second |
eleven_sound_effect | prediction | $0.0014 / second |
stability-ai/sdxl | prediction | $0.0023 / second |
stability-ai/stable-video-diffusion | prediction | $0.01 / second |
stable-audio-2 | prediction | $0.0017 / second |
sync/lipsync-2 | prediction | $0.0014 / second |
gemini_omni_flash | video | $0.05 / second |
gen3a_turbo | video | $0.05 / second |
gen4.5 | video | $0.05 / second |
happyhorse_1_0 | video | $0.05 / second |
kling2.5_turbo_pro | video | $0.05 / second |
kling3.0_4k | video | $0.05 / second |
kling3.0_pro | video | $0.05 / second |
kling3.0_standard | video | $0.05 / second |
klingO3_4k | video | $0.05 / second |
klingO3_pro | video | $0.05 / second |
klingO3_standard | video | $0.05 / second |
kwaivgi/kling-v2.1 | video | $0.03 / second |
minimax/video-01 | video | $0.08 / second |
minimax/video-01-live | video | $0.08 / second |
seedance2 | video | $0.05 / second |
seedance2_fast | video | $0.05 / second |
seedance2_mini | video | $0.05 / second |
veo-3.1-fast-generate-preview | video | $0.03 / second |
veo-3.1-generate-preview | video | $0.05 / second |
veo-3.1-lite-generate-preview | video | $0.01 / second |
veo3 | video | $0.05 / second |
veo3.1 | video | $0.05 / second |
veo3.1_fast | video | $0.03 / second |
video-01 | video | $0.08 / second |
video-01-live | video | $0.08 / second |