Models

Every model, with its real price.

135 models reachable from one base URL and one key. 41 of them are $0 input and $0 output. Premium providers bill at the provider's published rate with zero markup.

Pass an ID below as the model argument exactly as written, except for the 10 marked served by a default model, which we do not route yet and which answer from a Llama model instead. Premium IDs are shown with the provider prefix they need, for example openai/gpt-4o-mini or anthropic/claude-sonnet-4-6. Sent without that prefix an ID routes to the free NVIDIA tier instead of the provider you named. This list is generated from the live catalog and was last refreshed on 2026-08-05.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key="ifu_your_inferall_key",
)

resp = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",   # any id from the free table below
    messages=[{"role": "user", "content": "Hello"}],
)

Free open models (41)

$0 input and $0 output, served on NVIDIA NIM. A new trial key can call these without a card. If yours returns a 402 here, your account already started checkout at some point and needs it finished before free calls resume.

Model IDInput / MOutput / M
baai/bge-m3served by a default model$0$0
deepseek-ai/deepseek-v4-flashserved by a default model$0$0
deepseek-ai/deepseek-v4-proserved by a default model$0$0
google/diffusiongemma-26b-a4b-itserved by a default model$0$0
google/gemma-4-31b-itserved by a default model$0$0
meta/llama-3.1-70b-instruct$0$0
meta/llama-3.1-8b-instruct$0$0
meta/llama-3.2-11b-vision-instruct$0$0
meta/llama-3.2-1b-instruct$0$0
meta/llama-3.2-3b-instruct$0$0
meta/llama-3.2-90b-vision-instruct$0$0
meta/llama-3.3-70b-instruct$0$0
meta/llama-guard-4-12b$0$0
minimaxai/minimax-m3served by a default model$0$0
mistralai/mistral-medium-3.5-128b$0$0
mistralai/mistral-nemotron$0$0
nvidia/ai-synthetic-video-detector$0$0
nvidia/ising-calibration-1.5-31b$0$0
nvidia/llama-3.1-nemoguard-8b-content-safety$0$0
nvidia/llama-3.1-nemoguard-8b-topic-control$0$0
nvidia/llama-3.1-nemotron-nano-8b-v1$0$0
nvidia/llama-3.1-nemotron-nano-vl-8b-v1$0$0
nvidia/llama-3.1-nemotron-safety-guard-8b-v3$0$0
nvidia/llama-3.3-nemotron-super-49b-v1$0$0
nvidia/llama-3.3-nemotron-super-49b-v1.5$0$0
nvidia/nemoretriever-parse$0$0
nvidia/nemotron-3-nano-30b-a3b$0$0
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning$0$0
nvidia/nemotron-3-super-120b-a12b$0$0
nvidia/nemotron-3-ultra-550b-a55b$0$0
nvidia/nemotron-3.5-content-safety$0$0
nvidia/nemotron-mini-4b-instruct$0$0
nvidia/nemotron-nano-12b-v2-vl$0$0
nvidia/nemotron-parse$0$0
nvidia/nvidia-nemotron-nano-9b-v2$0$0
nvidia/riva-translate-4b-instruct-v1.1$0$0
nvidia/riva-translate-4b-instruct-v2$0$0
poolside/laguna-xs-2.1served by a default model$0$0
stepfun-ai/step-3.7-flashserved by a default model$0$0
thinkingmachines/inklingserved by a default model$0$0
z-ai/glm-5.2served by a default model$0$0

Premium text models (55)

Billed from your balance at the provider's published rate, cheapest input first. A trial key cannot call these until you add a card. The IDs below already include the provider prefix they need (openai/, anthropic/, or gemini/). Our catalog stores them bare, because that is the key we price against once the prefix has been read off. Sent that way, as gpt-4o rather than openai/gpt-4o, they route to the free NVIDIA tier. The reply does say so: its model field names the model that actually answered, so compare it against the ID you sent. What you will not get is an error or a warning header, so code that never reads response.model will not notice.

Model IDInput / MOutput / M
openai/gpt-4.1-nano$0.10$0.40
gemini/gemini-2.5-computer-use-preview-10-2025$0.15$0.60
gemini/gemini-2.5-flash$0.15$0.60
gemini/gemini-2.5-flash-image$0.15$0.60
gemini/gemini-2.5-flash-lite$0.15$0.60
gemini/gemini-3-flash-preview$0.15$0.60
gemini/gemini-3-pro-image$0.15$0.60
gemini/gemini-3-pro-image-preview$0.15$0.60
gemini/gemini-3.1-flash-image$0.15$0.60
gemini/gemini-3.1-flash-image-preview$0.15$0.60
gemini/gemini-3.1-flash-lite$0.15$0.60
gemini/gemini-3.1-flash-lite-image$0.15$0.60
gemini/gemini-3.1-flash-lite-preview$0.15$0.60
gemini/gemini-3.1-pro-preview$0.15$0.60
gemini/gemini-3.1-pro-preview-customtools$0.15$0.60
gemini/gemini-3.5-flash$0.15$0.60
gemini/gemini-3.5-flash-lite$0.15$0.60
gemini/gemini-3.6-flash$0.15$0.60
gemini/gemini-embedding-001$0.15$0.60
gemini/gemini-embedding-2$0.15$0.60
gemini/gemini-embedding-2-preview$0.15$0.60
gemini/gemini-flash-latest$0.15$0.60
gemini/gemini-flash-lite-latest$0.15$0.60
gemini/gemini-omni-flash-preview$0.15$0.60
gemini/gemini-pro-latest$0.15$0.60
gemini/gemini-robotics-er-1.6-preview$0.15$0.60
gemini/gemini-robotics-er-2-preview$0.15$0.60
openai/gpt-4o-mini$0.15$0.60
openai/gpt-5.4-nano$0.20$1.25
openai/gpt-4.1-mini$0.40$1.60
openai/gpt-3.5-turbo$0.50$1.50
openai/gpt-5.4-mini$0.75$4.50
anthropic/claude-haiku-4-5-20251001$0.80$4.00
openai/o3-mini$1.10$4.40
openai/o4-mini$1.10$4.40
gemini/gemini-2.5-pro$1.25$10.00
openai/gpt-5$1.25$10.00
openai/gpt-4.1$2.00$8.00
openai/gpt-4o$2.50$10.00
openai/gpt-5.4$2.50$15.00
anthropic/claude-sonnet-4-5-20250929$3.00$15.00
anthropic/claude-sonnet-4-6$3.00$15.00
openai/gpt-5.5$5.00$30.00
anthropic/claude-fable-5$10.00$50.00
openai/gpt-4-turbo$10.00$30.00
openai/o3$10.00$40.00
anthropic/claude-opus-4-1-20250805$15.00$75.00
anthropic/claude-opus-4-5-20251101$15.00$75.00
anthropic/claude-opus-4-6$15.00$75.00
anthropic/claude-opus-4-7$15.00$75.00
anthropic/claude-opus-4-8$15.00$75.00
openai/o1$15.00$60.00
openai/gpt-4$30.00$60.00
openai/gpt-5.5-pro$30.00$180.00
openai/o1-pro$150.00$600.00

Image, video, and audio (39)

Billed per image or per second of output rather than per token.

Model IDKindPrice
gpt-image-1image$0.04 / image
imagen-4.0-fast-generate-001image$0.04 / image
imagen-4.0-generate-001image$0.04 / image
imagen-4.0-ultra-generate-001image$0.04 / image
black-forest-labs/flux-1.1-proprediction$0.04 / second
black-forest-labs/flux-schnellprediction$0.0030 / second
bytedance/latentsyncprediction$0.0008 / second
devxpy/cog-wav2lipprediction$0.0001 / second
eleven_multilingual_v2prediction$0.0034 / second
eleven_sound_effectprediction$0.0014 / second
stability-ai/sdxlprediction$0.0023 / second
stability-ai/stable-video-diffusionprediction$0.01 / second
stable-audio-2prediction$0.0017 / second
sync/lipsync-2prediction$0.0014 / second
gemini_omni_flashvideo$0.05 / second
gen3a_turbovideo$0.05 / second
gen4.5video$0.05 / second
happyhorse_1_0video$0.05 / second
kling2.5_turbo_provideo$0.05 / second
kling3.0_4kvideo$0.05 / second
kling3.0_provideo$0.05 / second
kling3.0_standardvideo$0.05 / second
klingO3_4kvideo$0.05 / second
klingO3_provideo$0.05 / second
klingO3_standardvideo$0.05 / second
kwaivgi/kling-v2.1video$0.03 / second
minimax/video-01video$0.08 / second
minimax/video-01-livevideo$0.08 / second
seedance2video$0.05 / second
seedance2_fastvideo$0.05 / second
seedance2_minivideo$0.05 / second
veo-3.1-fast-generate-previewvideo$0.03 / second
veo-3.1-generate-previewvideo$0.05 / second
veo-3.1-lite-generate-previewvideo$0.01 / second
veo3video$0.05 / second
veo3.1video$0.05 / second
veo3.1_fastvideo$0.03 / second
video-01video$0.08 / second
video-01-livevideo$0.08 / second
Get an API key and call a free model