Models

Every model, with its real price.

130 models reachable from one base URL and one key. 17 of them are $0 input and $0 output. Premium providers bill at our published per-token rates.

Pass an ID below as the model argument exactly as written, except for the 7 marked served by a default model, which we do not route yet and which answer from a Llama model instead. Premium IDs are shown with the provider prefix they need, for example openai/gpt-4o-mini or anthropic/claude-sonnet-4-6. Sent without that prefix, a premium ID is refused with a message naming the exact spelling to use, rather than answered by a different model. This list is generated from the live catalog and was last refreshed on 2026-09-26.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key="ifu_your_inferall_key",
)

resp = client.chat.completions.create(
    model="nvidia/nemotron-3-super-120b-a12b",   # any id from the free table below
    messages=[{"role": "user", "content": "Hello"}],
)

Guides for popular free models

Free open models (17)

$0 input and $0 output, served on NVIDIA NIM. A new trial key can call these without a card. If yours returns a 402 here, your account already started checkout at some point and needs it finished before free calls resume.

Model IDInput / MOutput / M
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0
$0$0

Uncensored models (2)

Chat models without the refusal layer most providers add. Same API and key as everything else here; send the ID exactly as written. Billed from your balance, so a trial key needs the $5 starter pack first. inferall-uncensored has a 262k context window; inferall-uncensored-large has 1M. Reasoning is off by default; pass reasoning_effort to turn it on. You are responsible for how you use the output. Guide and examples.

Model IDInput / MOutput / M
$4.00$12.00
$12.00$20.00

Premium text models (68)

Billed from your balance at the provider's published rate, cheapest input first. A trial key cannot call these until you add a card. The IDs below already include the provider prefix they need (openai/, anthropic/, or gemini/). Our catalog stores them bare, because that is the key we price against once the prefix has been read off. Sent that way, as gpt-4o rather than openai/gpt-4o, they route to the free NVIDIA tier. The reply does say so: its model field names the model that actually answered, so compare it against the ID you sent. What you will not get is an error or a warning header, so code that never reads response.model will not notice.

Model IDInput / MOutput / M
$0.10$0.40
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.15$0.60
$0.20$1.25
$0.40$1.60
$0.50$1.50
$0.75$4.50
$1.00$5.00
$1.10$4.40
$1.10$4.40
$1.25$10.00
$1.25$10.00
$2.00$10.00
$2.00$8.00
served by a default model$2.00$4.00
$2.50$10.00
$2.50$15.00
served by a default model$2.50$5.00
served by a default model$2.50$5.00
served by a default model$2.50$5.00
served by a default model$2.50$5.00
$3.00$15.00
$3.00$15.00
$4.00$20.00
served by a default model$4.00$12.00
served by a default model$4.00$12.00
$5.00$25.00
$5.00$25.00
$5.00$25.00
$5.00$25.00
$5.00$25.00
$5.00$30.00
$10.00$50.00
$10.00$50.00
$10.00$30.00
$10.00$40.00
$15.00$60.00
$30.00$60.00
$30.00$180.00
$150.00$600.00

Image, video, and audio (43)

Billed per image or per second of output rather than per token.

Model IDKindPrice
image$0.04 / image
prediction$0.04 / second
prediction$0.0030 / second
prediction$0.0008 / second
prediction$0.0001 / second
prediction$0.0034 / second
prediction$0.0014 / second
prediction$0.0023 / second
prediction$0.01 / second
prediction$0.0017 / second
prediction$0.0014 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.03 / second
video$0.08 / second
video$0.08 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.05 / second
video$0.03 / second
video$0.05 / second
video$0.01 / second
video$0.05 / second
video$0.05 / second
video$0.03 / second
video$0.08 / second
video$0.08 / second
video$0.05 / second
video$0.05 / second
Get an API key and call a free model