Routing note
The model IDs deepseek-ai/deepseek-v4-flash, deepseek-ai/deepseek-v4-pro and poolside/laguna-xs-2.1 do not reach those models today. Requests for them are answered by a default Llama model instead, because we do not route those providers yet, and the reply still shows the ID you asked for. We are fixing this. Until then the free IDs under meta/, mistralai/ and nvidia/ on the models page do reach the model they name.
Update (2026-07-24): mistralai/codestral-22b-instruct-v0.1 is no longer available on the free
NVIDIA NIM tier — calls to it now return a 404 upstream. We found this by probing every model in
our catalog with a real inference call rather than trusting a status dashboard, and we'd rather say
so plainly than leave a code sample here that fails for you.
Below are the free, code-capable models that we verified are actually serving today — each was
called successfully before this post was updated. All are $0 input / $0 output within the
free-plan daily request limits.
The free coding models that work right now
| Model id | Good for |
|---|---|
poolside/laguna-xs-2.1 |
Code-specialised, the closest Codestral replacement |
deepseek-ai/deepseek-v4-pro |
Strong reasoning plus code |
deepseek-ai/deepseek-v4-flash |
Fast, low latency |
mistralai/mistral-medium-3.5-128b |
Mistral family, general purpose |
meta/llama-3.3-70b-instruct |
Strong general purpose |
The closest drop-in for what Codestral did, code-specialised and fast, is
poolside/laguna-xs-2.1.
Call it with the OpenAI SDK
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # get one at inferall.ai/keys — no card required
)
response = client.chat.completions.create(
model="poolside/laguna-xs-2.1",
messages=[{
"role": "user",
"content": "Write a Python function that validates email addresses using regex."
}],
max_tokens=512,
)
print(response.choices[0].message.content)
TypeScript is the same shape:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const r = await client.chat.completions.create({
model: "deepseek-ai/deepseek-v4-flash",
messages: [{ role: "user", content: "Refactor this function and explain the change." }],
});
Or curl:
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "poolside/laguna-xs-2.1",
"messages": [{"role": "user", "content": "Explain this stack trace."}]
}'
Get an ifu_... key at inferall.ai/keys — no card needed to start.
Check for yourself which models are live
Upstream availability changes without notice — which is exactly how the Codestral sample here went stale. You can list what's currently offered at any time:
npx @inferall/cli models --free
And confirm a specific model actually answers before you build on it:
npx @inferall/cli chat --model poolside/laguna-xs-2.1 "say hi"
Why the honesty matters: a code sample that 404s wastes your afternoon. We publish model ids we've verified, and when one dies we update the post rather than letting it rank forever on a promise it can't keep.