The Vercel AI SDK talks to InferAll through its OpenAI provider, so one baseURL and one key reach Claude, Gemini, OpenAI and the free NVIDIA models. You keep generateText, streamText, tools and generateObject as they are.
We ran every example here against the live API on 2026-09-27 with ai 7.0.118 and @ai-sdk/openai 4.0.78. Across openai/gpt-4o-mini, nvidia/nemotron-3-super-120b-a12b, anthropic/claude-haiku-4-5-20251001 and gemini/gemini-2.5-flash: text, streaming and a multi-step tool loop passed 24 of 24, and generateObject with a nested schema passed 8 of 8.
Install
npm i ai @ai-sdk/openai zod
export INFERALL_API_KEY=ifu_your_key_here # inferall.ai/keys
generateText (free model)
import { generateText } from "ai";
import { createOpenAI } from "@ai-sdk/openai";
const inferall = createOpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const { text } = await generateText({
model: inferall("nvidia/nemotron-3-super-120b-a12b"),
prompt: "Explain what an API is in one sentence.",
});
console.log(text);
streamText with a tool (Claude)
import { streamText, tool, stepCountIs } from "ai";
import { createOpenAI } from "@ai-sdk/openai";
import { z } from "zod";
const inferall = createOpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const result = streamText({
model: inferall("anthropic/claude-haiku-4-5-20251001"),
prompt: "What is the weather in Paris?",
tools: {
weather: tool({
description: "Current weather for a city",
inputSchema: z.object({ city: z.string() }),
execute: async ({ city }) => `${city}: 17C, light rain`,
}),
},
stopWhen: stepCountIs(3),
});
for await (const chunk of result.textStream) process.stdout.write(chunk);
console.log();
The model called the tool, got its result and answered from it, over two steps.
generateObject (Gemini)
import { generateObject } from "ai";
import { createOpenAI } from "@ai-sdk/openai";
import { z } from "zod";
const inferall = createOpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const { object } = await generateObject({
model: inferall("gemini/gemini-2.5-flash"),
schema: z.object({
items: z.array(z.object({ name: z.string(), qty: z.number().int() })),
note: z.string().nullable(),
}),
prompt: "Order 2 apples and 3 pears, note: fragile.",
});
console.log(object);
This printed { items: [ { name: 'apple', qty: 2 }, { name: 'pear', qty: 3 } ], note: 'fragile' } in our run.
Responses API or Chat Completions
inferall("model-id") uses the Responses API (/v1/responses), which is the provider's default. inferall.chat("model-id") uses Chat Completions (/v1/chat/completions). Both worked on all four models above; pick either.
Choosing a model
Use the full id with its prefix: anthropic/..., gemini/..., openai/... or nvidia/.... The NVIDIA models are free on a new trial key; Claude, Gemini and OpenAI models are paid, so add the $5 starter pack at /billing first. The full list is on the models page.
Get started
- Create a key at inferall.ai/keys.
- Run the first example with
INFERALL_API_KEYset.