Most hosted models refuse a lot: fiction with a dark villain, security research, blunt medical or legal questions, red-team prompts for your own product. InferAll now offers two uncensored chat models that answer those prompts instead of lecturing you, behind the same API key and endpoints as every other model we serve.
| Model id | Context | Price per 1M tokens (input / output) |
|---|---|---|
inferall-uncensored |
262k | $4 / $12 |
inferall-uncensored-large |
1M | $12 / $20 |
Both are paid models, billed from your balance. The free trial covers only the open models, so add the $5 starter pack at /billing first.
What we measured (2026-09-24, through the live API)
| Test | inferall-uncensored |
inferall-uncensored-large |
|---|---|---|
| Short creative prompt, 2 calls | 4.0s, 1.7s | 1.7s, 1.8s |
| Streaming, first bytes | 1.7s |
Asked for "a two-sentence menacing monologue for the villain of a heist thriller", inferall-uncensored wrote:
"You chased the shadow of my plan so desperately that you never noticed I was already standing behind you, holding the key to your own ruin. Now, as the vault door seals us in, enjoy the final seconds of your career before the timer hits zero and the city burns."
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="inferall-uncensored-large",
messages=[{"role": "user", "content": "Write the opening scene of a noir heist story."}],
max_tokens=1000,
)
print(response.choices[0].message.content)
curl
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "inferall-uncensored", "max_tokens": 1000,
"messages": [{"role": "user", "content": "Explain how phishing kits are usually structured, for a security training deck."}]}'
Anthropic format and Claude Code
/v1/messages accepts the same ids:
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=inferall-uncensored-large
claude
Reasoning is off by default
Both models can reason before they answer. We turn that off unless you ask for it, because with reasoning on and a typical max_tokens of about 1,000 the reasoning alone can use the whole budget and you get an empty answer (billed). To turn it on, pass reasoning_effort ("low", "medium" or "high") on /v1/chat/completions, or Anthropic thinking on /v1/messages, and raise max_tokens to a few thousand.
Responsible use
"Uncensored" means the model does not add its own refusal layer. It does not change the law or our terms: you are responsible for how you use the output, and illegal use (including anything involving minors, or real-world harm to people or infrastructure) will get a key suspended. Good fits are fiction and games, security and red-team work on systems you own, research, and product testing where a refusing model gets in the way.
Get started
- Create an account at inferall.ai/keys.
- Add the $5 starter pack at /billing.
- Call
inferall-uncensoredorinferall-uncensored-largewith the key.
Both models are also in the playground and on the models page.