GLM 5.2 was retired by its provider in August 2026, and no spelling of it reaches a live model any more. GLM 5.3 is available: InferAll serves it as z-ai/glm-5.3 at $0 input and $0 output on the open-model tier.
What we measured (2026-09-24, through InferAll)
Three short calls to z-ai/glm-5.3: all answered by GLM 5.3 itself, in 1.2s, 1.6s and 2.8s. Like every model on shared NVIDIA capacity it can be briefly busy; InferAll then answers from another free model and says so in the response's model field, so check that field if the exact model matters.
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="z-ai/glm-5.3",
messages=[{"role": "user", "content": "Write a Python one-liner that reverses the words in a sentence."}],
max_tokens=600,
)
print(response.model) # the model that answered
print(response.choices[0].message.content)
Claude Code
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=z-ai/glm-5.3
claude
Other free models
nvidia/nemotron-3-ultra-550b-a55b is the largest free model we serve (guide), and nvidia/nemotron-3-super-120b-a12b is the default. The live list is at GET /ai/v1/models.
Get a key
Sign up at inferall.ai/keys. New accounts get 25 free calls on open models, no card needed.