GLM-5.2 Fast Preview API via TokenMix

Use GLM-5.2 Fast Preview from Zhipu as a chat model through the TokenMix AI API relay and multi-model gateway.

GLM-5.2 Fast Preview is the high-throughput preview mode of Zhipu AI's GLM-5.2. It provides a 1M-token context window, 80-100 TPS output, controllable reasoning, tool calling, structured output in non-thinking mode, and implicit context caching for latency-sensitive coding, agents, and real-time applications.

API access

  • Base URL: https://api.tokenmix.ai/v1
  • Model ID: glm-5.2-fast-preview
  • OpenAI SDK compatible. Change the base URL and use your TokenMix API key.

Pricing

Input $2.235294/M tokens, output $7.823529/M tokens

Capabilities

Function calling, JSON mode, Streaming, Reasoning

Model specs

  • Context: 1000K tokens
  • Max output: 131K tokens

Availability

1/1 available API endpoints are healthy right now.

Recent performance

TTFT 910ms, latency 1702ms, throughput 99.7 tok/s.

Start using this model

Create an API key, top up from $1 when needed, and call this model through the TokenMix OpenAI-compatible endpoint.

Create API key · View pricing · Quickstart