GLM-4.7-FlashX pricing and specs
Efficient GLM model for fast reasoning, coding, and agent workflows
GLM-4.7-FlashX price per 1M tokens
| Type | ₽ / 1M | $ / 1M |
|---|---|---|
| Input tokens | 5,92 ₽ | $0.07 |
| Output tokens | 33,8 ₽ | $0.4 |
| Cache read | 0,85 ₽ | $0.01 |
| Cache write | 0 ₽ | $0 |
Prices in rubles as of 18.09.2026 at the CBR rate: 1 $ = 84.51 ₽.
GLM-4.7-FlashX specs and limits
- Context window
- 200K
- Max output
- 131K
- Input
- text
- Output
- text
- Family
- glm-flash
- Released
- 2026-01-19
- Updated
- 2026-01-19
- Knowledge cutoff
- 2025-04
- Also known as
- glm flashx, glm-4.5 flashx, glm-4.6 flashx, glm-4.7 flashx, glm-5 flashx, glm-5.1 flashx
What GLM-4.7-FlashX can do
- ✓Reasoning
- ✓Tool calling
- —Structured output
- —Attachments / files
- —Vision (images)
- ✓Temperature control
- ✓Open weights
Model weights
Frequently asked questions
How much does the GLM-4.7-FlashX API cost?
Input costs $0.07 per 1M tokens (5.92 ₽) and output costs $0.4 per 1M (33.8 ₽). Zhipu AI bills the prompt and the completion separately, so the price of a call depends on how long both are. Ruble figures use the Bank of Russia rate of 84.51 ₽ per dollar for September 2026.
How much do 1,000 GLM-4.7-FlashX tokens cost?
List prices are quoted per 1M tokens, so 1,000 tokens cost a thousandth of that: divide $0.07 for input and $0.4 for output by 1,000. The same rates in rubles are 5.92 ₽ and 33.8 ₽ per 1M tokens. Both prompt tokens and completion tokens are billed.
How many tokens does GLM-4.7-FlashX hold?
The GLM-4.7-FlashX context window is 200K tokens (200,000). Everything shares that budget: the system prompt, the dialogue history, attached files and the answer the model writes. The figure comes from the Zhipu AI model card as of September 2026.
What are the GLM-4.7-FlashX limits?
The context window is 200K tokens (200,000) and a single response is capped at 131K tokens, so longer output has to be generated in parts. Rate limits are set by Zhipu AI per account and depend on your plan rather than on the model, so they are not listed here.
How is the GLM-4.7-FlashX price in rubles calculated?
Providers quote their tariffs in US dollars, so the ruble amounts here are a conversion at the official Bank of Russia rate of 84.51 ₽ per dollar. For GLM-4.7-FlashX that gives 5.92 ₽ per 1M input tokens and 33.8 ₽ per 1M output tokens. When the rate moves, the ruble price moves with it.
How is GLM-4.7-FlashX different from other GLM models?
GLM-4.7-FlashX belongs to the GLM Flash family at Zhipu AI. Models in the GLM line differ in context size, price per 1M tokens and supported capabilities, so compare them by the numbers on their cards rather than by the name. Every GLM model is listed on the provider page.
Can GLM-4.7-FlashX reason?
Yes, GLM-4.7-FlashX declares a reasoning mode: it works through intermediate steps before it answers. Those steps add output tokens, so a call with a long reasoning chain costs more and takes longer than a plain completion. Capabilities declared for the model: reasoning, tool calling.
What capabilities does GLM-4.7-FlashX support?
The declared capabilities of GLM-4.7-FlashX are: reasoning, tool calling. It accepts text as input, and every attachment consumes tokens from the shared context window of 200K tokens. The list comes from the Zhipu AI model card and is refreshed with the catalog as of September 2026.
Does GLM-4.7-FlashX have open weights?
Yes, GLM-4.7-FlashX ships with open weights, so it can be downloaded and served on your own hardware instead of being used only through the Zhipu AI API. What you may do with it is set by the license shown on the card. Catalog prices cover hosted API access.