GLM-4.7-Flash pricing and specs

open weights

Budget GLM lane for fast coding help, routing, and everyday automation

0 ₽input / 1M$0
0 ₽output / 1M$0
200Kcontexttokens
131Kmax outputtokens

GLM-4.7-Flash price per 1M tokens

Type₽ / 1M$ / 1M
Input tokens0 ₽$0
Output tokens0 ₽$0
Cache read0 ₽$0
Cache write0 ₽$0

Prices in rubles as of 18.09.2026 at the CBR rate: 1 $ = 84.51 ₽.

GLM-4.7-Flash specs and limits

Context window
200K
Max output
131K
Input
text
Output
text
Family
glm-flash
Released
2026-01-19
Updated
2026-01-19
Knowledge cutoff
2025-04
Also known as
glm flash, glm-4.5 flash, glm-4.6 flash, glm-4.7 flash, glm-5 flash, glm-5.1 flash

What GLM-4.7-Flash can do

  • Reasoning
  • Tool calling
  • Structured output
  • Attachments / files
  • Vision (images)
  • Temperature control
  • Open weights

GLM-4.7-Flash benchmarks

BenchmarkScoreMetric
SWE-Bench Verified59.2resolved

Model weights

Frequently asked questions

How much does the GLM-4.7-Flash API cost?

Zhipu AI lists a zero token price for GLM-4.7-Flash: both input and output are marked free in the price feed, so converting to rubles at the Bank of Russia rate of 84.51 ₽ per dollar changes nothing. That covers token billing only — quotas, request limits and access terms are set by the provider.

How many tokens does GLM-4.7-Flash hold?

The GLM-4.7-Flash context window is 200K tokens (200,000). Everything shares that budget: the system prompt, the dialogue history, attached files and the answer the model writes. The figure comes from the Zhipu AI model card as of September 2026.

What are the GLM-4.7-Flash limits?

The context window is 200K tokens (200,000) and a single response is capped at 131K tokens, so longer output has to be generated in parts. Rate limits are set by Zhipu AI per account and depend on your plan rather than on the model, so they are not listed here.

How is GLM-4.7-Flash different from other GLM models?

GLM-4.7-Flash belongs to the GLM Flash family at Zhipu AI. Models in the GLM line differ in context size, price per 1M tokens and supported capabilities, so compare them by the numbers on their cards rather than by the name. Every GLM model is listed on the provider page.

Can GLM-4.7-Flash reason?

Yes, GLM-4.7-Flash declares a reasoning mode: it works through intermediate steps before it answers. Those steps add output tokens, so a call with a long reasoning chain costs more and takes longer than a plain completion. Capabilities declared for the model: reasoning, tool calling.

What capabilities does GLM-4.7-Flash support?

The declared capabilities of GLM-4.7-Flash are: reasoning, tool calling. It accepts text as input, and every attachment consumes tokens from the shared context window of 200K tokens. The list comes from the Zhipu AI model card and is refreshed with the catalog as of September 2026.

Does GLM-4.7-Flash have open weights?

Yes, GLM-4.7-Flash ships with open weights, so it can be downloaded and served on your own hardware instead of being used only through the Zhipu AI API. What you may do with it is set by the license shown on the card. Catalog prices cover hosted API access.

Other Zhipu AI models