Z.aiZ.aiยท๐Ÿ’ฌ Text Generation

GLM 4.7 Flash

๐Ÿง  Try in Intelligence โ†’Try on Venice.ai โ†—
Quick reference
GLM 4.7 Flash โ€” TLDR
  • ๐Ÿง  30B-A3B Mixture-of-Experts reasoning model, roughly 3B active parameters.
  • ๐Ÿ”ง Optimized for agentic coding, tool use, and long-horizon planning.
  • ๐Ÿ“ Context window near 200K tokens in this configuration.
  • ๐Ÿ”’ Runs inside a Trusted Execution Environment with hardware attestation.
  • ๐Ÿ†• Uses interleaved/"preserved" thinking before each tool call.
  • ๐ŸŒ Includes web search and reasoning capabilities; MIT-licensed.
  • ๐Ÿข Built by Z.ai (formerly Zhipu AI), released 2026.
๐Ÿ’ฐ Pricing
โ€”
๐Ÿ“… On Venice since
Apr 13, 2026
133 days ago
Provider

Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family ofโ€ฆ

Read full profile โ†’
13 models on Venice
12 text ยท 1 image
Since Apr 1, 2024

About this model

GLM 4.7 Flash is the efficiency-tier member of Z.ai's GLM 4.7 generation, built as a 30-billion-parameter Mixture-of-Experts model that activates only around 3 billion parameters per token, making it suited to lightweight and local deployment. This particular listing runs the model inside a Trusted Execution Environment, adding hardware attestation evidence so users can independently verify that inference happens in a sealed enclave โ€” a privacy-and-verifiability feature layered on top of the standard weights. It supports a context window approaching 200K tokens and is tuned specifically for coding, tool collaboration, and long-horizon agentic workflows.

Compared with the larger flagship GLM 4.7, this Flash variant trades raw capacity for speed and deployability while staying within the same release family. Against earlier GLM generations such as GLM 4.6, Z.ai positions the 4.7 line around stronger agentic coding, repository-level understanding, and "interleaved thinking," where the model reasons sequentially before each action rather than planning everything upfront.

Within Z.ai's broader catalog, it sits alongside successors like GLM 5 and GLM 5.1, remaining the compact, agent-focused option for developers prioritizing efficiency and verifiable execution.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies โ€” verify critical details against the sources listed above.

Data sources: Venice API ยท HuggingFace ยท Wikipedia โ€” enrichment updated 3d ago