All models

Qwen3 32B vs GLM 4.7 Flash

Qwen3 32B at $0.092 in and $0.322 out and GLM 4.7 Flash at $0.069 in and $0.460 out — per million tokens, from the same wallet.

Q

Qwen3 32B

Qwen

0.07¢

for a client proposal

$0.092 in · $0.322 out / 1M

41K context · released April 2025

Full page →
Z

GLM 4.7 Flash

Z.ai

0.07¢

for a client proposal

$0.069 in · $0.460 out / 1M

203K context · released 7 months ago

Full page →

Asking all both at once — identical prompt, identical context — costs about 0.14¢ for a typical client proposal.

At a glance

  • Cheapest input: GLM 4.7 Flash at $0.069 per million tokens.
  • Cheapest output: Qwen3 32B at $0.322 per million tokens.
  • Largest context: GLM 4.7 Flash at 203K tokens — about 290 pages.

Cost per task

Estimated totals at the rates above. Solid bar is input, lighter bar is output.

Quick code fix

2K in · 500 out
Qwen3 32B
0.03¢
GLM 4.7 Flash
0.04¢

Draft a client proposal from notes

5K in · 800 out
Qwen3 32B
0.07¢
GLM 4.7 Flash
0.07¢

Context-heavy session (Alyph workspace)

155K in · 4K out
Qwen3 32B
needs 155K ctx
GLM 4.7 Flash
1.3¢

Analyze a 50-page legal contract

35K in · 1K out
Qwen3 32B
0.35¢
GLM 4.7 Flash
0.29¢

Large architecture refactor

250K in · 3K out
Qwen3 32B
needs 250K ctx
GLM 4.7 Flash
needs 250K ctx

Solid part of each bar is input tokens, the lighter part is output. Models that don’t fit a task are marked instead of priced.

Cost vs Token Amount

Drag the slider to adjust the split between input and output workload.

80% Input20% Output
free$0.07$0.150 tokens250k500k750k1000k
Qwen3 32B
GLM 4.7 Flash

Adjust the ratio slider to change how output tokens influence the final price for equivalent workloads.

Pricing

Qwen3 32B
GLM 4.7 Flash
Input / 1M tokens
$0.092
$0.069
Output / 1M tokens
$0.322
$0.460
Cached input / 1M
$0.011

Specs

Qwen3 32B
GLM 4.7 Flash
Provider
Released
April 2025
7 months ago
Context window
41K
203K
Max output
16K
16K
Input
Text
Text
Output
Text
Text
Reasoning
Shows its thinking
Shows its thinking
Knowledge cutoff
2025-03-31

Try it with real numbers

2K tokens in · 500 tokens out.

Qwen3 32B

Qwen

0.03¢

for this task

53% input · 47% output

Full pricing →

GLM 4.7 Flash

Z.ai

0.04¢

for this task

38% input · 62% output

Full pricing →

All 2, same prompt and context: 0.07¢. That is the entire cost of the comparison.

Your $5 welcome credit covers about 7,012 of these.

Which should you pick?

For long documents, GLM 4.7 Flash has the largest context window here — 203K tokens.

On price, GLM 4.7 Flash is the cheapest of the two on a typical task, at about 0.07¢.

Or don’t pick. Send the identical prompt to all both, read the answers side by side, and keep the winner. How to run a fair bake-off →

Questions, answered

Which is cheaper, Qwen3 32B or GLM 4.7 Flash?

On a typical client proposal (5K tokens in, 800 out): GLM 4.7 Flash at 0.07¢; Qwen3 32B at 0.07¢. Long prompts can change the order when long-context rates apply.

Which has the largest context window?

GLM 4.7 Flash — 203K tokens, about 290 pages.

Can I run Qwen3 32B and GLM 4.7 Flash side by side on Alyph?

Yes — that is what Alyph is built for: the identical prompt with identical context to all 2, answers rendered side by side, billed from one wallet. This exact combination costs about 0.14¢ for a typical task.

Ask both at once

Identical prompt, identical context, answers side by side — about 0.14¢ for a typical client proposal. Free to start, with $5 of credit.