← All models

glm-5.3-flash

Run glm-5.3-flash through Flow AI's OpenAI-compatible API. We route to the cheapest provider hosting it and fail over automatically — you pay the live clearing price plus 2.5%.

Input $/M
—
Output $/M
—

Use it

python
from openai import OpenAI
client = OpenAI(base_url="https://api.flowaiapi.com/v1", api_key="fa-...")
r = client.chat.completions.create(
    model="glm-5.3-flash",          # or "auto" to let Flow AI pick the cheapest capable model
    messages=[{"role": "user", "content": "..."}],
)

How Flow AI makes glm-5.3-flash cheaper

Flow AI runs an open market for inference: it clears each request at the lowest price across every provider hosting glm-5.3-flash — direct APIs, inference providers, flat-rate subscriptions, and self-hosted nodes — and bills you the true (prompt-cache-aware) cost plus a flat 2.5%. The published rate is always the ceiling. See the docs or all models.

Similar models on Flow AI

nemotron-3-ultra · or:inclusionai/ling-3.0-flash-vl:free · or:nvidia/nemotron-3-super-120b-a12b:free · or:inclusionai/ling-3.1-flash · or:nvidia/nemotron-3-ultra-550b-a55b:free · glm-4.7 · glm-5.1 · glm-5.2

glm-5.3-flash API — the cheapest way to run glm-5.3-flash | Flow AI