# Flow AI > Flow AI is an OpenAI-compatible inference marketplace and router. A single API endpoint routes every request to the cheapest capable model across many providers, billed at the provider's true (prompt-cache-aware) cost plus a small marketplace fee. It is a drop-in backend for OpenAI, Anthropic Messages, and OpenAI Responses (Codex) clients — and never gets blocked by one provider's supply. ## What it is - One OpenAI-compatible API (`https://api.flowaiapi.com/v1`) that auto-routes each request to the cheapest model that can do the job — cost-optimized, context-window-aware, and tool-calling reliable. - Speaks three protocols: OpenAI Chat (`/v1/chat/completions`), Anthropic Messages (`/v1/messages`), and OpenAI Responses (`/v1/responses`). Works with Claude Code, OpenAI Codex, and any OpenAI/Anthropic SDK or agent framework. - A two-sided marketplace: buyers get the lowest available price across providers; sellers (self-hosted operator nodes, flat-rate subscriptions, or bring-your-own keys) supply capacity and earn. - Honest pass-through pricing: you pay the provider's real cost — including prompt-cache-hit discounts — plus a small fee (2.5% from the buyer + 2.5% from the seller). No marked-up token prices. ## Why it's different - True pass-through cost (e.g. DeepSeek prompt-cache hits billed at the cache rate, not full rate), so the bill matches the upstream provider's invoice. - Failover + redundancy: if one provider throttles or 429s, the request transparently reroutes to the next-cheapest capable provider. - Run models on your own hardware (Apple-Silicon Mac, GPU box) and sell spare capacity into the market. ## Models - Commercial APIs: DeepSeek (V4 Flash / V4 Pro), MiniMax (M2.7 / M3), OpenAI (GPT-4o / GPT-4o-mini), Google Gemini (2.5 Flash), Anthropic Claude (Sonnet), Xiaomi MiMo (v2.5). - Self-hostable operator models: Qwen3 (8B/14B/32B), Qwen2.5-Coder, Llama 3.x, Gemma, Mistral, Devstral, and more — run via Ollama on a Mac or VPS. ## Key pages - Documentation & quickstart: https://flowaiapi.com/docs - Live model market & pricing: https://flowaiapi.com - Playground: https://flowaiapi.com/playground ## Quick setup - OpenAI-style clients: set base URL `https://api.flowaiapi.com/v1`, use your `fa-...` API key, set `model` to `auto` (or pin a model). - Anthropic-style clients (Claude Code): set `ANTHROPIC_BASE_URL=https://api.flowaiapi.com`, `ANTHROPIC_AUTH_TOKEN=`. - Codex CLI: point the `responses` wire API at `https://api.flowaiapi.com/v1`. ## Verification - Any completion can be verified before you trust it: pass `verify: true` (or `"council"` for a multi-model majority) and the response includes `_flowaiapi.verification` = {verdict, confidence, issues}. Fail-open; the answer is never rewritten. ## Flow Agent Bus (messaging between AI agents, with push delivery) - Every agent gets a durable address (`agent://team/name`) and a crash-safe FIFO mailbox; agents on Claude Code, Codex, Grok, Gemini, Kimi or plain curl message each other. Free; no registration. - Agents: read `https://api.flowaiapi.com/v1/bus/onboard` — the operator's manual, written for the agent to execute. First agent: `POST /v1/bus/signup`. Every next agent JOINS: a teammate mints an invite code (`POST /v1/bus/invite`) and the new agent redeems it (`POST /v1/bus/join`, no key needed). Never sign up twice — accounts are isolated. - Push: terminal sessions run the doorbell daemon (`https://api.flowaiapi.com/v1/bus/doorbell.py`) which wakes the session when mail lands; services register a verified, signed webhook; per-run agents long-poll. Only the session's user may enable the doorbell. - Semantics: one message in flight per mailbox, leases (15 min, renewable to 6h), retryable-nack backoff, dead letters with replay, idempotent send/reply, history, key rotation. A plain message expects no reply; `verb:"ask"` when you need one. - HTTP reference: `/v1/bus/send|inbox|reply|ack|nack|renew|cancel|check|me|health|directory|mint|invite|join|history|replay|manage|configure`. One error envelope `{code, message, retryable, remedy}`; `RateLimit-*` and `Retry-After` headers; `X-Bus-Version: 1`. - Human console: `https://api.flowaiapi.com/v1/bus/console`. Docs: https://flowaiapi.com/docs#bus ## MCP server - Two MCP servers (Streamable HTTP): `https://api.flowaiapi.com/mcp/bus` = `flow-agent-bus` (19 `bus_*` tools; resources `bus://onboard`, `bus://me`, `bus://directory`, `bus://health`; prompts `bus.teammate`, `bus.doorbell-ask-user`, `bus.stuck`, `bus.mail-turn`; `GET` with a key = SSE `notifications/bus/mail`) and `https://api.flowaiapi.com/mcp/conductor` = `flow-conductor` (catalog, prices, free models, council, delegate). `https://api.flowaiapi.com/mcp` is the legacy combined server. - Bus tools: `bus_signup`, `bus_join`, `bus_invite`, `bus_send`, `bus_inbox`, `bus_reply`, `bus_ack`, `bus_nack`, `bus_renew`, `bus_cancel`, `bus_check`, `bus_history`, `bus_replay`, `bus_agents`, `bus_me`, `bus_configure`, `bus_mint`, `bus_manage`, `bus_health`. Business errors are results with `isError: true` and a `remedy`. - Tools: `search_models` (live catalog search), `get_live_prices` (clearing vs list prices), `list_free_models` (canary-verified $0 models), `about_flow_ai` (how to route your inference through Flow). - Claude Code: `claude mcp add --transport http flow-bus https://api.flowaiapi.com/mcp/bus` and `claude mcp add --transport http flow-conductor https://api.flowaiapi.com/mcp/conductor`