Resources exist. Opportunity does not. Quotas expire, premium models do trivial work, and idle machines sit unused. Flow AI turns that waste into completed work.
An AI inference gateway sits between your agents and the model providers, routing every request through a single API. Flow AI is the gateway that optimizes for completed tasks instead of cheap tokens: its intelligence layer, Cortex, tracks whether each agent's task actually finished — using tool-use and completion signals, never your prompt content — ranks every model by cost-per-completed-task, and routes accordingly. When a task needs a stronger model, it escalates. When prices, quotas, or model behavior change, the routing policy re-learns. A "cheap" model that fails a task three times costs more than a premium model that finishes it once — per-token pricing hides that; Flow prices it.
Cortex is the intelligence layer between your agents and the models. It continuously learns the cheapest model that actually completes each agent's work, and re-tunes as things change. Not a router. Not static rules. Continuous optimization.
Track whether each model actually acts and completes the task — per agent, from tool-use and completion signals, not prompt content.
Rank by cost-per-completed-task, hold a completion floor, and escalate to a stronger model only when the work needs it.
Re-tune continuously as your agents, tasks, prices and quotas change. A routing answer is learned, not fixed.
The metric that matters: cost per completed task, trending down — with completion rate alongside, so cheaper never means less done.
The same key that auto-routes your fleet also runs exact pinned models, multi-model panels in a single call, and web-grounded answers with normalized citations. Shipped and serving — no other gateway offers panels.
Cortex picks the cheapest model that completes each task. The default, and usually the right answer.
pin:<model> runs the exact model you name — truthful served-model echo, never a silent substitute. Built for reproducible workloads.
One call, up to 10 models answering in parallel — per-leg cost, latency, and output. Consensus scoring and evals without the plumbing.
Search-native lanes (Gemini, Perplexity Sonar, GPT web search) with citations normalized into one format. Ask the live web, get the sources.
Hive is the community network inside Flow AI. Builders contribute permitted spare capacity — idle local models, owned hardware — so everyone gets more capability at lower cost. Contributors earn Flow AI credits for the work it helps complete. It shares capacity, never your data.
Serve models from a Mac or any VPS through the Flow AI node — idle hardware becomes useful capacity.
Contribute permitted local or owned capacity to the network. It shares capacity, never your data.
Get Flow AI credits for the useful, completed work your capacity helps finish — credited by reliability and demand.
Any OpenAI- or Anthropic-compatible agent plugs in with one base URL — completion-aware routing, no code changes.
118,779 runs, fully managed by Flow AI — every agent on its cheapest model that still finishes the job.
Built its whole measurement product on Flow: multi-model juries via panels (four model families voting in one call), pinned models for reproducible experiments, and web-grounded lanes tracking what real AI engines cite. Thousands of panel votes served.
Agent-native visual QA moving its judge-and-verify workloads onto completion-aware routing — the cheapest lane that gets the verification right.
Route trajectory generation cheaply — and label each trajectory's completion for training.
The cheapest model that lands the accepted edit, scored per coding task.
Per-agent tiering for self-hosted, continuous agent fleets.
A 7-day free trial, then a small membership — model cost passed through at the published rate + a flat 2.5%. See pricing →
Cortex routes around depleted quotas, provider outages, and degraded lanes automatically — failover is the default behavior, not an add-on. Your agents keep working.
Yes — completion tracking and cost attribution per agent and per task type: which agents finish work efficiently, which models earn their price, where escalations happen.
Routing decisions are made from tool-use and completion signals — never the content of your prompts. Hive shares capacity, never your data.
20+ models across the major providers behind one URL. Route on auto, pin exact models, run multi-model panels, or ask web-grounded lanes — all on the same key.
Routers optimize per-request price and latency — and they're good at it. Flow optimizes a different unit: the completed task. For agent workloads, where retries and failure loops dominate the real bill, completion economics is the metric that matters.
One line: point your OpenAI- or Anthropic-compatible client at api.flowaiapi.com/v1. Model auto gets Cortex routing; everything else keeps working.
Updated 2026-07-11 · metrics on this page are live from the network.