gemma2-2bself-hostable
Run gemma2-2b through Flow AI's OpenAI-compatible API. We route to the cheapest provider hosting it and fail over automatically — you pay the live clearing price plus 2.5%.
Input $/M
$0.04
Output $/M
$0.08
Use it
from openai import OpenAI
client = OpenAI(base_url="https://api.flowaiapi.com/v1", api_key="fa-...")
r = client.chat.completions.create(
model="gemma2-2b", # or "auto" to let Flow AI pick the cheapest capable model
messages=[{"role": "user", "content": "..."}],
)How Flow AI makes gemma2-2b cheaper
Flow AI runs an open market for inference: it clears each request at the lowest price across every provider hosting gemma2-2b — direct APIs, inference providers, flat-rate subscriptions, and self-hosted nodes — and bills you the true (prompt-cache-aware) cost plus a flat 2.5%. The published rate is always the ceiling. See the docs or all models.
Similar models on Flow AI
or:inception/mercury-2.5-preview · or:sao10k/l3-lunaris-8b · or:cohere/command-r7b-12-2024 · or:openai/gpt-oss-120b · or:tencent/hy-mt2-1.8b