deepinfra-gemma-4-26b-a4b-it · Google
A Mixture-of-Experts model that activates only 4B parameters per inference,delivering high-performance reasoning with a fraction of the memory cost - idealfor cost-efficient, high-throughput server deployments.
A Mixture-of-Experts model that activates only 4B parameters per inference,delivering high-performance reasoning with a fraction of the memory cost - idealfor cost-efficient, high-throughput server deployments.
deepinfra-gemma-4-26b-a4b-it has a 262,100 token context window.
On AIHubMix, deepinfra-gemma-4-26b-a4b-it costs $0.088 per million input tokens and $0.385 per million output tokens. Cached input reads are billed at $0.011 per million tokens.
deepinfra-gemma-4-26b-a4b-it is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to deepinfra-gemma-4-26b-a4b-it — no other code changes needed.
deepinfra-gemma-4-26b-a4b-it is developed by Google. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Gemini 3.7 Flash is Google’s natively multimodal reasoning model for coding, agents, web…
Gemini 3.7 Flash free version: fFree model resources are limited and provided only for…
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…
Google's newest, most compact, and most cost-effective image generation and editing…
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for…
Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only…
Use deepinfra-gemma-4-26b-a4b-it via the AIHubMix unified API — one interface for every major LLM.