Gemini-2.5-flash defaults to thinking enabled; to disable thinking, request the name gemini-2.5-flash-nothink, which only supports OpenAI-compatible format calls and does not support Gemini SDK; for the native Gemini SDK, please set the parameter budget=0 directly.
Gemini-2.5-flash defaults to thinking enabled; to disable thinking, request the name gemini-2.5-flash-nothink, which only supports OpenAI-compatible format calls and does not support Gemini SDK; for the native Gemini SDK, please set the parameter budget=0 directly.
gemini-2.5-flash-nothink has a 1,047,576 token context window.
On AIHubMix, gemini-2.5-flash-nothink costs $0.3 per million input tokens and $2.499 per million output tokens. Cached input reads are billed at $0.03 per million tokens.
gemini-2.5-flash-nothink accepts text, image, audio and video input.
gemini-2.5-flash-nothink supports tool calling, function calling, structured outputs and long context. Per-protocol parameter support is listed in the capability table on this page.
gemini-2.5-flash-nothink is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gemini-2.5-flash-nothink — no other code changes needed.
gemini-2.5-flash-nothink is developed by Google. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Gemini 3.7 Flash is Google’s natively multimodal reasoning model for coding, agents, web…
Gemini 3.7 Flash free version: fFree model resources are limited and provided only for…
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…
Google's newest, most compact, and most cost-effective image generation and editing…
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for…
Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only…
Use gemini-2.5-flash-nothink via the AIHubMix unified API — one interface for every major LLM.