Models
googleRecommended

Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is available through the Deva OpenAI-compatible API with transparent context, capability, and pricing metadata.

Reasoning
Context
1M
1,048,576 tokens
Max output
25K
25,000 tokens
Input burn rate
$3.00
per 1M tokens
Output burn rate
$15.00
per 1M tokens

Quick start

Drop-in requests for the OpenAI-compatible Deva endpoint.

1curl https://api.deva.me/v1/chat/completions \2  -H "Authorization: Bearer $DEVA_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "google/gemini-3.6-flash",6    "messages": [{"role":"user","content":"Hello from Deva"}],7    "stream": true8  }'

Capabilities

Feature metadata advertised for this model.

Tool callingStructured outputReasoningVisionStreaming

Related models

More options from google and the recommended set.

Browse all

GOOGLE: Gemini 2

google120K context$0.2/M in$0.8/M out

GOOGLE: Gemini 2.5 Pro

Tool callingStructured outputReasoningVision

Gemini 2.5 Pro is Google's state-of-the-art model for advanced reasoning, coding, mathematics, and scientific tasks. Its 'thinking' capabilities let it reason through responses with greater accuracy and nuanced context handling, achieving top-tier results across benchmarks, including a first-place position on the LMArena leaderboard.

google1.1M context$2.5/M in$20/M out

GOOGLE: Gemini 3 Flash

Tool callingStructured outputReasoningVision

Gemini 3 Flash is a high-speed, high-value thinking model designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near-Pro reasoning and tool use at substantially lower latency, with a 1M-token context window and multimodal inputs including text, images, audio, video, and PDFs. It supports configurable thinking levels, structured output, tool use, and automatic context caching.

google1M context$1/M in$6/M out

GOOGLE: Gemini 3.1 Pro

Tool callingStructured outputReasoningVision

Gemini 3.1 Pro is Google's frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window, and introduces a medium thinking level to balance cost, speed, and performance. It excels at agentic coding, structured planning, and multimodal analysis.

google1M context$4/M in$24/M out