Models
moonshotai

Kimi K2

moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts model from Moonshot AI with 1 trillion total parameters (32 billion active per forward pass), optimized for agentic capabilities including advanced tool use, reasoning, and code synthesis. It excels across coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use benchmarks, and supports long-context inference up to 128K tokens.

Tool calling
Context
128K
128,000 tokens
Max output
25K
25,000 tokens
Input burn rate
$1.14
per 1M tokens
Output burn rate
$4.60
per 1M tokens

Quick start

Drop-in requests for the OpenAI-compatible Deva endpoint.

1curl https://api.deva.me/v1/chat/completions \2  -H "Authorization: Bearer $DEVA_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "moonshotai/kimi-k2",6    "messages": [{"role":"user","content":"Hello from Deva"}],7    "stream": true8  }'

Capabilities

Feature metadata advertised for this model.

Tool callingStructured outputReasoningVisionStreaming

Related models

More options from moonshotai and the recommended set.

Browse all

MOONSHOTAI: Kimi K2.6

Tool callingStructured outputReasoningVision

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding across Python, Rust, and Go and turns prompts and visual inputs into production-ready interfaces, with an agent-swarm architecture that scales to hundreds of parallel sub-agents for autonomous task decomposition.

moonshotai262K context$1.37/M in$6.84/M out

MOONSHOTAI: Kimi K3

Reasoning
moonshotai1M context$6/M in$30/M out

X AI: Grok 4.3

Tool callingStructured outputReasoningVision

Grok 4.3 is a reasoning model from xAI that accepts text and image inputs with text output, suited to agentic workflows, instruction-following, and applications requiring high factual accuracy. Reasoning effort is configurable (none, low, medium, or high), and a 1M-token context window with no output limit makes it well-suited to long-document analysis, deep research, and multi-step agentic tasks.

x-ai1M context$2.5/M in$5/M out

ANTHROPIC: Claude Opus 4.7

Tool callingStructured outputReasoningVision

Claude Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on Opus 4.6's coding and agentic strengths, it delivers stronger performance on complex, multi-step tasks such as large codebases, multi-stage debugging, and end-to-end project orchestration, plus improved knowledge work from document drafting to data analysis, maintaining coherence across very long outputs and extended sessions.

anthropic1M context$10/M in$50/M out