Kimi K3 API
Call Kimi K3 with a Wokey API key through an OpenAI-compatible API. Find key setup steps, the base URL, model ID, a request example, and usage prices on this page.
Price check: official, OpenRouter and Wokey
USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-23, excluding tax; the Wokey rate is read from the live price list.
| Channel | Input | Cache read | Output |
|---|---|---|---|
| Official API | $3 | $0.3 | $15 |
| OpenRouter | $3 | $0.3 | $15 |
| Wokey | $0.9 | $0.09 | $4.5 |
The cheapest OpenRouter host is Sail Research's fp4-quantized deployment ($1.4989 input / $10.758 output). Quantized output can differ from the official weights, so it is not in the table.
· Official pricing page · OpenRouter endpoint list
Use it in Claude Code
This model supports /v1/messages on Wokey. Set these variables, then start claude.
export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="kimi-k3"Kimi K3 API quickstart
Get a Wokey API key and follow these three steps to send your first request.
Get an API key
Register or sign in to Wokey, open the API console, and create and copy a key in API key management. Check your available balance before calling.
Configure your client
Use the base URL below in your OpenAI-compatible client, enter your Wokey API key, and select kimi-k3 as the model.
Send your first request
Replace YOUR_WOKEY_API_KEY in the example with your own key and run it in a terminal. Read the answer in choices, then review usage and charges in the API console.
Get a Wokey API key · Read API docs
- Client base URL
https://api.wokey.ai/v1- Model ID
kimi-k3- Full request URL
https://api.wokey.ai/v1/chat/completions
The client base URL includes /v1. The curl example uses the full request URL.
Minimal request
export WOKEY_API_KEY="YOUR_WOKEY_API_KEY"
curl https://api.wokey.ai/v1/chat/completions \
-H "Authorization: Bearer $WOKEY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "Hello, Kimi!"}],
"max_tokens": 1024
}'- Model ID
kimi-k3- Vendors
- moonshot, volcengine
- Supply status
- Currently callable
Catalog pricing snapshot
Prices are per 1M tokens; request-time pricing controls settlement.
| Meter | Wokey | Official reference | Savings |
|---|---|---|---|
| Input | $0.9 | $3 | 70% |
| Output | $4.5 | $15 | 70% |
| Cache read | $0.09 | $0.3 | 70% |
| Cache write | — | — | — |
Official price source: https://platform.kimi.ai/docs/pricing/chat-k3
Model capabilities
- Context
- 1,048,576 tokens
- Max output
- 1,048,576 tokens
- Streaming
- Supported
- Tools
- Supported
- Vision
- Supported
- Supported API forms
/v1/chat/completions·/v1/messages
Send a request
/v1/chat/completions
curl https://api.wokey.ai/v1/chat/completions \
-H "Authorization: Bearer $WOKEY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"kimi-k3","messages":[{"role":"user","content":"Hello"}]}'/v1/messages
curl https://api.wokey.ai/v1/messages \
-H "x-api-key: $WOKEY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"kimi-k3","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'Use Kimi K3 in your tools
The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is kimi-k3. Open the guide for the client you use.
- Use Kimi K3 in Claude Code: Anthropic Messages · base URL and API key setup
- Use Kimi K3 in Claude Desktop: Anthropic Messages · base URL and API key setup
- Use Kimi K3 in opencode: OpenAI Chat Completions · base URL and API key setup
- Use Kimi K3 in OpenClaw: OpenAI Chat Completions · base URL and API key setup
- Use Kimi K3 in Hermes Agent: OpenAI Chat Completions · base URL and API key setup
- OpenRouter alternative: Wokey vs OpenRouter: Compare token prices, payment fees, API forms, and upstream sources.
- Verifiable AI API: which models and API forms carry response proofs: Response proofs currently cover only raw passthrough Claude Messages and GPT Responses on official routes.
Example production call
These usage values come from a completed Kimi K3 call that succeeded on its first attempt and was reconciled with its bill.
- Total input
- 76023
- Output
- 574
- Cached share of input
- 99.3%
- Call duration
- 25.60 s
Chat Completions usage
Compiled from the billing record into API field examples. prompt_tokens is total input, including cached tokens.
{
"model": "kimi-k3",
"usage": {
"prompt_tokens": 76023,
"prompt_tokens_details": {
"cached_tokens": 75520
},
"completion_tokens": 574,
"total_tokens": 76597
}
}Token accounting
These are the final per-million-token rates recorded for this call.
| Meter | Tokens | Rate for this call / 1M |
|---|---|---|
| Uncached input | 503 | $0.9 |
| Cache read | 75520 | $0.09 |
| Output | 574 | $4.5 |
- Billed amount
- $0.009833
Cost = sum of tokens × their rate ÷ 1,000,000, rounded to six decimals after summing.
Source of this call
This call used the Volcengine channel at ark.cn-beijing.volces.com with API Key authorization. Usage was metered upstream and the request returned HTTP 200.
About Provider Node · Integration docsVerify Prompt Cache
- Keep the shared prefix of long prompts identical and reuse context according to the selected API’s cache rules.
- Check prompt_tokens_details.cached_tokens; zero means no cache reads were recorded for this call.
- Uncached input = prompt_tokens minus cached tokens. Price each bucket separately and add the output charge.
Frequently asked questions
How do I get an API key for Kimi K3?
To call Kimi K3 through Wokey, use a Wokey API key. Register or sign in, open the /api console, then create and copy a key in API key management. Make sure your account has available balance before calling.
What base URL and model ID should I use for Kimi K3?
In an OpenAI-compatible client, use https://api.wokey.ai/v1 as the base URL and kimi-k3 as the model ID. The full HTTP endpoint is https://api.wokey.ai/v1/chat/completions. Pass your Wokey API key using Authorization: Bearer.
How do a Kimi web subscription and a Wokey API key work?
This example uses an API key created in your Wokey account and the Wokey gateway. Usage is billed against your Wokey balance. Kimi web subscriptions and keys issued by other platforms are managed by their respective platforms.
How much does the Kimi K3 API cost, and how much does it save versus the official API?
Wokey lists input at $0.9 per 1M tokens and output at $4.5 per 1M tokens; the official references are $3 and $15. That is 70% lower for input and 70% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.
Which API forms does Kimi K3 support, and how do I call it through Wokey?
Kimi K3 supports /v1/chat/completions (Chat Completions requests) and /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID kimi-k3; you can start with /v1/chat/completions. The Send a request section includes a curl example for every supported API form.
How do I use Kimi K3 in Claude Code?
Claude Code calls Kimi K3 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="kimi-k3", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.
What are the context window and maximum output for Kimi K3?
The catalog context window is 1,048,576 tokens and the maximum output is 1,048,576 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.
How are cache reads and cache writes priced for Kimi K3?
Cache read: Wokey lists $0.09 per 1M tokens and the official reference is $0.3 per 1M tokens. Cache write: the vendor does not publish a separate meter, so Wokey also shows — as the public price.
How does calling Kimi K3 through Wokey differ from OpenRouter?
Wokey lists Kimi K3 at $0.9 input / $4.5 output per 1M tokens, against an official reference of $3 / $15. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.
How do Kimi K3 and Kimi K2.8 Preview differ in price and context?
In the published price comparison, Kimi K3 lists input/output at $0.9 / $4.5 with a 1,048,576-token context window; Kimi K2.8 Preview lists $0.3 / $1.2 with a 1,048,576-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.
Related models and documentation
Get API key · kimi-k3