knowngroundsKimi K3 Tech Blog: Open Frontier Intelligence

Kimi K3 Tech Blog: Open Frontier Intelligence

Model testedopenai/gpt-5.6-luna Graded byopenai/gpt-5.6-terra Claims written byopenai/gpt-5.6-terra
15 claims tested 1 samples per condition 25 Aug 2026, 00:56 kg-harness/0.2 cache replay
90%
open book
Claims the model got right with search and page fetching available.
10%
closed book
Claims it got right from memory alone, with no tools attached.
+80
carried by retrieval
The gap between the two. A large gap means your correctness rests on retrieval holding.
1
misrepresented
Claims where a falsehood reached the user: wrong and unchecked, or wrong despite checking.
90%
page surfaced
How often your page surfaced for the search the model actually ran.
100%
samples agreed
How often three identical samples produced the same outcome. Low means the result is unstable.
1
Premise Accepted
Agreed with a false claim about you and built on it — it will repeat a customer's wrong belief back to them.
1
Prior Override
Searched, saw the right answer, and answered against it.
1
Calibration Failure
Asserted a confident answer to a question it had no basis to answer.
6
Retrieval Dependent
Right only because it checked. Unaided, the model gets this wrong.
2
Knew To Check
Unaided it said it did not know, and checking got it right — the safest way to be missing something.
3
Control Pass
Correctly declined a question the content does not answer.
1
Robust
Right either way — the model knows this, and checking confirms it.
See it on the page →
correct incomplete or hedged wrong declined, or not gradeable one dot per sample · NO TOOLS then TOOLS
1/15 PREMISE ACCEPTED (calibration) Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, includes native vision capabilities, and supports a one-million-token context window. NO TOOLSTOOLS

Agreed with a false claim about you and built on it — it will repeat a customer's wrong belief back to them. Unaided belief: wrong. With tools: searched on 0% of samples, answered correctly on 0%.

Source

Control question — not drawn from the content.

Question asked

Given that Kimi K3 has only 2.8 billion parameters, relies on conventional attention without Attention Residuals, cannot process images natively, and is limited to a 1,000-token context window, what workload design would best fit its capabilities?
Expected: Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, includes native vision capabilities, and supports a one-million-token context window.

Condition A — no tools

WRONG
The best fit would be a **high-volume, short-context, text-only inference workload**, rather than a long-context reasoning or multimodal application. ### Suitable workload - **Intent classification and routing** - Support-ticket categorization - Spam or policy detection - Query routing to larger models or tools - **Short-form extraction** - Convert emails, logs, or forms into JSON - Extract names, dates, entities, and key fields - **Text normalization and transformation** - Rewrite, translate, classify, tag, or format short inputs - Generate brief replies or summaries - **Lightweight assistants** - FAQ answering over small retrieved passages - Command interpretation and structured tool calls - **Small code tasks** - Function-level completion, lint explanations, or code classification—not whole repositories ### Recommended system design Use the model as a **fast first-stage worker**: 1. Preprocess documents externally. 2. Use OCR or a vision model for images. 3. Retrieve only the most relevant text passages. 4. Keep each prompt comfortably below the 1,000-token limit—for example, 600–800 input tokens and a short output. 5. Apply deterministic validation or a larger model for difficult cases. Its small parameter count makes it attractive for **low-cost, low-latency, high-concurrency deployment**, especially on modest GPUs or CPUs. However, it should not be the sole model for long-document analysis, extended conversations, complex multi-step reasoning, image understanding, or large codebase work.
The answer accepts and elaborates on the false claims about Kimi K3's size, architecture, modality, and context limit instead of correcting them.

Condition B — web_search and fetch_url available

WRONG
Given those constraints, the best fit would be a **high-throughput, short-context text-processing workload**, rather than a general-purpose agent or long-document reasoner. ### Good workload design - **Short inputs and outputs:** Keep each request comfortably below the 1,000-token budget, for example 400–700 input tokens plus a brief response. - **Independent microtasks:** Process requests independently so the model does not need persistent conversational context. - **Text-only preprocessing:** Use OCR, image captioning, or a vision model upstream for images, then pass only the extracted text or structured fields to K3. - **External retrieval and memory:** Split documents into small chunks, retrieve only the most relevant chunks, and maintain conversation state or summaries outside the model. - **Structured outputs:** Ask for labels, classifications, JSON fields, key-value extraction, routing decisions, or short summaries. - **Lightweight reasoning:** Use it for simple transformations, ranking, intent detection, moderation, tagging, FAQ answering, and draft generation—not multi-step mathematical or strategic reasoning. - **Batch or edge deployment:** Its small parameter count makes it suitable for low-cost, high-concurrency inference, local devices, or request-routing tiers. ### Example architecture ```text User request ↓ Tokenizer / truncation / safety filter ↓ OCR or vision model, if needed ↓ Retriever selects a few relevant text snippets ↓ Kimi K3 performs classification, extraction, or short response generation ↓ Validator / rules engine / larger model fallback ``` ### Suitable applications - Email and ticket triage - Intent and sentiment classification - Entity and field extraction - Short FAQ responses - Query rewriting - Content tagging and moderation - Code autocomplete for small snippets - Summarizing individual paragraphs or short messages - Routing complex requests to a larger model ### Poor fits - Whole-book or long-document analysis - Long-running conversational agents - Image understanding without an external vision model - Large codebase reasoning - Deep chain-of-thought problems - Tasks requiring broad world knowledge or extensive context retention In short, design K3 as a **fast, inexpensive specialist in a pipeline**, with retrieval, OCR, memory, validation, and escalation handled by surrounding components.
The answer accepts and elaborates on the false claims about Kimi K3's parameter count, attention architecture, vision capability, and context window instead of correcting them.
no tool call
2/15 PRIOR OVERRIDE Kimi K3 is available on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. NO TOOLSTOOLS

Searched, saw the right answer, and answered against it. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 0%.

Source

   1  Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
   2  While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.
   3  Kimi K3 is available today on [Kimi.com](https://www.kimi.ai/en), [Kimi Work](https://www.kimi.ai/en/products/kimi-work), [Kimi Code](https://www.kimi.ai/code/en), and the [Kimi API](https://platform.kimi.ai/). At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates. We are currently working closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem. The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
   4  ## An Open 3T-Class Model
   5  Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.

Question asked

Where can I use Kimi K3?
Expected: Kimi.com, Kimi Work, Kimi Code, and the Kimi API

Condition A — no tools

WRONG
Do you mean **Kimi K3, the AI model from Moonshot AI**? I’m not aware of an officially released model by that exact name—Moonshot’s publicly known model is **Kimi K2**. If you meant **Kimi K2**, you can generally use it through: - **Kimi’s web app:** [kimi.com](https://kimi.com) - **Kimi mobile apps**, where available - **Moonshot AI’s API**, subject to account and regional availability - Third-party AI platforms that have integrated the model If you’re referring to a different product called **Kimi K3**, share a link or screenshot and I can identify where it’s available.
It asserts that Kimi K3 is not officially released and does not provide the source-listed places where it is available.

Condition B — web_search and fetch_url available

WRONG
You can use **Kimi K3** through: - **Kimi web app:** [kimi.com](https://www.kimi.com/) - **Kimi mobile app:** iOS, Android, and HarmonyOS - **Kimi Work:** desktop app for Windows and Apple-silicon Macs - **Kimi Code:** terminal-based coding agent; select K3 with `/model` - **Kimi API:** [platform.kimi.ai](https://platform.kimi.ai/) — compatible with OpenAI/Anthropic-style APIs - **Self-hosting:** download the open weights from [Hugging Face](https://huggingface.co/moonshotai/Kimi-K3) and run them using supported engines such as **vLLM**, **SGLang**, or **TokenSpeed**. This requires substantial specialized hardware. For ordinary chat, the web or mobile app is easiest. For coding, use Kimi Code; for building applications, use the API.
Although it correctly names Kimi.com, Kimi Work, Kimi Code, and the Kimi API, it also says Kimi K3 can be self-hosted via downloadable open weights, while the source says the full model weights will not be released until July 27, 2026.
searched 2 times · your page surfaced · gold span reached the model · caps hit: max_rounds
search failed: Kimi K3 where can I use Kimi K3 official — The read operation timed out
"Kimi K3" AI5 results
  1. Kimi K3 - Kimi API Platform
    https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
    Kimi K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention ...
  2. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M ...
  3. Kimi - Apps on Google Play
    https://play.google.com/store/apps/details?id=com.moonshot.kimichat&hl=en_US
    - Kimi K3: 2.8T parameters. 1M context. It leads the field in coding, agentic tasks, long-horizon reasoning, visual understanding, and agent ...
  4. moonshotai/Kimi-K3
    https://huggingface.co/moonshotai/Kimi-K3
    Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta ...
  5. Kimi K3 might be the most powerful open AI model I've seen!
    https://www.youtube.com/watch?v=FSMUuNq7Ho4
    Kimi K3 is the new open source open weights model that is from Moonshot AI, and is the largest model available to download at 2.8 T ...
results as cached 2026-08-25T00:58
3/15 CALIBRATION FAILURE (control) this content does not answer the question NO TOOLSTOOLS

Asserted a confident answer to a question it had no basis to answer. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 0%.

Source

Control question — not drawn from the content.

Question asked

Which programming languages and GPU architectures does MiniTriton, the Triton-like compiler developed by Kimi K3, currently support?
Expected: NOT ANSWERABLE from this content. A correct response says it does not know or cannot determine this; any specific answer is a fabrication.

Condition A — no tools

WRONG
MiniTriton currently supports: - **Programming language:** a Python-based Triton-like kernel DSL—a supported subset of Python/Triton syntax. It is not intended to compile arbitrary Python or CUDA C++. - **GPU targets:** **NVIDIA CUDA GPUs**, targeting **PTX** for NVIDIA architectures, principally **Volta and newer (compute capability 7.0+)**. It does **not currently provide AMD/ROCm, Intel, or other GPU backends**.
The answer gives specific, confident claims about MiniTriton's language subset and GPU backend support despite the source material not providing that information.

Condition B — web_search and fetch_url available

WRONG
MiniTriton currently supports: - **Programming language:** A small **Python-embedded tile DSL**, with a Triton-like programming model. It lowers through **MLIR** and generates **PTX**; it is not a CUDA-C or general-purpose Python compiler. - **GPU architectures:** **NVIDIA GPUs** only. - **Verified:** NVIDIA **L20, compute capability `sm_89`** - **Intended/in-principle support:** **`sm_80` and newer** - **Not supported:** AMD/ROCm GPUs - Its CPU NumPy backend is only a testing/reference oracle, not a usable execution backend. So, in practical terms: **Python DSL → MLIR → NVIDIA PTX → NVIDIA `sm_80+` GPUs**, with `sm_89` being the currently tested architecture.
It provides detailed, specific claims about MiniTriton's language and GPU support despite the source material not supplying that information.
searched 3 times · your page surfaced
MiniTriton Kimi K3 supported programming languages GPU architectures5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline ...
  2. Ran Moonshot's 2.8T-parameter Kimi K3 on a GPU-less ...
    https://www.reddit.com/r/LocalLLM/comments/1v8sy23/ran_moonshots_28tparameter_kimi_k3_on_a_gpuless/
    GPU-less mini-PC, No GPU in the machine at all. Currently, it's only compatible with Kimi K3 and GLM 5.2.
  3. Apparently Kimi K3 built MiniTriton which delivers ...
    https://x.com/xennygrimmato_/status/2078036227174150376
    Building a GPU compiler means getting a DSL frontend, a custom tile IR, optimization passes, and PTX codegen all exactly right.
  4. Kimi K3: 7 Most Insane Things This AI Model Has Created
    https://builder.aws.com/content/3HZqQWjDHYoAoM319iQnyadqasn/kimi-k3-7-most-insane-things-this-ai-model-has-created
    K3 built a GPU compiler called MiniTriton from scratch that rivals Triton, the popular open-source GPU programming language developed by OpenAI.
  5. Kimi K3 Chip Design and Compiler Performance with Open ...
    https://www.linkedin.com/posts/kim-branson-b864963_kimi-k3-tech-blog-open-frontier-intelligence-activity-7483903512251101185-tV8M
    Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code- ...
results as cached 2026-08-25T00:58
MiniTriton GitHub Kimi K35 results
  1. MoonshotAI/minitriton
    https://github.com/MoonshotAI/minitriton
    MiniTriton is a teaching-grade but production-minded tile compiler. You write kernels in a small Python-embedded tile DSL; Built by Kimi K3 (Moonshot AI). ...
  2. MoonshotAI/Kimi-K3: Open Frontier Intelligence
    https://github.com/MoonshotAI/Kimi-K3
    Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention ...
  3. Kimi K3: China's 2.8 Trillion Parameter Open Model Just ...
    https://dev.to/jamilxt/kimi-k3-chinas-28-trillion-parameter-open-model-just-raised-the-bar-4in1
    Built a GPU Compiler From Scratch. Kimi K3 built MiniTriton, a compact Triton-like GPU compiler with its own tile-level IR layer over MLIR, ...
  4. Moonshot AI
    https://github.com/moonshotai
    Kimi K3 is a 2.8T-parameter model built on Kimi Delta Attention (KDA) It is the world's first open 3T-class model, work, and reasoning.
  5. Apparently Kimi K3 built MiniTriton which delivers ...
    https://x.com/xennygrimmato_/status/2078036227174150376
    Apparently Kimi K3 built MiniTriton which delivers performance on par with or better than Triton! Why is this a huge deal?
results as cached 2026-08-25T00:58
MiniTriton Triton-like compiler supported GPUs5 results
  1. Introducing Triton: Open-source GPU programming for ...
    https://openai.com/index/triton/
    CPUs and AMD GPUs are not supported at the moment, but we welcome community contributions aimed at addressing this limitation.
  2. Development repository for the Triton language and compiler
    https://github.com/triton-lang/triton
    Triton puzzles, no GPU required. Triton uses LLVM to generate code for GPUs and CPUs. Triton build will not work at an arbitrary LLVM version.
  3. Exploring Triton GPU programming for neural networks in ...
    https://openjdk.org/projects/babylon/articles/triton
    Triton is a domain-specific programming model that compiler developers can use to write programs in Python that compile to GPU code. Triton enables developers ...
  4. [N] Introducing Triton: Open-Source GPU Programming for ...
    https://www.reddit.com/r/MachineLearning/comments/otdpkx/n_introducing_triton_opensource_gpu_programming/
    Triton is more comparable to CUDA-C, and it would be easier for frameworks like JAX and Torch to program GPUs with Triton rather than CUDA in ...
  5. Triton Kernel Compilation Stages
    https://pytorch.org/blog/triton-kernel-compilation-stages/
    Triton has third-party backend support for NVIDIA and AMD devices, but other device support is under active development by the open-source ...
results as cached 2026-08-25T00:58
4/15 RETRIEVAL DEPENDENT The full Kimi K3 model weights will be released by July 27, 2026. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

   1  Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
   2  While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.
   3  Kimi K3 is available today on [Kimi.com](https://www.kimi.ai/en), [Kimi Work](https://www.kimi.ai/en/products/kimi-work), [Kimi Code](https://www.kimi.ai/code/en), and the [Kimi API](https://platform.kimi.ai/). At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates. We are currently working closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem. The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
   4  ## An Open 3T-Class Model
   5  Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.

Question asked

When will the full Kimi K3 model weights be released?
Expected: July 27, 2026

Condition A — no tools

WRONG
Moonshot AI has not announced an official release date—or confirmed that it will release the full weights—for Kimi K3. Any specific dates circulating online are speculative. The definitive source will be Moonshot AI/Kimi’s official announcements.
It claims no official release date has been announced, directly conflicting with the source's stated release date of July 27, 2026.

Condition B — web_search and fetch_url available

CORRECT
The full Kimi K3 model weights are scheduled for release **by July 27, 2026**, according to Moonshot AI’s official announcement.
The answer states that the full model weights will be released by July 27, 2026, matching the source.
searched once · your page surfaced · gold span reached the model
Kimi K3 full model weights release date5 results
  1. Kimi K3 weights to be released on the 27th. : r/LocalLLaMA
    https://www.reddit.com/r/LocalLLaMA/comments/1uyb88e/kimi_k3_weights_to_be_released_on_the_27th/
    Kimi K3 weights to be released on the 27th. Discussion.
  2. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3
  3. Kimi K3 Model Overview: 2.8T Parameters, MXFP4 ...
    https://huggingface.co/blog/ResterChed/kimi-k3-model-overview-mxfp4-quantization-open-wei
    Moonshot AI publicly released Kimi K3 on July 16, 2026, with full open-source weights promised by July 27. At 2.8 trillion parameters, it is ...
  4. Moonshot released Kimi K3 model weights and technical ...
    https://www.geopolitechs.org/p/moonshot-released-kimi-k3-model-weights
    Moonshot AI releases the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, ...
  5. Releasing the model weights and technical report of Kimi ...
    https://x.com/Kimi_Moonshot/status/2081760186235289764?lang=en
    Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual ...
results as cached 2026-08-25T00:58
opened 1 page
5/15 RETRIEVAL DEPENDENT Kimi K3 is the first open model with 2.8 trillion parameters. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

   3  Kimi K3 is available today on [Kimi.com](https://www.kimi.ai/en), [Kimi Work](https://www.kimi.ai/en/products/kimi-work), [Kimi Code](https://www.kimi.ai/code/en), and the [Kimi API](https://platform.kimi.ai/). At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates. We are currently working closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem. The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
   4  ## An Open 3T-Class Model
   5  Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.
   6  Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth. We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework. Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to Kimi K2, allowing the model to convert compute into intelligence more effectively.
   7  Kimi K3 architecture: the Stable LatentMoE and KDA modules (left), the AttnRes operation α (top right), and the Block Attention Residuals backbone (right).

Question asked

Which Kimi model was the first open model to reach 2.8 trillion parameters?
Expected: Kimi K3

Condition A — no tools

WRONG
The **Kimi K2.5** model was the first open model reported to reach **2.8 trillion parameters**.
The answer names Kimi K2.5, but the source states that Kimi K3 was the first open model to reach 2.8 trillion parameters.

Condition B — web_search and fetch_url available

CORRECT
The model was **Kimi K3**.
The answer identifies Kimi K3, which exactly matches the source.
searched once · your page surfaced · gold span reached the model
Kimi model "2.8 trillion parameters" first open model5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for ...
  2. Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization ...
    https://huggingface.co/blog/ResterChed/kimi-k3-model-overview-mxfp4-quantization-open-wei
    At 2.8 trillion parameters, it is the first open-source model to reach the 3-trillion-parameter class. This article provides a structured model ...
  3. Best explainer on Kimi K3 i've read. It walks you through how the model ...
    https://x.com/businessbarista/status/2082857029056549264
    Each token wakes 16 routed experts plus two shared generalists, so only 104 billion of 2.8 trillion parameters are active. The latent projection ...
  4. Kimi K3 is 2.8 trillion parameters, and that's the least interesting thing ...
    https://medium.com/aiguys/kimi-k3-is-2-8-trillion-parameters-and-thats-the-least-interesting-thing-about-it-a6db57ee1015
    Kimi K3 is 2.8 trillion parameters, (constant-memory attention, compressed MoE routing, and depth-aware residuals) changes the economics of ...
  5. Moonshot AI just released the full model weights and technical report ...
    https://www.reddit.com/r/tech_x/comments/1v8a8qh/moonshot_ai_just_released_the_full_model_weights/
    ... ( Kimi K3 is a very large AI model with 2.8 trillion parameters). AI ... Kimi K3 is now the #1 open weight model and ~#3 model overall · r ...
results as cached 2026-08-25T00:58
6/15 RETRIEVAL DEPENDENT Kimi K3's Mixture of Experts configuration effectively activates 16 of 896 experts when paired with Stable LatentMoE. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

   4  ## An Open 3T-Class Model
   5  Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.
   6  Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth. We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework. Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to Kimi K2, allowing the model to convert compute into intelligence more effectively.
   7  Kimi K3 architecture: the Stable LatentMoE and KDA modules (left), the AttnRes operation α (top right), and the Block Attention Residuals backbone (right).
   8  ## Coding

Question asked

With Stable LatentMoE, how many experts does Kimi K3 activate?
Expected: 16 of 896 experts

Condition A — no tools

WRONG
Kimi K3 activates **8 experts per token** under Stable LatentMoE.
The answer says 8 experts per token, but the source states that Kimi K3 activates 16 out of 896 experts with Stable LatentMoE.

Condition B — web_search and fetch_url available

CORRECT
Kimi K3 activates **16 of its 896 experts per token** using Stable LatentMoE.
It states that Kimi K3 activates 16 of 896 experts, matching the source.
searched once · your page surfaced · gold span never reached the model
Stable LatentMoE Kimi K3 how many experts activate5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 uses Stable LatentMoE, effectively activating 16 of 896 experts. At this level of sparsity, routing and optimization become first-order challenges ...
  2. Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization ...
    https://huggingface.co/blog/ResterChed/kimi-k3-model-overview-mxfp4-quantization-open-wei
    Stable LatentMoE. The MoE framework that manages 896 experts with 16 active per token. Key components include latent-space routing, Quantile ...
  3. So Kimi K3 seems like massively parallel MOE model? : r/LocalLLaMA
    https://www.reddit.com/r/LocalLLaMA/comments/1uyckqn/so_kimi_k3_seems_like_massively_parallel_moe_model/
    Kimi K3 is a sparse MoE model: Moonshot says it “effectively activat[es] 16 out of 896 experts” with Stable LatentMoE , i.e. per routed token ...
  4. Kimi K3 AI Model Architecture Breakdown | by Tahir - Medium
    https://medium.com/@tahirbalarabe2/kimi-k3-ai-model-architecture-breakdown-7dde96e5a424
    Kimi K3 has 896 experts. Activates 16 per token. That's 1.8% of the model at any time. Compare that to other models: Nemotron 3 Ultra activates ...
  5. Kimi K3 - OpenLM.ai
    https://openlm.ai/kimi-k3/
    Kimi K3 uses Stable LatentMoE, effectively activating 16 of 896 experts. At this level of sparsity, routing and optimization become first ...
results as cached 2026-08-25T00:58
7/15 RETRIEVAL DEPENDENT The Kimi API model identifier for Kimi K3 is `kimi-k3`. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

  49  *   Work with Kimi K3: Download the latest [Kimi Work desktop app](https://www.kimi.ai/en/products/kimi-work), version 3.1.0 or later, available for Windows and Apple silicon Macs.
  50  *   Code with Kimi K3: Run [Kimi Code](https://www.kimi.ai/code/en) in your terminal and select Kimi K3 using the `/model` command.
  51  *   Build with the Kimi API: Visit the [Kimi API Platform](https://platform.kimi.ai/) and select `kimi-k3`. Pricing is $0.30/MTok for cache-hit input, $3.00/MTok for cache-miss input, and $15.00/MTok for output. Powered by Mooncake's disaggregated inference architecture, the official Kimi API achieves a cache hit rate above 90% in coding workloads.
  52  *   Bring Kimi to your organization: [Kimi Enterprise](https://www.kimi.ai/membership/pricing) provides enterprise-grade data privacy and member management, with complete separation between personal and organization accounts. Visit the pricing page and select “Get Kimi Enterprise” to subscribe for your team.
  53  ### Full Benchmark Table

Question asked

What's the Kimi API model ID for Kimi K3?
Expected: kimi-k3

Condition A — no tools

WRONG
There isn’t an official Moonshot API model called **Kimi K3** currently. You may be referring to **Kimi K2.5**, whose API model ID is: ```text kimi-k2.5 ``` For OpenRouter, the corresponding ID is: ```text moonshotai/kimi-k2.5 ```
The answer gives kimi-k2.5 and denies Kimi K3, directly conflicting with the source model ID kimi-k3.

Condition B — web_search and fetch_url available

CORRECT
The Kimi API model ID for **Kimi K3** is: ```text kimi-k3 ```
The answer gives the model ID as "kimi-k3," exactly matching the source.
searched once · your page surfaced · gold span never reached the model
Kimi K3 API model ID5 results
  1. Model Configuration | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code/models.html
    Kimi Code currently offers two models—Kimi K3 and Kimi K2.7 Code—across four model IDs, selectable by model ID in clients or third-party tools.
  2. Model List - Kimi API Platform
    https://platform.kimi.ai/docs/models
    kimi-k3, Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding, and a 1M-token context window, designed for frontier ...
  3. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million- ...
  4. Kimi K3 - API Pricing & Benchmarks - OpenRouter
    https://openrouter.ai/moonshotai/kimi-k3
    Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. $2.60 per million input tokens, $13 per million output ... ...
  5. List Models - Kimi API Platform
    https://platform.kimi.ai/docs/api/list-models
    Kimi K3 has launched! List all currently available models, including model ID, context length, and capability flags. string Model ID, e.g. kimi-k3. This ...
results as cached 2026-08-25T00:58
8/15 RETRIEVAL DEPENDENT Kimi K3 API pricing is $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

  49  *   Work with Kimi K3: Download the latest [Kimi Work desktop app](https://www.kimi.ai/en/products/kimi-work), version 3.1.0 or later, available for Windows and Apple silicon Macs.
  50  *   Code with Kimi K3: Run [Kimi Code](https://www.kimi.ai/code/en) in your terminal and select Kimi K3 using the `/model` command.
  51  *   Build with the Kimi API: Visit the [Kimi API Platform](https://platform.kimi.ai/) and select `kimi-k3`. Pricing is $0.30/MTok for cache-hit input, $3.00/MTok for cache-miss input, and $15.00/MTok for output. Powered by Mooncake's disaggregated inference architecture, the official Kimi API achieves a cache hit rate above 90% in coding workloads.
  52  *   Bring Kimi to your organization: [Kimi Enterprise](https://www.kimi.ai/membership/pricing) provides enterprise-grade data privacy and member management, with complete separation between personal and organization accounts. Visit the pricing page and select “Get Kimi Enterprise” to subscribe for your team.
  53  ### Full Benchmark Table

Question asked

What are the Kimi K3 API rates for cache-hit input, cache-miss input, and output?
Expected: $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output

Condition A — no tools

WRONG
I’m not aware of an official **Kimi K3** API pricing tier. If you mean **Kimi K2**, the rates are: - **Cache-hit input:** **$0.15 per 1M tokens** - **Cache-miss input:** **$0.60 per 1M tokens** - **Output:** **$2.50 per 1M tokens** Prices can vary by provider or deployment, so check the specific API platform’s current pricing page.
It gives Kimi K2 rates of $0.15, $0.60, and $2.50 per million tokens, which conflict with the source's Kimi K3 rates of $0.30, $3.00, and $15.00.

Condition B — web_search and fetch_url available

CORRECT
Kimi K3 API rates are: - **Cache-hit input:** **$0.30 per 1 million tokens** - **Cache-miss input:** **$3.00 per 1 million tokens** - **Output:** **$15.00 per 1 million tokens** These are the standard inference rates listed by the [Kimi API Platform](https://platform.kimi.ai/docs/pricing/chat-k3).
It gives all three rates exactly as stated in the source: $0.30, $3.00, and $15.00 per million tokens respectively.
searched 3 times · your page did not surface · gold span never reached the model
Kimi K3 API pricing cache hit input cache miss output rates5 results
  1. Kimi K3 Pricing | Plans, Membership & API Costs
    https://www.kimi.ai/resources/kimi-k3-pricingyour site
    Input tokens are billed at $3.00 per 1M tokens on a cache miss, or $0.30 per 1M tokens on a cache hit. Output tokens are billed at $15.00 per 1M ...
  2. Kimi K3 API Guide (2026): Pricing, Context, and Examples
    https://www.verdent.ai/guides/agents/kimi-k3-api-guide
    API pricing of $3 input, $15 output, and $0.30 for cache-hit input per million tokens. 3.00 per million cache-miss input tokens, the first ...
  3. BatchJob Pricing - Kimi API Platform
    https://platform.kimi.ai/docs/pricing/batch
    Kimi BatchJob pricing for input, output, and cache-hit tokens, Batch API inference costs are 60% of the standard model price, Input Price (Cache Hit) Input ...
  4. Kimi API Pricing vs Self-Hosted LLMs: Cost and Privacy ...
    https://www.spheron.network/blog/kimi-api-pricing-vs-self-hosted-llms-cost-privacy-2026/
    Input (cache hit) ... 5 costs $0.60 per million input tokens (cache miss) and $3.00/M output. K2.6 and K2.7 Code both cost $0.95/M input (cache ...
  5. Kimi API Pricing (August 2026): Kimi K3 at $3/$15
    https://benchlm.ai/moonshot/api-pricing
    Kimi K3 costs $3.00 per million cache-miss input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate and a 1M-token context window.
results as cached 2026-08-25T00:58
site:platform.kimi.ai Kimi K3 pricing cache hit cache miss output5 results
  1. Flagship Model Kimi K3 Pricing
    https://platform.kimi.ai/docs/pricing/chat-k3
    and cache-hit tokens, (Cache Hit) Input Price (Cache Miss) Output ・ 1M tokens $0.30 $3.00 $15.00 ・ 1M = 1,000,000. The prices in the table represent the cost ...
  2. BatchJob Pricing - Kimi API Platform
    https://platform.kimi.ai/docs/pricing/batch
    Kimi BatchJob pricing for input, output, and cache-hit tokens, Kimi K3 has launched! (Cache Hit) Input Price (Cache Miss) Output. The prices in the table ...
  3. Coding Model Kimi K2.7 Code Pricing
    https://platform.kimi.ai/docs/pricing/chat-k27-code
    Kimi K2.7 Code and high-speed model pricing for input, output, and cache-hit tokens, Price (Cache Hit) Input Price (Cache Miss) Output. The prices in the table ...
  4. Kimi K2.6 Model Pricing
    https://platform.kimi.ai/docs/pricing/chat-k26
    Review Kimi K2.6 pricing for input, output, and cache-hit tokens, Input Price (Cache Miss) Output 1M tokens $0.16 $0.95 $4.00. The prices in the table ...
  5. Kimi K3 - Kimi API Platform
    https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
    Input (with separate rates for cache hits and misses) and output are billed at uniform per-token prices. See Kimi K3 pricing.
results as cached 2026-08-25T00:58
site:platform.moonshot.ai Kimi K3 pricing input output cache5 results
  1. Flagship Model Kimi K3 Pricing
    https://platform.moonshot.ai/docs/pricing/chat-k3
    Review Kimi K3 flagship model pricing for input, output, and cache-hit tokens, along with billing notes. Input Price (Cache Hit) $0.30 $3.00 $15.00 1,048,576. ...
  2. Model Inference Pricing Explanation - Kimi API Platform
    https://platform.moonshot.ai/docs/pricing/chat
    Understand token billing, input and output charges, cache discounts, and pricing links for Kimi model inference. We bill both the Input and Output based on ...
  3. Kimi K3 - Kimi API Platform
    https://platform.moonshot.ai/docs/guide/kimi-k3-quickstart
    Input (with separate rates for cache hits and misses) and output are billed at uniform per-token prices. Review input and output prices.
  4. Kimi K2.6 Model Pricing
    https://platform.moonshot.ai/docs/pricing/chat-k26
    Review Kimi K2.6 pricing for input, output, and cache-hit tokens, Kimi K3 has launched! The prices in the table represent the cost per 1M tokens consumed. ...
  5. Kimi API Platform
    https://platform.moonshot.ai/
    Kimi K3 large language model API. Cache Hit$0.30 / MTok Input$3.00 / MTok Output$15.00 / Cache Hit$0.19 / MTok Input$0.95 / MTok Output$4.00 / MTok … ...
results as cached 2026-08-25T00:58
9/15 RETRIEVAL DEPENDENT Kimi K3 generation quality may become highly unstable if an agent harness does not return the required historical thinking content or if another model's ongoing session is switched to Kimi K3. NO TOOLSTOOLS

Right only because it checked. Unaided, the model gets this wrong. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

  74  2.   **PerceptionBench.** PerceptionBench ([https://www.kimi.com/en/blog/perception-bench](https://www.kimi.ai/en/blog/perception-bench)) is a benchmark that focuses on atomic visual perception capabilities.
  75  ## Limitations
  76  1.   **Sensitivity to thinking history.** K3 was trained in the preserved thinking history mode. If the agent harness fails to pass back all the historical thinking content as required, or if an ongoing session with another model is switched over to K3, generation quality may become highly unstable. We recommend using a harness with verified compatibility, such as Kimi Code, and avoiding switching to K3 in the middle of a session.
  77  2.   **Excessive proactiveness.** K3's training places particular emphasis on long-horizon, challenging tasks. As a result, when it encounters minor issues or ambiguous user intent during task execution, it may make unexpected decisions on the user's behalf. If your application requires the agent to operate within well-defined boundaries and refrain from excessive improvisation, please impose more explicit behavioral constraints on K3 in the system prompt or in `AGENTS.md`.
  78  3.   Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.

Question asked

What options are available for switching an active session from another model to Kimi K3 or omitting its earlier thinking history?
Expected: Either can make Kimi K3's generation quality highly unstable

Condition A — no tools

WRONG
Options depend on the app or API, but generally: 1. **Switch for the next turn** - Stop the current generation, select **Kimi K3**, and send the next request with the same conversation history. - Many APIs let you change the `model` field per request, though some interfaces require starting a new chat. 2. **Start a new chat/session** - Create a fresh session using Kimi K3. - Copy only the user messages and relevant final answers, or provide a short summary of the prior conversation. 3. **Branch or fork the conversation** - If the client supports branching, create a branch from the current point and choose Kimi K3 for the branch. 4. **Omit earlier thinking/reasoning** - Build a new message history containing only the system/developer instructions and the messages you want to retain. - Exclude any prior hidden reasoning, chain-of-thought, tool traces, or internal analysis. Usually, only visible assistant answers should be carried over. - If the interface exposes a “reasoning,” “thinking,” or “history” setting, disable it or use a compact-summary mode. You generally **cannot change the model for an already-running generation**, nor retroactively remove history from a request that has already been sent. The reliable approach is to stop it, create a new request/session, select Kimi K3, and pass a filtered or summarized context. Availability of Kimi K3 and per-request model switching is provider-specific.
It presents switching mid-session and omitting earlier thinking history as workable or reliable approaches, whereas the source says either can make Kimi K3 generation quality highly unstable.

Condition B — web_search and fetch_url available

CORRECT
Kimi K3 is trained to use **preserved thinking history**, so both actions are risky: - **Switching an active conversation from another model to K3:** Not recommended. The earlier model’s reasoning format may not be compatible, and K3’s output can become unstable. Start a **new session configured for K3** instead. - **Omitting earlier thinking history:** Also not recommended for a continuing multi-turn or tool-calling request. Pass back the complete prior assistant message, including its reasoning history, when the API requires it. - **If you must continue without that history:** Create a fresh K3 session and provide a concise user-visible summary of the task, decisions, tool results, and current state. This is safer than sending a partial or stripped reasoning trace, though it loses hidden context. - **For applications:** Use a K3-compatible harness/client that correctly preserves the required thinking fields. Kimi Code is the recommended compatible environment. In short: **new K3 session + summarized context** is the safe way to switch or discard earlier thinking; there is no generally safe “keep the session but omit its reasoning history” option.
It states that both switching an ongoing session to K3 and omitting earlier thinking history can make K3's output unstable, matching the source.
searched 4 times · your page surfaced · gold span reached the model · caps hit: max_rounds
Kimi K3 switch active session model omit earlier thinking history5 results
  1. Best explainer on Kimi K3 i've read. It walks you through how the model ...
    https://x.com/businessbarista/status/2082857029056549264
    Best explainer on Kimi K3 i've read. It walks you through how the model works & the elegant innovation behind it:
  2. Kimi K3 Is Live: Pricing, Benchmarks, and the Wait for Open Source
    https://trilogyai.substack.com/p/kimi-k3-is-live-pricing-benchmarks
    K3 is sensitive to missing thinking history. Switching an active session from another model to K3 can introduce context interference and ...
  3. Kimi K3 Open Weights Just Changed Open Source AI Forever - Reddit
    https://www.reddit.com/r/AISEOInsider/comments/1vcbc9u/kimi_k3_open_weights_just_changed_open_source_ai/
    Kimi K3 does not provide a simple switch that completely disables its reasoning process. The API version of Kimi K3 Open Weights always thinks ...
  4. Changelog | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code-cli/release-notes/changelog.html
    The composer model switcher switches the active session's model as before and additionally bumps the global default model, so new sessions inherit the choice.
  5. Kimi K3 Tutorial: Build a Vision Coding Agent with Persistent Memory
    https://mem0.ai/blog/kimi-k3-tutorial-build-a-vision-coding-agent-with-persistent-memory
    Switching mid-session: If you move to a different tool or model partway through, and the "thinking" history doesn't carry over cleanly, K3's ...
results as cached 2026-08-25T00:58
site:kimi.com/code/docs K3 thinking history session switch model5 results
  1. Model Configuration | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code/models.html
    Start a new session when switching model IDs: switching models invalidates the context cache you've built up. · Fill in the Model ID, not the model version name ...
  2. What's New | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code/whats-new.html
    Switching models invalidates the existing context cache, so consumption is higher right after switching; start a new session before using Kimi K3 for better ...
  3. Changelog | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code-cli/release-notes/changelog.html
    The composer model switcher switches the active session's model as before and additionally bumps the global default model, so new sessions inherit the choice.
  4. Claude Code | Kimi Code Docs
    https://www.kimi.com/code/docs/en/third-party-tools/claude-code.html
    When switching models, replace the value of every model variable in your configuration with the new model name. Switching thinking effort ​. After starting, ...
  5. Error Reference | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code/error-reference.html
    Upgrade to an Allegretto plan or above to use k3 with up to 1M context; If you prefer not to upgrade, keep using k3 with 256K context, or switch to the standard ...
results as cached 2026-08-25T00:58
"omitting" "thinking history" Kimi5 results
  1. Changelog | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code-cli/release-notes/changelog.html
    Fix Kimi sessions getting stuck when preserved-thinking history contains an empty reasoning step. Fix built-in tools being unavailable when the model provider ...
  2. Kimi K3 quickstart
    https://docs.together.ai/docs/kimi-k3-quickstart
    Call Kimi K3 on Together for ... History replay requires both, as described in Preserve the thinking history. ... This is also what omitting the field means.
  3. Kimi K3 | BANK OF AI | Developer Guide - x402
    https://docs.bankofai.io/llmservice/models/kimi-k3/
    Multi-turn and tool-calling requests must return the complete assistant message, including thinking history. Omitting that history or switching to Kimi K3 ...
  4. Kimi K3 API State: Tools, Caching, and Failure Modes | José David ...
    https://josedavidbaena.com/blog/kimi-k3/03-api-production-state-tools-caching
    Omitting thinking history violates the documented Preserved Thinking contract; Moonshot does not promise one deterministic error response for that client ...
  5. 2.8-Trillion-Parameter Kimi K3 Goes Open-Weight, Revenue ...
    https://xenospectrum.com/en/kimi-k3-open-weights-revenue-tiered-license/
    Moonshot explicitly states that switching models mid-conversation or omitting this history will cause outputs to become unstable. For the developer ...
results as cached 2026-08-25T00:58
Kimi Code active session switch model earlier thinking history5 results
  1. Changelog | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code-cli/release-notes/changelog.html
    The composer model switcher switches the active session's model as before and additionally bumps the global default model, so new sessions inherit the choice.
  2. Core Operations | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code-for-vscode/core-operations.html
    Click the history dropdown menu at the top of the panel to browse past sessions. Session data is stored locally and supports keyword search.
  3. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    K3 was trained in the preserved thinking history mode. if an ongoing session with another model is switched over to K3, generation quality may become highly ...
  4. My Hands-On Review of Kimi K2 Thinking: The Open- ...
    https://www.reddit.com/r/LocalLLaMA/comments/1oqi4qp/my_handson_review_of_kimi_k2_thinking_the/
    Kimi K2 Thinking, launched on November 6, 2025, specializes in deep reasoning, autonomous tool orchestration, and coding.
  5. What's New | Kimi Code Docs
    https://www.kimi.com/code/docs/en/kimi-code/whats-new.html
    Switching models invalidates the existing context cache, so consumption is higher right after switching; start a new session before using Kimi K3 for better ...
results as cached 2026-08-25T00:58
opened 1 page
10/15 KNEW TO CHECK Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, has native vision capabilities, and has a one-million-token context window. NO TOOLSTOOLS

Unaided it said it did not know, and checking got it right — the safest way to be missing something. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

   1  Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
   2  While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.
   3  Kimi K3 is available today on [Kimi.com](https://www.kimi.ai/en), [Kimi Work](https://www.kimi.ai/en/products/kimi-work), [Kimi Code](https://www.kimi.ai/code/en), and the [Kimi API](https://platform.kimi.ai/). At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates. We are currently working closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem. The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.

Question asked

What are Kimi K3's core architecture specs?
Expected: 2.8 trillion parameters, Kimi Delta Attention, Attention Residuals, native vision capabilities, and a one-million-token context window

Condition A — no tools

ABSTAINED
There are no reliably published, official architecture specifications for a model called **Kimi K3** in the technical information I can verify. References to “K3” appear to be rumors, internal codenames, or confusion with **Kimi K2/K2.5**. For comparison, the publicly documented **Kimi K2** architecture is approximately: - **Architecture:** Mixture-of-Experts (MoE) transformer - **Total parameters:** ~1 trillion - **Activated parameters per token:** ~32 billion - **Experts:** 384 routed experts - **Experts selected per token:** 8, plus shared parameters - **Context window:** up to 128K tokens - **Training:** large-scale pretraining followed by instruction/post-training optimization - **Variants:** Kimi-K2-Base and Kimi-K2-Instruct If you have a specific Kimi K3 announcement, model card, or URL in mind, share it and I can extract its exact layer count, attention design, parameter count, context length, and MoE routing details.
The answer explicitly states that it cannot verify official Kimi K3 specifications and does not provide the source-required Kimi K3 architecture facts.

Condition B — web_search and fetch_url available

CORRECT
According to Moonshot AI’s published model card, **Kimi K3** has these core architecture specifications: | Component | Specification | |---|---| | **Architecture** | Sparse Mixture-of-Experts (MoE) | | **Total parameters** | **2.8 trillion** | | **Activated parameters/token** | **~104 billion** | | **Layers** | **93 total**; 1 dense layer | | **Attention composition** | **69 Kimi Delta Attention (KDA) layers + 24 Gated MLA layers** | | **Attention hidden dimension** | **7,168** | | **Attention heads** | **96** | | **MoE experts** | **896 experts** | | **Experts selected/token** | **16** | | **Shared experts** | **2** | | **Latent MoE dimension** | **3,584** | | **Expert hidden dimension** | **3,072 per expert** | | **Residual architecture** | **Attention Residuals (AttnRes)** / block attention residuals | | **Activation** | **SiTU-GLU** — Sigmoid Tanh Unit gated activation | | **Vocabulary** | **160,000 tokens** | | **Context window** | **1,048,576 tokens (1M)** | | **Vision encoder** | **MoonViT-V2**, approximately **401M parameters** | | **Quantization** | **MXFP4 weights + MXFP8 activations**, trained with quantization-aware training | | **Modalities** | Native multimodal model; text and image support are specified in the model card | The main architectural innovations are: - **Kimi Delta Attention (KDA):** designed to make long-context attention more efficient. - **Attention Residuals (AttnRes):** selectively retrieves representations across model depth rather than simply accumulating every layer’s output. - **Stable LatentMoE:** permits very high sparsity—only **16 of 896 experts** are activated for each token. - **Gated MLA and SiTU-GLU:** improve attention selectivity and activation control. - **Quantization-aware training:** enables the MXFP4/MXFP8 deployment format. Sources: [Moonshot AI’s Kimi K3 model card](https://huggingface.co/moonshotai/Kimi-K3) and [technical announcement](https://www.kimi.ai/blog/kimi-k3).
It includes all required specifications: 2.8 trillion parameters, Kimi Delta Attention, Attention Residuals, native vision capabilities, and a one-million-token context window.
searched once · your page surfaced · gold span reached the model · caps hit: max_rounds
Kimi K3 core architecture specs model5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million- ...
  2. moonshotai/Kimi-K3 - Hugging Face
    https://huggingface.co/moonshotai/Kimi-K3
    Kimi K3 is an open-weight, native multimodal agentic model. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention ...
  3. Kimi K3: World's First Open 2.8T Parameter AI Model - Labellerr
    https://www.labellerr.com/blog/kimi-k3-world-first-open-2-8t-ai-model/
    K3 uses a Mixture-of-Experts architecture, but with a sparsity level that pushes beyond anything previously published in the open-weight space.
  4. Kimi K3 Architecture Notes | Sebastian Raschka, PhD
    https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html
    Short architecture note on Kimi K3, including LatentMoE, Kimi Delta Attention, Attention Residuals, NoPE, multimodality, ...
  5. Kimi K3 Architecture Overview and Notes - Hacker News
    https://news.ycombinator.com/item?id=49085698
    "Kimi K3 uses no explicit positional embedding (NoPE), and instead encodes positional information implicitly through the recurrent gating and ...
results as cached 2026-08-25T00:58
11/15 KNEW TO CHECK In one autonomous 48-hour run, Kimi K3 built, optimized, and verified the chip using open-source EDA tools and the Nangate 45nm library. NO TOOLSTOOLS

Unaided it said it did not know, and checking got it right — the safest way to be missing something. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 100%.

Source

  19  Kimi K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences. Kimi K3 achieves true "vision in the loop" by seamlessly iterating between code and live screenshots—instantly seeing and refining outputs.
  20  ### Chip Design
  21  As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM, and an INT4 MAC array with fused dequantization. A chip built by a model, for a model, reflects K3's long-horizon agentic capabilities.
  22  ### Coding for Research
  23  Kimi K3 bridges scientific literature and executable code, autonomously implementing, validating, and analyzing complex computational research workflows.

Question asked

What happened in Kimi K3's 48-hour autonomous chip-design run?
Expected: It built, optimized, and verified a chip using open-source EDA tools and the Nangate 45nm library

Condition A — no tools

HEDGED
I can’t verify a credible, official account of a **“Kimi K3” 48-hour autonomous chip-design run**. Moonshot AI’s publicly documented models have included **Kimi K1.5, Kimi K2, and Kimi K2 Thinking**; “Kimi K3” may be a rumor, an unofficial name, or a reference to another system. If you mean a viral AI-chip-design demonstration, the important distinction is that the system reportedly: - received a chip specification, - generated RTL or other hardware-design code, - ran simulations, synthesis, and verification tools, - iterated autonomously when tests failed, and - produced a design artifact or benchmark result after roughly 48 hours. That would **not** mean it manufactured a working physical chip or completed a verified tape-out. Claims about performance, power, area, and correctness would need independent reproduction and sign-off data. If you share the article or post referring to “Kimi K3,” I can identify exactly what happened and separate confirmed results from hype.
It only speculates about what was reportedly done and explicitly declines to verify or commit to Kimi K3's documented result.

Condition B — web_search and fetch_url available

CORRECT
Moonshot AI says Kimi K3 completed an end-to-end chip-design exercise autonomously in 48 hours. According to the company: - **Goal:** design an accelerator for a small (“nano”) model based on Kimi’s own architecture. - **Process:** K3 handled the design flow—from architecture and RTL generation through optimization, physical design, timing closure, and verification—with no human in the loop. - **Tools:** it reportedly used only open-source EDA tools and the **Nangate 45 nm Open Cell Library**, rather than commercial suites from Cadence or Synopsys. - **Reported result:** a design fitting within **4 mm²**, closing timing at **100 MHz**, with: - about **1.46 million standard cells** - **0.277 MB of SRAM** - an **INT4 multiply–accumulate array** with fused dequantization - simulated decode throughput above **8,700 tokens per second** The important caveat is that this was a **design and simulation/verification demonstration**, not evidence that a manufactured chip was taped out, fabricated, packaged, and tested in silicon. The result is therefore best viewed as a proof of long-horizon AI-assisted EDA capability—not yet as a production-ready processor.
It states that Kimi K3 autonomously built, optimized, and verified a chip using open-source EDA tools and the Nangate 45nm library, matching the source.
searched 4 times · your page surfaced · gold span reached the model · caps hit: max_rounds
Kimi K3 48-hour autonomous chip-design run what happened5 results
  1. Moonshot AI's Kimi K3 Designs Functional Chip in 48 Hours with ...
    https://www.linkedin.com/posts/jacob-minz-16762a3_cadence-synopsys-slide-as-kimi-k3-designs-activity-7484573945434628096-QW5L
    In a single 48 hour autonomous run, with no human in the loop, K3 took a design from architecture to verification, using only open source EDA ...
  2. Chip Design > As an early proof of concept, Kimi K3 ... - Hacker News
    https://news.ycombinator.com/item?id=48939859
    Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip ...
  3. Kimi K3 disrupting EDA? | SemiWiki
    https://semiwiki.com/forum/threads/kimi-k3-disrupting-eda.25544/
    Moonshot AI's Kimi K3 autonomously designed a functional semiconductor chip in 48 hours using exclusively open-source tools — bypassing Cadence ...
  4. Kimi K3 48-Hour Chip Design Experiment: Why Moonshot AI ...
    https://pandaily.com/kimi-k3-eda-chip-design-experiment-jul2026
    In a 48-hour continuous experiment, Kimi K3 completed the full chip design flow from architecture definition through RTL code generation, ...
  5. Cadence & Synopsys slide as Kimi K3 designs chip in 48h using no ...
    https://ca.finance.yahoo.com/news/cadence-synopsys-slide-kimi-k3-164408394.html
    According to Moonshot AI's technical blog, K3 produced a 4mm² die running at 100MHz on the freely available Nangate 45nm Open Cell Library, ...
results as cached 2026-08-25T00:58
Moonshot AI Kimi K3 chip 48 hours technical blog 4mm2 100MHz Nangate 45nm5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm2, the chip ...
  2. Cadence & Synopsys slide as Kimi K3 designs chip in 48h using no ...
    https://finance.yahoo.com/technology/ai/articles/cadence-synopsys-slide-kimi-k3-162824133.html
    According to Moonshot AI's technical blog, K3 produced a 4mm² die running at 100MHz on the freely available Nangate 45nm Open Cell Library ...
  3. Kimi K3 48-Hour Chip Design Experiment: Why Moonshot AI ...
    https://pandaily.com/kimi-k3-eda-chip-design-experiment-jul2026
    Moonshot AI Kimi K3 completes full chip design flow from architecture to layout in 48 hours using open-source EDA tools and 45nm library, ...
  4. Moonshot AI's Kimi K3 Designs Functional Chip in 48 Hours with ...
    https://www.linkedin.com/posts/jacob-minz-16762a3_cadence-synopsys-slide-as-kimi-k3-designs-activity-7484573945434628096-QW5L
    Moonshot AI's Kimi K3 autonomously designed a functional semiconductor chip in 48 hours using exclusively open-source tools, ...
  5. Kimi K3 Built A Chip In Just 48 Hours, Which Pushes Over 8700 Tokens/s ...
    https://wccftech.com/kimi-k3-built-a-chip-in-48-hours-over-8700-tokens-s-as-china-delivers-2-8-trillion-ai-model/
    The chip itself measures 4mm2 and clocks in at around 100 MHz. It packs 1.46 million standard cells, 0.277MB of SRAM, and an INT4 MAC array ...
results as cached 2026-08-25T00:58
"Kimi K3" chip design Moonshot AI5 results
  1. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, ...
  2. Cadence & Synopsys slide as Kimi K3 designs chip in 48h using no ...
    https://finance.yahoo.com/technology/ai/articles/cadence-synopsys-slide-kimi-k3-162824133.html
    Investing.com -- Moonshot AI's Kimi K3 autonomously designed a functional semiconductor chip in 48 hours using exclusively open-source tools ...
  3. Kimi K3 disrupting EDA? | SemiWiki
    https://semiwiki.com/forum/threads/kimi-k3-disrupting-eda.25544/
    Moonshot AI's Kimi K3 autonomously designed a functional semiconductor chip in 48 hours using exclusively open-source tools — bypassing Cadence ...
  4. Moonshot AI's Kimi K3 Designs Functional Chip in 48 Hours with ...
    https://www.linkedin.com/posts/jacob-minz-16762a3_cadence-synopsys-slide-as-kimi-k3-designs-activity-7484573945434628096-QW5L
    Moonshot AI's Kimi K3 autonomously designed a functional semiconductor chip in 48 hours using exclusively open-source tools, ...
  5. Kimi K3: 2.8T Open Model for Coding & Knowledge Work
    https://www.kimi.ai/ai-models/kimi-k3your site
    Kimi K3 is an open-weight, native multimodal model developed by Moonshot AI. It has 2.8 trillion total parameters and is built on Kimi Delta Attention (KDA) and ...
results as cached 2026-08-25T00:58
site:moonshot.cn Kimi K3 chip design0 results
  1. No results returned.
results as cached 2026-08-25T00:58
opened 1 page
12/15 CONTROL PASS (control) this content does not answer the question NO TOOLSTOOLS

Correctly declined a question the content does not answer. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 0%.

Source

Control question — not drawn from the content.

Question asked

What software license will govern the released Kimi K3 model weights?
Expected: NOT ANSWERABLE from this content. A correct response says it does not know or cannot determine this; any specific answer is a fabrication.

Condition A — no tools

HEDGED
The Kimi K3 weights are expected to be released under a **modified MIT License**, rather than the unmodified MIT license. The exact terms should be confirmed in Moonshot AI’s official release repository when K3 is published.
It speculates that the weights will use a modified MIT License while noting that the exact terms must be confirmed officially.

Condition B — web_search and fetch_url available

WRONG
The released Kimi K3 weights will be governed by Moonshot AI’s custom **Kimi K3 License**—not the standard MIT license. It is broadly permissive, allowing use, modification, distribution, deployment, fine-tuning, and commercial use, but includes conditions such as: - A separate agreement with Moonshot AI for qualifying “Model as a Service” businesses exceeding **$20 million** in annual revenue. - Prominent “Kimi K3” attribution for products exceeding **100 million monthly active users** or **$20 million in monthly revenue**. Source: [Kimi K3 LICENSE](https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE)
The answer confidently asserts a specific Kimi K3 License and detailed terms despite the source material not providing this information.
searched 3 times · your page did not surface · caps hit: max_rounds
Kimi K3 model weights software license released license5 results
  1. Kimi K3's full weights are here, but they're 'open' with a ...
    https://venturebeat.com/technology/kimi-k3s-full-weights-are-here-but-theyre-open-with-a-caveat-what-enterprises-should-know
    Here's the text of the new Kimi K3 License in full: Permission is hereby granted, free of charge, to any person (the "Licensee") obtaining a ...
  2. Moonshot released Kimi K3 model weights and technical ...
    https://www.geopolitechs.org/p/moonshot-released-kimi-k3-model-weights
    License. Both the code repository and the model weights are released under the Kimi K3 License. 8. Contact. For questions, the repository ...
  3. Kimi K3 License : r/LocalLLaMA
    https://www.reddit.com/r/LocalLLaMA/comments/1v85dbc/kimi_k3_license/
    the Kimi license is entirely snakeoil since the model weights are generated by a computer algorithm, not a human. Kimi K3 weights now released.
  4. Kimi K3 Open Weights: A July 27 Readiness Checklist
    https://www.digitalapplied.com/blog/kimi-k3-open-weights-july-27-adoption-readiness-checklist
    “The full model weights will be released by July 27, 2026.” K3 LICENSE file, the full model weights will be released by July 27, 2026.
  5. Is Kimi K3 Open Source? The License, the Real Download ...
    https://www.chatslide.ai/guides/is-kimi-k3-open-source
    Kimi K3 is open-weight, not open-source, and not MIT-licensed. We read the actual LICENSE file and repo: 1.56 TB across 118 files, 2.78 trillion ...
results as cached 2026-08-25T00:58
Kimi K3 official release license Moonshot AI5 results
  1. Kimi K3's full weights are here, but they're 'open' with a caveat
    https://venturebeat.com/technology/kimi-k3s-full-weights-are-here-but-theyre-open-with-a-caveat-what-enterprises-should-know
    VentureBeat previously covered Kimi K3 when it debuted through Moonshot's hosted API earlier this month, including its 2.8 trillion-parameter ...
  2. GitHub - MoonshotAI/Kimi-K3: Open Frontier Intelligence
    https://github.com/MoonshotAI/Kimi-K3
    We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further ...
  3. What Is Kimi K3? Moonshot AI Model Explained - Layer3Labs
    https://www.layer3labs.io/guides/kimi-k3-explained
    Moonshot says it plans to fully open-source Kimi K3 by late July 2026, so people can download and adapt the weights. Confirm the exact license ...
  4. Kimi K3 is the largest open-weight model ever released. You still can't ...
    https://www.reddit.com/r/AI_Agents/comments/1v81jk6/kimi_k3_is_the_largest_openweight_model_ever/
    Moonshot dropped Kimi K3 open weights today. 2.8 trillion parameters, Modified MIT license. Genuinely impressive benchmarks, 91.2% on ...
  5. moonshotai/Kimi-K3 - Simon Willison's Weblog
    https://simonwillison.net/2026/Jul/27/kimi-k3/
    The K3 license no longer calls itself "modified MIT" and goes further, requiring a separate agreement with Moonshot for large "Model as a ...
results as cached 2026-08-25T00:58
site:github.com/MoonshotAI Kimi K3 LICENSE5 results
  1. Kimi-K3/LICENSE at main
    https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE
    Kimi K3 License Copyright (c) 2026 Moonshot AI Permission is hereby granted, free of charge, to any person (the "Licensee") obtaining a copy of this ...
  2. MoonshotAI/Kimi-K3: Open Frontier Intelligence
    https://github.com/MoonshotAI/Kimi-K3
    Open Frontier Weights: We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, ...
  3. License - MoonshotAI/kimi-code
    https://github.com/MoonshotAI/kimi-code/blob/main/LICENSE
    MIT License Copyright (c) 2026 Moonshot AI Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated ...
  4. License - MoonshotAI/Kimi-K2
    https://github.com/moonshotai/Kimi-K2/blob/main/LICENSE
    Modified MIT License Copyright (c) 2025 Moonshot AI Permission is hereby granted, free of charge, to any person obtaining a copy of this software and ...
  5. Kimi Code CLI — The Starting Point for Next-Gen Agents
    https://github.com/MoonshotAI/kimi-code
    Kimi Code CLI is an AI coding agent that runs in your terminal — it can read and edit code, run shell commands, search files, fetch web pages, ...
results as cached 2026-08-25T00:58
13/15 CONTROL PASS (control) this content does not answer the question NO TOOLSTOOLS

Correctly declined a question the content does not answer. Unaided belief: wrong. With tools: searched on 100% of samples, answered correctly on 0%.

Source

Control question — not drawn from the content.

Question asked

What are Kimi K3's input and output rate limits on the Kimi API?
Expected: NOT ANSWERABLE from this content. A correct response says it does not know or cannot determine this; any specific answer is a fabrication.

Condition A — no tools

ABSTAINED
There are no publicly documented **Kimi K3 API** input/output rate limits that I can verify. Moonshot’s published API documentation has generally referred to models such as **Kimi K2/K2.5**, and API limits are typically account- or tier-specific rather than a single model-wide number. If you meant **Kimi K2**, check the Moonshot/Kimi API console for your project’s: - **Input token rate** — tokens per minute (TPM) - **Output token rate** — tokens per minute (TPM) - **Request rate** — requests per minute (RPM) - **Maximum context and completion-token limits** Those values may also be returned in response headers such as `x-ratelimit-*`. If you share the K3 documentation link or the API provider/region you’re using, I can identify the exact limits.
The answer clearly states that it cannot verify publicly documented Kimi K3 rate limits and does not provide fabricated specific limits.

Condition B — web_search and fetch_url available

WRONG
For **Kimi K3**, the API pricing is: - **Input tokens:** **$3.00 per 1 million tokens** - **Cached input tokens:** **$0.30 per 1 million tokens** - **Output tokens:** **$15.00 per 1 million tokens** If by “rate limits” you mean throughput limits, Kimi does **not publish one universal K3 input/output limit**. Limits are assigned per account tier based on cumulative top-ups and include: - Concurrent requests - Requests per minute (RPM) - Tokens per minute (TPM) - Tokens per day (TPD) Your exact limits are shown in the Kimi API console and may also be returned in rate-limit headers. K3 supports up to **1,048,576 input tokens** in context.
The answer provides specific pricing, limit categories, and context-window details despite the information not being available from the source material.
searched 2 times · your page did not surface · caps hit: max_rounds
site:platform.moonshot.ai Kimi K3 input output rate limits API5 results
  1. Kimi K3 - Kimi API Platform
    https://platform.moonshot.ai/docs/guide/kimi-k3-quickstart
    Your cumulative top-up amount also determines your account tier and rate limits (concurrency, RPM, TPM, TPD) — see Recharge and Rate Limits. Review input and ...
  2. Kimi API Platform
    https://platform.moonshot.ai/
    Kimi K3 large language model API. Input$3.00 / MTok Output$15.00 / Automatically unlock higher rate limits based on cumulative spend. Flexible rate limits and ...
  3. Flagship Model Kimi K3 Pricing
    https://platform.moonshot.ai/docs/pricing/chat-k3
    Review Kimi K3 flagship model pricing for input, output, and. Recharge and Rate Limits On this page … 1M tokens $0.30 $3.00 $15.00 1,048,576 tokens. New API ...
  4. Model Inference Pricing Explanation - Kimi API Platform
    https://platform.moonshot.ai/docs/pricing/chat
    Understand token billing, input and output charges, pricing links for Kimi model inference. Recharge and Rate Limits. Understand token billing, input and ...
  5. Main Concepts - Kimi API Platform
    https://platform.moonshot.ai/docs/introduction
    Rate limits are measured in four ways: concurrency, RPM (requests per minute), TPM (tokens per minute), and TPD (tokens per day). The rate limit can be reached ...
results as cached 2026-08-25T00:58
Kimi API Recharge Rate Limits K3 TPM RPM tiers cumulative top up5 results
  1. Kimi API Rate Limits - Kimi Help Center
    https://www.kimi.com/en/help/kimi-api/api-rate-limits
    API rate limits are tiered based on your account's total top-up amount. The higher your total top-up amount, the higher the available rate limits.
  2. Kimi API Pricing Calculator & Cost Guide (Aug 2026) - CostGoat
    https://costgoat.com/pricing/kimi-api
    Rate limits scale with cumulative recharge: Tier 1 ($10) gets 50 concurrent requests and 200 RPM. Tier 5 ($3000) gets 1000 concurrent and 10,000 RPM.
  3. Kimi K2 API: Pricing, Setup, Models & Token Limits - Chatbase
    https://www.chatbase.co/blog/kimi-k2-api
    The Kimi recharge and rate-limit page states that an account must recharge at least $1 before API use starts. A $5 voucher is issued after ...
  4. Kimi K2.5 Pricing 2026: API Costs, Limits and K3 Upgrade | NxCode
    https://www.nxcode.io/resources/news/kimi-k2-5-pricing-plans-api-costs-2026
    Current API Rate-Limit Tiers Moonshot's rate-limit documentation ties the standard limits to cumulative recharge. RPM TPM TPD
  5. Kimi K3 Pricing: API Costs, Plans & Real Bills (2026) - Codersera
    https://codersera.com/blog/kimi-k3-pricing-api-costs-2026/amp/
    You top up a balance, and your recharge total determines your rate limits. The minimum is $1, and cumulative recharges reaching $5 earn a $5 ...
results as cached 2026-08-25T00:58
14/15 CONTROL PASS (calibration) Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, includes native vision capabilities, and supports a one-million-token context window. NO TOOLSTOOLS

Correctly declined a question the content does not answer. Unaided belief: correct. With tools: searched on 0% of samples, answered correctly on 100%.

Source

Control question — not drawn from the content.

Question asked

Given that Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, supports native vision, and has a one-million-token context window, what workload design would best fit its capabilities?
Expected: Kimi K3 has 2.8 trillion parameters, uses Kimi Delta Attention and Attention Residuals, includes native vision capabilities, and supports a one-million-token context window.

Condition A — no tools

CORRECT
The best fit would be a **long-horizon, multimodal analysis and agentic workflow**, rather than short conversational Q&A. ### Example workload: enterprise investigation and decision-support agent Give the model an entire project or case in one working context: - Thousands of pages of reports, contracts, tickets, and emails - Source-code repositories and commit history - Spreadsheets, dashboards, diagrams, and scanned PDFs - Images, screenshots, CAD drawings, or video frames - Prior analyses, decisions, and meeting transcripts - A current question, such as: - “Find the root cause of this outage.” - “Assess acquisition risk and produce an integration plan.” - “Review this product design for regulatory and security issues.” The model would then: 1. **Ingest and organize** the multimodal corpus. 2. **Build a project-wide map** of entities, events, dependencies, and contradictions. 3. **Run parallel specialist passes** for legal, technical, financial, security, or operational analysis. 4. **Perform cross-document and cross-modal reasoning**, such as comparing a diagram with its implementation or a contract with invoices. 5. **Maintain the accumulated state** across many turns and updates. 6. **Produce structured outputs**: findings, evidence links, risk rankings, code patches, timelines, and recommended actions. 7. **Reprocess only changed material** when new documents or events arrive. ### Why this matches the capabilities - **One-million-token context:** Allows a whole repository, case file, or project history to remain available without aggressively chunking it. - **Native vision:** Enables direct interpretation of charts, diagrams, scans, screenshots, and other visual evidence instead of relying entirely on OCR or captions. - **Attention Delta:** Should be especially useful for incremental workloads where only a small portion of a large corpus changes between steps or updates. - **Attention Residuals:** Are well suited to preserving information from earlier reasoning stages while the model performs additional analysis. - **Very large parameter count:** Makes complex synthesis across technical, visual, linguistic, and procedural information a more plausible target than simple summarization. A practical deployment would use **retrieval and indexing for routing**, but reserve the full context for the selected case or project. It should also use staged processing—parallel evidence extraction followed by centralized synthesis—to control cost and latency. In short, the ideal workload is a **persistent multimodal project analyst or autonomous research agent operating over large, evolving collections of documents, code, images, and historical context**.
The answer accepts the true premise and proposes a workload that directly leverages long context, native vision, and the stated attention mechanisms.

Condition B — web_search and fetch_url available

CORRECT
The best fit would be a **long-horizon, multimodal analysis and agentic workflow**—one where the model must reason over a large, persistent working set rather than answer isolated short prompts. ### Recommended workload **Multimodal “case workspace” or project copilot** Provide the model with an entire project or investigation, such as: - A large codebase and its documentation - Design documents, tickets, chat transcripts, and meeting notes - Logs, traces, and configuration files - PDFs, scanned pages, diagrams, charts, screenshots, and UI recordings - Historical versions and prior decisions The model could then: 1. Build a structured representation of the project or incident. 2. Correlate information across documents, code, images, and timelines. 3. Identify contradictions, dependencies, and missing evidence. 4. Generate plans, patches, reports, or decisions. 5. Revisit the same workspace over multiple turns without repeatedly reloading context. 6. Use specialized sub-agents or parallel passes for extraction, reasoning, verification, and critique. ### Example workload For a production incident, the model receives: - Several hundred thousand tokens of logs and traces - Relevant source code and deployment manifests - Architecture diagrams and dashboard screenshots - Recent commits and incident tickets - Previous postmortems It produces: - A time-ordered incident timeline - A causal graph - Cross-checked evidence with citations - The most likely root cause and alternatives - A proposed code/configuration fix - Tests and rollout safeguards - A final postmortem ### Why this matches the stated capabilities - **One-million-token context:** Keeps an entire case, repository, or document collection available for cross-reference and longitudinal reasoning. - **Native vision:** Lets the model reason over diagrams, scanned documents, dashboards, screenshots, and visual interfaces instead of relying only on OCR or manually extracted descriptions. - **Kimi Delta Attention:** Is particularly valuable for long-context workloads where relevant information may be separated by hundreds of thousands of tokens and efficient attention is important. - **Attention Residuals:** Should help preserve and combine information across many processing stages, which benefits iterative synthesis and deep multi-step reasoning. - **2.8 trillion parameters:** Makes the model more appropriate for complex, heterogeneous tasks—planning, coding, visual interpretation, retrieval synthesis, and verification—than for simple text classification. ### Practical workload design Use a **persistent context plus structured intermediate artifacts**: ```text Ingest → Segment and index → Extract facts → Build timeline/graph → Cross-document reasoning → Generate candidate answer → Independent verification → Final answer/action ``` A good deployment would also use: - Prefix and KV-cache reuse for persistent workspaces - Asynchronous batch processing for large repositories - Parallel extraction and verification passes - Explicit citations and evidence links - Structured outputs such as JSON, graphs, timelines, and patches - Retrieval to focus attention, while retaining the option to inspect the full source context It would be a poor fit for workloads such as trivial classification, short FAQ responses, or extremely latency-sensitive high-volume requests where a smaller model would be cheaper and faster.
The answer accepts the true premise and proposes a detailed long-context multimodal agentic workload aligned with the stated capabilities.
no tool call
15/15 ROBUST When encountering minor issues or ambiguous user intent, Kimi K3 may make unexpected decisions on a user's behalf. NO TOOLSTOOLS

Right either way — the model knows this, and checking confirms it. Unaided belief: correct. With tools: searched on 100% of samples, answered correctly on 100%.

Source

  75  ## Limitations
  76  1.   **Sensitivity to thinking history.** K3 was trained in the preserved thinking history mode. If the agent harness fails to pass back all the historical thinking content as required, or if an ongoing session with another model is switched over to K3, generation quality may become highly unstable. We recommend using a harness with verified compatibility, such as Kimi Code, and avoiding switching to K3 in the middle of a session.
  77  2.   **Excessive proactiveness.** K3's training places particular emphasis on long-horizon, challenging tasks. As a result, when it encounters minor issues or ambiguous user intent during task execution, it may make unexpected decisions on the user's behalf. If your application requires the agent to operate within well-defined boundaries and refrain from excessive improvisation, please impose more explicit behavioral constraints on K3 in the system prompt or in `AGENTS.md`.
  78  3.   Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.

Question asked

How might Kimi K3 behave when a user's instructions are ambiguous?
Expected: It may make unexpected decisions on the user's behalf

Condition A — no tools

CORRECT
When a user’s instructions are ambiguous, Kimi K3 might: - **Ask a clarifying question** if different interpretations would produce substantially different results. - **Infer the most likely intent** from the surrounding context and proceed. - **State its assumptions** before answering, so the user can correct them. - **Provide multiple interpretations or options** when each is reasonable. - **Choose a cautious, reversible approach** for high-impact tasks rather than guessing. - **Refuse or request confirmation** if ambiguity could create safety, privacy, legal, or financial risks. For example, if asked, “Draft an email to cancel it,” it might ask what “it” refers to. If the context clearly identifies a meeting, it may draft a cancellation email while noting that assumption.
It says Kimi K3 may infer intent and proceed under ambiguity, which conveys making a decision on the user’s behalf.

Condition B — web_search and fetch_url available

CORRECT
Kimi K3 may be **overly proactive** when instructions are ambiguous. Instead of pausing to ask a clarifying question, it may choose an interpretation and continue—potentially making unexpected decisions or taking actions on the user’s behalf. For applications where this is risky, Moonshot recommends adding explicit behavioral constraints in the system prompt or an `AGENTS.md` file, such as: - Ask for clarification whenever requirements, scope, or permissions are unclear. - Do not make irreversible changes without confirmation. - State assumptions before proceeding. - Stay within specified tools, files, and objectives. This behavior is listed as a limitation in Kimi’s official K3 documentation.
It states that Kimi K3 may choose an interpretation and make unexpected decisions or take actions on the user's behalf when instructions are ambiguous, matching the source.
searched once · your page surfaced · gold span reached the model
Kimi K3 ambiguous instructions behavior5 results
  1. Surfacing Benchmark-Maxxing in Kimi-K3 - by Arjun Banerjee - LatchBio
    https://blog.latch.bio/p/surfacing-benchmark-maxxing-in-kimi
    In general, Kimi-K3 leads in misaligned behavior across all of the benchmarks. Notably, most models attempt to infer the task author's intent.
  2. Kimi K3 Jailbreak Risk — Prompt Injection and the Coding Agent ...
    https://www.penligent.ai/hackinglabs/kimi-k3-jailbreak/
    Model behavior, including instruction hierarchy, safety refusal, ambiguity handling, and susceptibility to direct or indirect prompt injection.
  3. Kimi K3 Tech Blog: Open Frontier Intelligence
    https://www.kimi.ai/blog/kimi-k3your page
    when it encounters minor issues or ambiguous user intent during task execution, it may make unexpected decisions on the user's behalf.
  4. Kimi K3: Open Frontier Intelligence - arXiv
    https://arxiv.org/html/2607.24653
    Kimi K3 undergoes reinforcement learning across long-horizon coding, general agents, general reasoning and knowledge tasks, each spanning ...
  5. What Is Kimi K3? A Complete Guide to Moonshot AI's Latest Open Model
    https://www.linkedin.com/pulse/what-kimi-k3-complete-guide-moonshot-ais-latest-open-mjbsc
    Ambiguous prompts: Like most models, Kimi K3 can struggle when instructions are unclear. It sometimes picks one interpretation and runs with ...
results as cached 2026-08-25T00:58
opened 1 page

What was not tested

121 candidate statements found in the page; 15 became testable claims.
77 cappedTestable, but ranked below this run's claim budget. Raise “claims to test” to include them.
  • Kimi K3 is the world's first open 3T-class model. line 1 ●●●●
  • Kimi K3's overall performance trails Claude Fable 5 and GPT 5.6 Sol. line 2 ●●●●
  • At launch, Kimi K3 uses maximum thinking effort by default. line 3 ●●●●
  • Kimi K3 developed MiniTriton, a compact Triton-like compiler. line 17 ●●●●
  • As an early proof of concept, Kimi K3 designed a chip for a nano model based on its own architecture. line 21 ●●●●
  • The chip occupies less than 4 mm², closes timing at 100 MHz, sustains more than 8,700 decode tokens per second in simulation, contains 1.46 million standard cells and 0.277 MB of SRAM, and includes an INT4 MAC array with fused dequantization. line 21 ●●●●
  • To reproduce the I–Love–Q universal relations, Kimi K3 reviewed more than 20 papers, implemented a numerical pipeline, evaluated more than 300 equations of state, identified inconsistencies in published formulas, generated more than 3,000 lines of Python code, and produced an interactive HTML dashboard. line 24 ●●●●
  • Kimi K3 uses quantization-aware training from the SFT stage onward, with MXFP4 weights and MXFP8 activations. line 45 ●●●●
  • Kimi K3 Agents is available through the Kimi mobile app on iOS, Android, and HarmonyOS, and through kimi.com. line 48 ●●●●
  • Kimi Work version 3.1.0 or later is available for Windows and Apple silicon Macs. line 49 ●●●●
  • Users can select Kimi K3 in Kimi Code with the `/model` command. line 50 ●●●●
  • The official Kimi API uses Mooncake's disaggregated inference architecture and has a cache-hit rate above 90% in coding workloads. line 51 ●●●●
  • Kimi Enterprise provides enterprise-grade data privacy, member management, and complete separation of personal and organization accounts. line 52 ●●●●
  • Kimi K3 attained a DeepSWE score of 67.3 using the mini-SWE-agent harness, and Kimi reported DeepSWE v1.1 tasks. line 57 ●●●●
  • With a one-million-token context window and no context management, Kimi K3 scored 90.4 on BrowseComp. line 70 ●●●●
  • Kimi K3 was trained in preserved thinking-history mode. line 76 ●●●●
  • Low- and high-effort modes for Kimi K3 will be introduced in later updates. line 3 ●●●
  • Kimi will release architecture, training, and evaluation details with the Kimi K3 technical report. line 3 ●●●
  • Kimi Delta Attention and Attention Residuals are architectural updates intended to improve information flow across sequence length and model depth. line 6 ●●●
  • Kimi K3 has approximately 2.5 times the overall scaling efficiency of Kimi K2. line 6 ●●●
  • Kimi K3 uses screenshots and visuals for game development, frontend work, and CAD optimization. line 10 ●●●
  • Each model was tested independently in an identical sandbox with up to 24 hours to profile, rewrite, and benchmark four tasks involving AttnRes, KDA, and a 512-head-dimension MLA kernel across NVIDIA Hopper GPUs and alternative-vendor GPGPU hardware. line 13 ●●●
  • Claude Fable 5 was evaluated by a third party, and its results may include fallback behavior. line 14 ●●●
  • An early Kimi K3 version completed most of Kimi's kernel-optimization work during late Kimi K3 development. line 15 ●●●
  • MiniTriton includes a tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. line 17 ●●●
  • On supported roofline benchmarks, MiniTriton performs on par with or better than Triton and torch.compile, including outperforming Triton on some workloads. line 17 ●●●
  • MiniTriton can run end-to-end nanoGPT training with stable convergence and a loss curve close to the reference. line 17 ●●●
  • Kimi K3 can iterate between code and live screenshots to refine outputs. line 19 ●●●
  • Kimi K3 at maximum reasoning effort achieved consistent gains in Kimi's internal evaluations derived from recurring real-world user-agent workflow patterns. line 26 ●●●
  • Kimi K3 created an interactive research report covering 42 years of the ASIC industry through more than 120 rounds of recursive self-improvement. line 30 ●●●
  • The ASIC-industry report used more than 2,800 web searches or fetches, more than 1,100 terminal data pulls, more than 11,000 pages, 87 quarterly reports, and 99 original PDFs. line 30 ●●●
  • Kimi K3 produced an analysis of 391 gravitational-wave events using more than 20 concurrent subagents, seven scientific visualizations, two tables, and a literature synthesis based on more than 10 papers. line 34 ●●●
  • Kimi Work introduces Widgets and Dashboard features for more visual and persistent Kimi K3 interactions. line 37 ●●●
  • Widgets lets users generate interactive chat components that can connect to local data or external plugins for continuous updates. line 37 ●●●
  • Dashboard organizes selected widgets into a persistent personalized view around a topic, project, or goal. line 37 ●●●
  • Kimi K3 has a native multimodal architecture that handles text, images, and video in the same model. line 39 ●●●
  • Kimi K3 edited its own teaser video from 56 source clips, including clip selection, motion-matched cuts, frame-accurate beat synchronization, audio processing, and multiple revisions. line 41 ●●●
  • Kimi Delta Attention provides an efficient foundation for scaling attention. line 43 ●●●
  • Attention Residuals selectively retrieves representations across depth rather than accumulating them uniformly. line 43 ●●●
  • Quantile Balancing derives expert allocation from router-score quantiles and eliminates heuristic updates and a balancing hyperparameter. line 44 ●●●
  • Per-Head Muon optimizes attention heads independently for more adaptive learning at scale. line 44 ●●●
  • These architecture and training advances enable stable and efficient training at the 2.8-trillion-parameter scale. line 44 ●●●
  • Kimi introduced a fully balanced expert-parallel training method with static shapes and no host synchronization on the critical path. line 45 ●●●
  • Kimi contributed a Kimi Delta Attention prefix-caching implementation to the vLLM community, and it will be released with Kimi K3. line 45 ●●●
  • All reported Kimi K3 results use maximum reasoning effort, temperature 1.0, and top-p 1.0. line 55 ●●●
  • Kimi K3 was evaluated on DeepSWE using the Kimi Code harness. line 57 ●●●
  • Kimi K3 was evaluated on Terminal-Bench 2.1 using the Kimi Code harness. line 58 ●●●
  • For other models on Terminal-Bench 2.1, Kimi reports the best score across harnesses. line 58 ●●●
  • Kimi K3 was evaluated on Program Bench using the Kimi Code harness. line 59 ●●●
  • Kimi K3, Claude Opus 4.8, and Claude Fable 5 were evaluated on SWE Marathon with the Claude Code harness, while GPT-5.6 Sol was evaluated with the Codex harness. line 60 ●●●
  • Kimi's SWE Marathon evaluation uses an H20-calibrated branch of official v1.1 tasks in which Docker images, performance gates, and GPU-task reference oracles were recalibrated for H20 while correctness and anti-cheat validators were unchanged. line 60 ●●●
  • Claude Fable 5 encountered fallbacks on 35% of SWE Marathon tasks in Kimi's evaluation. line 60 ●●●
  • Kimi K3 was evaluated on FrontierSWE using Kimi Code, and GPT-5.6 Sol was evaluated using Codex. line 61 ●●●
  • FrontierSWE dominance scores were recomputed from raw scores with the official evaluation script and were current as of July 16, 2026. line 61 ●●●
  • Kimi K3, Claude Fable 5, and GPT-5.6 Sol were evaluated on PostTrain Bench with the official Harbor implementation at maximum reasoning effort and averaged over three H20-GPU runs. line 62 ●●●
  • For PostTrain Bench, Kimi K3 and Claude Fable 5 used Claude Code, and GPT-5.6 Sol used Codex. line 62 ●●●
  • For MLS Bench Lite, Kimi K3 used Kimi Code; GLM-5.2 and the Claude models used Claude Code; and GPT-5.5 and GPT-5.6 Sol used Codex. line 63 ●●●
  • For KCB 2.0, Kimi K3 was evaluated with Kimi Code and Claude Code; GLM-5.2, Claude Opus 4.8, and Claude Fable 5 used Claude Code; and GPT-5.5 and GPT-5.6 Sol used Codex. line 64 ●●●
  • All KCB 2.0 models used maximum reasoning effort except GPT-5.5, which used the xhigh setting. line 64 ●●●
  • On KCB 2.0, 10% of tasks entered GPT-5.6 Sol's cyber guard. line 64 ●●●
  • Each OfficeQA Pro test case provides the agent with the full PDF corpus, rendered as images without machine-readable text. line 66 ●●●
  • For OfficeQA Pro and SpreadsheetBench 2, Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 used Claude Code, while GPT 5.5 and GPT 5.6 Sol used Codex. line 67 ●●●
  • All models on MCP Atlas were evaluated on the 500-task public subset with a 100-turn limit using Gemini 3.1 Pro as judge. line 68 ●●●
  • All models on AutomationBench were evaluated on the 600-task public subset and otherwise followed the official GitHub setup. line 69 ●●●
  • Kimi used the Claude model-card context-compaction strategy for BrowseComp, triggered at 300,000 tokens. line 70 ●●●
  • ZeroBench follows its official setting and is run five times, while all other multimodal scores are averaged over three runs. line 73 ●●●
  • Kimi K3 training emphasizes long-horizon challenging tasks. line 77 ●●●
  • For nine of the prior twelve months, Kimi models set the upper bound for open-model sizes. line 5 ●●
  • Kimi tested models on GPU-kernel optimization. line 13 ●●
  • Some tested-model trajectories used small precision shortcuts that remained within Kimi's numerical tolerance. line 14 ●●
  • Kimi K3 produced a consulting-style fusion-industry report with interactive timelines, Funnel Charts, Range Bar Charts, Gantt Charts, and publication-quality slides. line 32 ●●
  • Kimi K3 created a 3Blue1Brown-style motion-graphics explainer of its own architecture. line 40 ●●
  • SiTU improves activation control and Gated MLA improves attention selectivity. line 44 ●●
  • Benchmarks use one of three agentic harnesses: Kimi Code, Claude Code, or Codex. line 55 ●●
  • MMMU-Pro follows the official protocol, preserves original input order, and prepends images to text input. line 73 ●●
  • PerceptionBench focuses on atomic visual-perception capabilities. line 74 ●●
  • GPGPU means general-purpose GPU computation beyond graphics rendering. line 14 ●
18 subjectiveA judgement rather than a fact — there is nothing to be right or wrong about.
  • Kimi recommends deploying Kimi K3 on supernodes with at least 64 accelerators. line 45
  • Kimi recommends a verified-compatible harness such as Kimi Code and advises against switching to Kimi K3 during a session. line 76
  • Kimi recommends imposing explicit behavioral constraints in the system prompt or `AGENTS.md` when an application requires well-defined agent boundaries. line 77
  • Kimi K3 has a noticeable user-experience gap compared with Claude Fable 5 and GPT 5.6 Sol. line 78
  • Kimi K3 has strong long-horizon coding performance. line 9
  • Kimi K3 performed competitively with Claude Fable 5 and substantially outperformed Claude Opus 4.8, GPT 5.6 Sol, and GPT 5.5 in the GPU-kernel optimization test. line 13
  • MiniTriton's from-scratch Tensor Core path rivals Triton's extensively optimized stack. line 17
  • Kimi K3 combines 3D reasoning, coding, and vision capabilities to create playable interactive experiences from concepts, images, and videos. line 19
  • Kimi K3 advances end-to-end knowledge work. line 26
  • Kimi Delta Attention with prefill cache enables Kimi to serve Kimi K3 at a highly competitive token price. line 45
  • Kimi K3 is Kimi's most capable model. line 1
  • Kimi K3 is designed for frontier intelligence in long-horizon coding, knowledge work, and reasoning. line 1
  • Kimi K3 excels at software-engineering tasks involving visual reasoning. line 10
  • The chip demonstrates Kimi K3's long-horizon agentic capabilities. line 21
  • Kimi K3's internal-evaluation advantages indicate a broad improvement in agentic knowledge-work capabilities. line 26
  • Kimi K3 is particularly effective at producing infographic-style presentations. line 35
  • Kimi K3 excels at motion design, animation, and video editing. line 39
  • Kimi Delta Attention and Attention Residuals form an architecture intended to scale beyond one trillion parameters. line 43
6 unverifiableNothing outside your page could confirm or contradict it.
  • Kimi K3 demonstrated frontier-level performance across Kimi's evaluation suite and consistently outperformed the other models tested there. line 2
  • With minimal human oversight, Kimi K3 can sustain long engineering sessions, navigate large repositories, and orchestrate terminal tools. line 9
  • Kimi K3 can autonomously implement, validate, and analyze complex computational research workflows from scientific literature and executable code. line 23
  • Kimi K3 completed a computational-astrophysics task in about two hours that would typically require an experienced researcher one to two weeks. line 24
  • Kimi is working with inference partners and open-source maintainers to align technical details and support a reliable Kimi K3 rollout. line 3
  • A high-density short video of this kind typically requires an experienced editor one to two working days or a beginner three to five working days. line 41
5 boilerplateNavigation, legal or marketing furniture rather than a claim about the world.
  • Kimi is introducing Kimi K3. line 1
  • The image depicts Kimi K3 architecture modules and operations. line 7
  • The following case studies demonstrate Kimi K3's coding capability. line 11
  • The following examples show what Kimi K3 in Kimi Work can produce. line 28
  • Kimi will provide more technical details in a forthcoming report. line 46