InsightsPublished 5 min read

Kimi K3: working with more documents and images at once

Kimi K3 can take in more material in one request. Learn what that helps with, what the tests show and what to check before running it on your own systems.

By Monolith

A wall of documents and diagrams with one highlighted panel
A wall of documents and diagrams with one panel highlighted.

Moonshot introduced Kimi K3 in July 2026. It can work with images and a large amount of text in one request. That may be useful for a long research brief or a software project spread across many files. The key question is whether it finds and uses the right details. Teams that want to run it themselves also need to check its custom license.3

When more room for documents helps

Imagine taking over a client account with a long brief, meeting notes and several versions of a proposal. You want to know which promises are still current. If only one document fits into the AI’s working space, you may have to split the task and manually connect the answers.

A larger context window gives you more room to include related material together. You could ask the model to identify conflicting deadlines and quote the passages behind its answer. That is a useful trial for a long-context model: does the extra material help it find the relationship you need?

More room also makes careful instructions more valuable. Tell it which document takes priority, ask it to flag disagreements and keep the original files available for checking. A large collection of outdated drafts can create confusion for AI just as it can for a new colleague.

What changed from the K2 generation

Moonshot says K3 is designed to help with lengthy coding and research tasks. Its internal design uses selected parts of the model for each step, rather than activating everything at once. That approach is called mixture of experts. The word “experts” describes parts of the software, not people or a guarantee of specialist judgment.4

DetailWhat it means
Size2.8 trillion parameters in total; 104 billion active.
Named design changesKimi Delta Attention and Attention Residuals. These describe internal processing, not business outcomes.
Architecture reference from Moonshot.4

A context window is the amount of information a model can work with in one request. K3's listed window is 1,048,576 tokens, compared with 262,144 for K2.7 Code. Tokens are small pieces of information rather than pages, so document capacity varies. More room can help you include connected material, but does not guarantee that the model notices every important fact.2

Benchmark comparisons, not a universal ranking

MeasureKimi K3GPT-5.5GLM-5.2
GPQA Diamond93.593.591.2
Terminal-Bench 2.188.383.482.7
Moonshot's K3 repository table. K3 and GLM use max effort; GPT-5.5 uses xhigh. These are publisher-reported comparisons.4

GPQA Diamond tests difficult science questions. Terminal-Bench tests work done through a computer's command interface. The terminal scores here used different supporting software: Kimi Code, Codex and Claude Code. Moonshot selected competitors' best published results across these setups. Use the table to identify models worth trying, rather than as a clean test of the models alone.4

To compare models, give each the same files and instructions and allow the same tools and time. Decide what a successful answer must contain before looking at the results. Include the attempts that fail, since those still use time and money.

Practical uses with clear boundaries

For a software handover, ask K3 to explain how one feature works and point to the relevant files. For a website review, give it screenshots and a brief, then ask it to identify specific differences. Check the explanation before permitting changes. Image understanding can help with a review, but it cannot replace testing with users or checking accessibility.4

Long tasks need the software to retain the earlier steps. If a system forgets which files it already read or what a tool returned, it may repeat work or miss information. This is a connection detail for your developer to check before relying on a long-running task.

DetailWhat it means
Conversation stateKeep tool calls and required reasoning information between steps.
Thinking modeAlways enabled. Low, high and max effort are available; max is the default.
Developer notes from the K3 guide.1

Access, cost and the license distinction

Using Moonshot’s hosted service means sending requests to its servers and paying usage fees. Eligible reference material stored for reuse can be cheaper to read again. Ask your developer whether your task actually qualifies for that discount rather than assuming every repeated document will cost less.

DetailWhat it means
Rates per million tokensUncached input $3.00; cached input $0.30; output $15.00, excluding applicable taxes.
Starting accessThe quickstart lists a successful top-up of at least $1.
Usage limitsAccount tiers set request-rate limits.
Hosted service details from the cited documentation.12

The downloadable model uses the custom Kimi K3 License, rather than the standard MIT license. A team planning to run it or build a service around it should review those terms for its intended use. Public availability does not mean every form of distribution has the same conditions.3

The license includes conditions for some businesses selling model access and attribution requirements at specified scale, with stated exceptions. The source review verified that the files were available; it did not establish the exact day they first appeared.3

Start with an explanation that someone on your team can check. For example, ask which documents support a conclusion and where they disagree. If that works well, try a small edit or a more involved research question. This lets you see whether the extra document capacity is useful before rebuilding a process around it.

Terms used in this article

TermPlain-language meaning
TokenA small unit of information counted by the model. Usage fees often depend on how many are read and generated.
Input / outputThe information you send / the response the model generates.
Cached inputEligible information stored for reuse, sometimes charged at a lower rate.
BenchmarkA defined test. Its score depends on the tasks, settings and supporting tools.
Context windowHow much information fits into one request; it is not a guarantee of perfect recall.
A quick guide to terms used in this article.
Questions, answered
How much context does Kimi K3 support?
Moonshot lists 1,048,576 tokens for K3, compared with 262,144 for K2.7 Code.2
Is Kimi K3 under the MIT license?
No. It uses the custom Kimi K3 License, which includes conditions for certain model-as-a-service uses and attribution at specified scale.3
What does the K3 API cost?
The listed rates are $3.00 uncached input, $0.30 cached input and $15.00 output per million tokens, excluding applicable taxes.2
Sources

Read for this feature. The numbers match the markers in the text.

  1. Kimi K3 API quickstartplatform.kimi.ai
  2. Kimi API pricingplatform.kimi.ai
  3. Kimi K3 Licenseraw.githubusercontent.com
  4. Moonshot Kimi K3 model cardraw.githubusercontent.com
Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.