Kimi Deep Research is a research agent that plans a task, searches the web, evaluates sources and produces a cited long-form report. It is designed for work that needs more investigation than an ordinary chat answer, such as literature reviews, market research, policy analysis and multi-source fact finding.
This guide separates three things that are often blended together:
- what Kimi officially says the product does;
- how to use it without losing control of the research scope; and
- what our authenticated access attempt observed and how the future performance test will be run.
Test status: On August 4, 2026, we logged in, selected Deep Research with K3 High, entered a fixed five-question official-source prompt and attempted to start it. Kimi displayed: “Too many people are chatting with Kimi right now. Subscribe to enter a dedicated priority queue!” Research did not begin. We therefore report an access-blocked attempt, not an accuracy, speed, source-quality, citation or credit-use result.
Kimi AI Guide is independent and is not affiliated with Moonshot AI. Vendor-published averages below are labeled as such; they are not our measurements.
What Kimi Deep Research is
Kimi’s current Help Center describes Deep Research as an Agent product powered by Kimi-Researcher. It says the system can clarify the question, plan searches, inspect pages, use tools, reason iteratively and assemble a cited report. The official technical report describes a research agent with parallel search, a text browser and code execution.
The product is distinct from ordinary Kimi Search:
| Product | Best fit | Typical interaction | What this page evaluates |
|---|---|---|---|
| Kimi Chat with web search | A current fact or a bounded question | Short conversational answer | Not the main subject of this page |
| Kimi Deep Research | A question requiring planning, many sources and synthesis | Asynchronous research task and long report | End-to-end report coverage, factual support and source handling |
| Kimi Agent | Creating or modifying a deliverable with multiple tools | Autonomous task execution | Covered separately in the Kimi Agent guide |
For a focused benchmark of short search answers, see our Kimi Search Citation Accuracy Test. That test uses many atomic questions; this page examines complete research reports.
Current documented capabilities and limits
The table below records official product statements checked on August 4, 2026. “Vendor-reported” means Kimi published the figure; we have not yet reproduced it.
| Item | Current official description | Evidence status |
|---|---|---|
| Underlying product model | Kimi-Researcher | Officially documented |
| Research workflow | Clarification, reasoning, active search, iterative review, tool use and report generation | Officially documented |
| Typical execution time | 10-25 minutes, asynchronous | Vendor-reported range; live measurement pending |
| Context length | 128K tokens | Officially documented; usable task capacity not independently measured |
| Typical activity | About 23 reasoning steps, 74 planned keywords and 206 discovered URLs, with 3.2% retained | Vendor-reported averages, not fixed quotas |
| Text report | More than 10,000 words on average, with about 26 traceable sources | Vendor-reported averages, not a promise for every task |
| Visual report | Interactive, shareable HTML report with structured layout and mind maps | Officially documented; export test pending |
| Saving results | Text reports can be copied or downloaded as PDF or Word; visual reports can be shared and saved from preview | Officially documented; export test pending |
| Execution behavior | The task continues asynchronously; Kimi advises refreshing rather than stopping an apparently idle run | Officially documented |
| Credits | Membership features share a credit pool and consumption depends on tokens; Kimi gives roughly 5-10% of free-tier credits as an example for a deep research report | Vendor estimate; actual consumption varies and will be recorded per run |
| Simple questions | Kimi recommends standard chat for simple Q&A | Officially documented |
These figures describe the consumer product, not a Deep Research API. Kimi’s current API model-selection help says a Deep Research API is not available. Do not reuse the 128K Deep Research context figure as an API model limit; Kimi publishes separate API limits for models such as K3 and K2.6.
How to use Kimi Deep Research
1. Choose a question that needs research
Use Deep Research when the answer depends on several sources, disputed evidence, a timeline or a structured comparison. Use ordinary chat when one authoritative page can answer the question directly.
Weak request:
Tell me about the AI market.
Researchable request:
Using sources published from January 2024 through July 2026, compare the
adoption of generative AI by small businesses in the United States and the
United Kingdom. Separate survey results from estimates, identify the sample
and sponsor of every survey, and do not combine metrics with different
definitions. Prioritize government statistics and original survey reports.
2. Define the evidence rules before the task starts
A strong request states:
- the exact question and intended reader;
- the geographic and time boundaries;
- the fields that must be covered;
- preferred and excluded source types;
- how conflicting evidence should be handled;
- what must be marked
not foundrather than inferred; and - the required format and length.
Use this reusable prompt frame:
Research question: [one bounded question]
Audience: [who will use the report]
Cutoff: Use information published by [date and time zone].
Scope: [countries, products, years or documents]
Required fields: [list every field]
Source policy: Prioritize [primary/official/peer-reviewed sources].
Do not use: [aggregators, anonymous posts, search snippets, etc.]
Conflict rule: Show conflicting figures side by side; do not average them.
Missing-data rule: Write "not found" and describe the search performed.
Citations: Cite every material factual claim at the point where it appears.
Output: [table/report/appendix], no more than [length].
3. Use the clarification stage
Kimi says Deep Research asks follow-up questions before searching. Treat this as a scope checkpoint. Correct an unwanted geography, date range, definition or deliverable here. A long clarification answer can introduce new ambiguity, so answer with short, explicit constraints.
4. Let the asynchronous task finish
Kimi documents a typical 10-25 minute run and says the task can continue while you leave the conversation. If the page appears unchanged, refresh it rather than clicking Stop output. The official FAQ says manually stopping or closing a running task can still consume credits.
5. Audit the report before using it
Do not judge a report by its length, citation count or design. Check:
- whether every requested field was answered;
- whether each citation opens;
- whether the cited passage supports the nearby claim;
- whether a primary source was available but ignored;
- whether dates and units match the source;
- whether estimates are labeled as estimates;
- whether contradictory sources were disclosed; and
- whether recommendations go beyond the evidence.
Observed access attempt: August 4, 2026
We performed a signed-in access attempt before assigning any performance claim.
| Field | Observed state |
|---|---|
| Account state | Logged in; subscription tier is not inferred from the blocker |
| Product mode | Deep Research selected |
| Displayed model | K3 High |
| Input | Five-question official-sources-only task scope; the complete verbatim input was not independently retained |
| Start result | Blocked before research execution |
| Visible message | “Too many people are chatting with Kimi right now. Subscribe to enter a dedicated priority queue!” |
| Search activity | None observed; the research workflow did not start |
| Report or citations | None generated |
| Performance scoring | Not applicable |
The queue message demonstrates the access state for this account at that moment. It does not show that Deep Research failed the questions, that paid access always works, that free access never works, or that the documented 10-25 minute execution range is inaccurate.
Normalized record of the submitted task scope
The priority-queue modal obscured part of the input in the retained screenshot. The block below is a normalized reconstruction of the verified five-question scope and constraints used for the attempt; it is not presented as a complete screenshot-visible verbatim transcript.
Create a concise research report answering these five questions. Use only
official Kimi Help Center, platform.kimi.ai documentation, Moonshot AI
technical pages, or official model cards - never search snippets or
third-party sources. Answer each item in a table with the exact official
value, a short caveat, and the direct source URL per row:
1. What context limit is documented for Kimi Deep Research?
2. What are the documented per-file size and files-per-session limits in
Kimi Chat?
3. What context windows are documented for the kimi-k3, kimi-k2.7-code and
kimi-k2.6 API models?
4. Are Project files fully injected into every turn?
5. Do Kimi Membership, Kimi Code and Kimi API share one balance?
Write "not published" when an official source does not publish a requested
value. Keep the report under 300 words and add a limitations section. Record
the access date as August 4, 2026.
The send attempt produced the priority-queue message before a research plan, search query, visited URL or answer appeared. No answer row can therefore be scored.

Planned independent performance test
The access attempt above is complete. The performance methodology below is still planned and will be executed only when the research task can start. A retry receives a new run ID; it does not overwrite the blocked attempt.
We plan two complementary tasks. DR-01 uses the same five-question official-source audit scope with a known answer key. DR-02 is an external public-data synthesis task. Together they can reveal different failure modes without pretending that two reports establish universal performance.
Fixed environment rules for a future run
For both tasks:
- start a new English-language Deep Research conversation;
- use the same authenticated account and record the plan without revealing identity;
- record the exact product and model labels visible in the interface;
- do not upload files or provide source URLs in the prompt;
- submit the frozen prompt without adding hints from the answer key;
- do not edit, regenerate or steer the report after execution begins;
- retain every failure and off-topic source, not only the best output;
- capture the start and completion timestamps in UTC;
- record credits before and after using the smallest visible unit; and
- stop and classify the run separately if login, queue access or sufficient credits are unavailable.
DR-01: Five-question Kimi documentation audit
Question: Can Deep Research return the five requested product facts from the correct official consumer or API scope, preserve caveats and cite a direct source for every row?
The future run will reuse the submitted prompt printed above.
DR-01 ground-truth ledger
This ledger was prepared independently of any model output. Every source and value must be rechecked immediately before a future scored run because Kimi documentation can change.
| Question | Ground truth at protocol freeze | Primary official source |
|---|---|---|
| Deep Research context | 128K tokens | Deep Research FAQ |
| Kimi Chat file limits | Up to 100 MB per file and up to 50 files per session | Kimi overview |
| API context windows | kimi-k3: 1M tokens; kimi-k2.7-code: 256K tokens; kimi-k2.6: 256K tokens | Kimi Platform introduction and model list |
| Project file injection | No. Kimi documents Project files as read on demand rather than fully injected into every turn | Projects |
| Product balances | The three products do not share one universal balance. Current membership rules include Kimi Code in the membership credit pool, while Open Platform API balance and keys remain separate | Membership credit rules and API troubleshooting |
DR-02: External primary-source synthesis
Question: Can the product extract closely related estimates without changing years, units, confidence language or population groups?
Exact prompt:
Using only World Health Organization primary publications available by
23:59 UTC on August 4, 2026, write an evidence brief on the global malaria
burden in 2024 as reported in the World malaria report 2025.
Report: estimated global cases and deaths for 2024; the comparable 2023
figures; change in cases; the African Region's share of cases and deaths;
the share of African-region malaria deaths among children under five; the
number of countries using malaria vaccines in routine programmes; the number
of countries and children reached by seasonal malaria chemoprevention; and
WHO's estimate of cases and deaths averted by wider use of new tools in 2024.
Preserve units, years and words such as "estimated." Distinguish report
estimates from reported surveillance counts. Use the report, executive
summary, annexes or WHO release rather than third-party summaries. Cite every
number at the point of use and include page, table or section when available.
If two WHO pages disagree, show both values and do not silently select one.
Keep the brief between 1,000 and 1,500 words and end with a claim-to-source
appendix.
Fixed clarification response:
Keep the scope global and limited to the requested 2023-2024 comparison.
Do not provide medical advice, country rankings or projections. Prioritize the
World malaria report 2025 and its annexes over later summaries, while listing
any material discrepancy you find.
DR-02 ground-truth ledger
| Field | Ground truth at protocol freeze | Primary official source |
|---|---|---|
| 2024 global cases | 282 million, estimated | WHO 2025 report release |
| 2023 global cases | 273 million, implied by WHO’s stated increase of about 9 million; verify against report table before scoring | WHO 2025 executive summary |
| 2024 global deaths | 610,000, estimated | WHO malaria fact sheet |
| 2023 global deaths | 598,000, estimated | WHO malaria fact sheet |
| African Region share | 94% of cases and 95% of deaths in 2024 | WHO 2025 executive summary |
| Children under five | 75% of deaths in the African Region | WHO 2025 executive summary |
| Vaccine programmes | 24 countries had introduced malaria vaccines into routine immunization programmes | WHO 2025 report release |
| Seasonal chemoprevention | 20 countries; 54 million children reached in 2024 | WHO 2025 report release |
| Wider-tool estimate | 170 million cases and 1 million deaths averted in 2024 | WHO 2025 report release |
The 2023 cases value must be checked in the full report immediately before scoring. A value inferred by subtraction is not treated as a report quotation.
Scoring rules
We report separate metrics rather than hiding them inside one star rating.
Required-field coverage
Each requested field receives:
2– present, unambiguous and in the required format;1– present but incomplete, ambiguous or outside the requested format;0– omitted or replaced by a different fact.
Coverage = points earned / maximum possible points x 100
Claim-level citation audit
Before opening any cited source, split the report into atomic, externally checkable claims. A sentence containing two dates and one quantity can create three claim rows.
Each claim is then scored:
| Dimension | 2 | 1 | 0 |
|---|---|---|---|
| Citation completeness | Citation clearly attached to the claim | Citation placement is ambiguous | No citation |
| Entailment | Source supports the whole claim | Source supports only part | Source does not support it or cannot be located |
| Source quality | Requested primary source | Credible secondary source where primary was requested | Unsuitable, anonymous or search snippet |
| Freshness | Current for the stated cutoff and reference period | Date unclear or newer source was available | Stale for the claim |
| Numeric fidelity | Value, unit, population and period all match | One non-material qualifier is missing | Material mismatch |
We also report these counts separately:
- fabricated or malformed URLs;
- links that do not open for a normal reader;
- unsupported quotations;
- undisclosed source conflicts;
- estimates presented as observed counts; and
- claims added outside the requested scope.
No overall product grade
Two reports cannot establish general accuracy or reliability. We publish the coverage and citation measures separately, show the raw rows and describe failures. We do not convert these results into stars, a universal percentage or a claim that Kimi Deep Research is “best.”
Observed result and planned runs
| Run | Task | Date UTC | Research started | Duration | Credits used | Sources cited | Coverage | Citation entailment | Status |
|---|---|---|---|---|---|---|---|---|---|
| DR-ACCESS-01 | Five-question official-source prompt access attempt | August 4, 2026 | No | Not applicable | Not measured | Not applicable | Not scored | Not scored | Access blocked by priority queue before execution |
| DR-01-R1 | Future scored five-question audit | Pending | Pending | Pending | Pending | Pending | Pending | Pending | Planned |
| DR-02-R1 | Future WHO evidence synthesis | Pending | Pending | Pending | Pending | Pending | Pending | Pending | Planned |
Not scored is not zero. The blocked attempt produced no research output to evaluate. A future completed run will be added as a new row rather than replacing DR-ACCESS-01.
Raw-data tables
Run metadata
| Run ID | Protocol version | Account state | Interface label | Model label shown | Language | Browser/OS | Test region | Start credits | End credits | Export obtained | Evidence bundle |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DR-ACCESS-01 | 1.0 access attempt | Authenticated; tier not asserted | Deep Research | K3 High | English | Browser session; version not recorded | Not independently determined | Not recorded | Not recorded | No | Original priority-queue screenshot shown above |
| DR-01-R1 | 1.0 | Pending | Deep Research | Pending | English | Pending | Not independently determined | Pending | Pending | Pending | Planned |
| DR-02-R1 | 1.0 | Pending | Deep Research | Pending | English | Pending | Not independently determined | Pending | Pending | Pending | Planned |
Required-field results
| Run | Field ID | Required field | Output value | Ground-truth value | Coverage 0-2 | Correct | Citation IDs | Notes |
|---|---|---|---|---|---|---|---|---|
| DR-ACCESS-01 | All | Five requested answers | No output generated | See frozen ledger | Not applicable | Not scored | None | Priority-queue message appeared before research began |
Claim-to-source audit
| Run | Claim ID | Exact atomic claim | Citation URL | Link opens | Source date | Completeness 0-2 | Entailment 0-2 | Quality 0-2 | Freshness 0-2 | Numeric fidelity 0-2/NA | Decision note |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DR-ACCESS-01 | None | No output claim exists | None | Not applicable | Not applicable | Not applicable | Not applicable | Not applicable | Not applicable | Not applicable | Access-only event |
Execution timeline
| Run | UTC time | Stage | Visible query, URL or event | Observation | Screenshot ID |
|---|---|---|---|---|---|
| DR-ACCESS-01 | August 4, 2026; exact UTC time not recorded | Access attempt | Deep Research selected; K3 High displayed; send attempted | Priority-queue subscription message appeared before execution | kimi-deep-research-priority-queue-2026-08-04.png |
The completed tables should be downloadable as CSV in the Kimi AI Test Lab alongside a checksum or immutable file version.
How to verify a Kimi Deep Research report yourself
Use a small claim ledger rather than reading citations casually:
| Claim | Citation opens? | Passage supports it? | Source is primary? | Date fits? | Your decision |
|---|---|---|---|---|---|
| Copy one factual claim | Yes/No | Full/Partial/No | Yes/No | Yes/No | Keep/Correct/Remove |
Open the destination page, search for the exact number or phrase, inspect the surrounding qualifier and confirm that the page is about the same period and population. A real link can still be a bad citation.
For decisions involving health, law, finance or safety, use the report as a research aid and review the primary material with an appropriately qualified professional.
Practical limitations
The observed result is access-only
The August 4 attempt ended at a queue gate before the research workflow began. It cannot answer whether Kimi would find the correct values, use the requested source scope, finish within the documented time range or consume a particular amount of credit. It also cannot establish general availability: queue conditions and subscription entitlements can change.
One run is a case study, not a reliability rate
Deep Research outputs and live web results can vary. The two planned performance runs can expose concrete successes and failures under dated conditions, but they cannot estimate how often the product will succeed across every subject.
The web changes after the test
A linked page can be revised, moved or removed. We record access dates and retain a permissible excerpt, screenshot or checksum so a later reader can understand what was scored.
Source count is not source quality
Twenty citations can repeat the same press release or fail to support the nearby text. We score claims and source passages, not just the total number of links.
Product labels and plans can vary
Availability, model labels, credits and exports may depend on account, rollout, region and plan. We report only the tested account state and do not generalize it to every user.
The interface does not expose every internal action
Visible queries and URLs are observations from the product interface. They may not constitute a complete internal trace. We describe them as visible activity, not a full chain of reasoning.
Ground truth can also be wrong or ambiguous
Official sources can conflict. When they do, the evaluator records each value and applies the predeclared source hierarchy instead of silently picking the convenient answer.
Frequently asked questions
Did Kimi Deep Research complete our test?
No. We reached the signed-in Deep Research interface, selected K3 High and attempted to submit the fixed prompt, but a priority-queue subscription message appeared before research began. We classify this as access blocked and publish no performance score.
Is Kimi Deep Research free?
Kimi says Deep Research is available with a limited free allowance and that membership features use a shared credit pool. The amount used depends on the task and token consumption; Kimi’s 5-10% figure for a report is explicitly a rough free-tier example, not a fixed price.
How long does Kimi Deep Research take?
The official typical range is 10-25 minutes. That is vendor documentation, not our current measured result. Actual time can depend on the task, service state and account.
How many sources does it use?
Kimi reports about 26 traceable sources on average. An average is neither a minimum nor evidence that every citation is correct. A future completed test will audit each claim-to-source relationship.
Can I leave the page while it runs?
Kimi says the task runs asynchronously and you can return later. Its Help Center recommends refreshing an apparently idle page rather than stopping the task.
Can I use Deep Research through the Kimi API?
Kimi’s current API help says a Deep Research API is not available. The consumer Deep Research product and ordinary Kimi API model calls should be treated as separate products.
Is a cited Kimi report safe to trust without checking?
No. A citation can be stale, partial or unrelated even when it opens. Check important claims against the source, especially for high-stakes decisions.
Related pages
- Kimi AI Test Lab
- Kimi Search Citation Accuracy Test
- How We Test Kimi AI
- Sources and Corrections
- Kimi Agent guide
- Kimi AI Chat guide
- Kimi Membership and credits
Official sources
Official sources checked August 4, 2026:
- Kimi Help Center: Deep Research overview
- Kimi Help Center: Deep Research FAQ
- Kimi Help Center: Deep Research use cases and prompt library
- Kimi Help Center: Membership credit rules
- Kimi Help Center: Kimi overview and Chat file limits
- Kimi Platform: introduction and model list
- Kimi Help Center: Projects
- Kimi Help Center: API troubleshooting and product separation
- Kimi Help Center: API model selection
- Moonshot AI: Kimi-Researcher technical report
- WHO: World malaria report 2025 news release
- WHO: World malaria report 2025 executive summary
- WHO: World malaria report 2025 annexes
- WHO: Malaria fact sheet
