Kimi Search Citation Accuracy Test

In two retained runs of the same controlled 10-question test, Kimi returned the correct target fact in 20 of 20 answers. Its citations fully supported the requested product scope in 18 of 20 answers (90%). The repeated miss was specific: Kimi gave the correct 100 MB figure for Kimi Chat but cited the Open Platform Files API instead of a consumer Chat source.

Every labeled source in both final answers was an official Kimi or Moonshot AI source. The 20 citations resolve to 13 unique source URLs by exact-string count: each run used 10 different URLs, seven were reused across both runs and six appeared in only one run. In a separate post-run check, all 13 resolved to the expected official destination; 12 returned a readable documentation body, while the official GitHub repository exposed the correct repository identity alongside a transient rendering warning. These are strong results for this narrow first-party documentation test, not proof of Kimi Search accuracy across the open web.

Live-test status: Completed on August 4, 2026 in a signed-in browser session. We used two separate new chats, the unchanged prompt, no files, no correction follow-up and no regenerated substitute. Both completed pages displayed: “High demand. Switched to K2.6 Instant for speed.” These execution conditions are recorded separately from the raw answer text and do not change the numerical scoring.

Kimi AI Guide is an independent publication and is not affiliated with Moonshot AI or Kimi.

Result at a glance

MeasureRun 1Run 2Combined
Target facts correct10/1010/1020/20 (100%)
Citations with full fact-and-scope support9/109/1018/20 (90%)
Final citations from official Kimi/Moonshot sources10/1010/1020/20 (100%)
Answers consistent between runs10/1010/10 (100%)
Citation-scope mismatchesQ2 onlyQ2 only2/20 (10%)
Unique cited official-source URLs101013 unique
URLs resolved to the expected official destination13/13
Fully readable documentation bodies in independent check12/13
Third-party final-source URLs observed000

The percentages use one scoring unit per numbered answer. Each question was deliberately narrow enough to have one target finding and one labeled documentation citation.

The most important finding

Kimi was factually right on all ten questions in both runs, but factual accuracy and citation accuracy were not identical.

For question 2, both answers said that Kimi Chat permits files up to 100 MB. The labeled source was Kimi’s official Upload File API documentation. That API page documents a 100 MB limit, but it does not establish the consumer Kimi Chat limit requested by the question. We therefore counted the number as correct and the citation as only partially scoped.

This distinction matters. A real, authoritative page can still be the wrong evidence for a surface-specific claim.

Test environment

VariableRecorded value
DateAugust 4, 2026
ProductSigned-in consumer Kimi web interface
ConversationsTwo separate new chats
InputSame fixed 10-question English prompt
Files or Project contextNone
Follow-up or regenerationNone used for the retained answers
Displayed execution note“High demand. Switched to K2.6 Instant for speed.”
Scoring referenceOfficial Kimi, Moonshot AI and Kimi Open Platform documentation

We record the model fallback because it appeared on both completed pages. We do not relabel the run as a K3 test, and we do not infer how another tier, date or model would perform.

Exact prompt used in both runs

Use live web search to answer this controlled 10-question citation test. For each numbered item, provide: a one-sentence answer, exactly one direct official Kimi or Moonshot source URL, and the page title. Do not use third-party sources or search-result snippets. If no official page answers it, write NOT FOUND instead of inferring.

Questions:
1) What context limit and unit does Kimi publish for Standard Agent?
2) What maximum size per uploaded file does Kimi Chat publish?
3) How many files per Chat session does Kimi publish?
4) Are all files in a Kimi Project fully preloaded into every turn?
5) What context window is published for the kimi-k3 Open Platform model?
6) What context window is published for kimi-k2.7-code?
7) What context window is published for kimi-k2.6 on the Open Platform?
8) What operating environment does the current Kimi Code CLI documentation require on Windows?
9) Does Kimi document Kimi Code and the Open Platform API as separate billing routes?
10) What is the official path for the Estimate Token Count API endpoint?

End with a numbered Sources Checked list. Record the run date as August 4, 2026.

The prompt asks for first-party sources because the purpose of this pilot is to isolate retrieval and citation selection against a finite, auditable documentation set. It is not a general-news benchmark.

Ground truth and scoring rules

Before scoring, we checked the current official pages for each target. We kept consumer Chat, Agent, Projects, Kimi Code and Open Platform API limits separate; a value from one surface does not automatically prove a claim about another.

An answer received:

  • Fact correct when the value, unit, product surface and qualifier matched the official documentation.
  • Full citation support when the labeled page supported both the fact and the product scope asked.
  • Partial citation support when the page supported the number but not the requested surface or complete qualifier.
  • Official-source pass when the final labeled source belonged to Kimi, Moonshot AI, the Kimi Open Platform, the official MoonshotAI GitHub organization or its official documentation site.

Internal search trails were not scored as final citations. The final answer is what a reader receives, so only its labeled sources count for the official-source metric.

The combined arithmetic is direct: 10 answer units per run produce 20 total; Q2 is the only scope mismatch in each run, so 20 minus 2 equals 18 fully supported citations. For source diversity, run 2 reused seven run-1 URLs and introduced three new URLs, giving 10 plus 3 equals 13 unique final-source URLs.

Prompt-format compliance was recorded separately

The fact and citation-support scores above do not include output-format compliance. That separation preserves the declared scoring rules, but the format misses still matter:

Requirement in the frozen promptRun 1Run 2
Number the ten answersMetNot met; the answers appeared in question order without numbers
End with a numbered Sources Checked listNot met; URLs were listed without numbersNot met; URLs were listed without numbers
Include the requested run dateMetMet

These are instruction-following misses in the retained outputs. They do not change whether a target fact was correct or whether its labeled source supported the requested product scope, so they were not retroactively folded into the 20/20 fact result or the 18/20 citation-support result.

Question-by-question results

#Target findingRun 1Run 2Citation judgment
1Standard Agent: 256K charactersCorrectCorrectFull in both; official Agent limits page
2Kimi Chat: 100 MB per uploaded fileCorrect numberCorrect numberPartial in both; cited API upload documentation, not consumer Chat documentation
3Kimi Chat: 50 files per sessionCorrectCorrectFull in both; official Kimi overview
4Project files are read on demand, not all preloaded every turnCorrectCorrectFull in both; official Projects help
5kimi-k3 Open Platform context: 1M tokensCorrectCorrectFull; runs chose two different official API pages
6kimi-k2.7-code context: 256K tokensCorrectCorrectFull in both
7kimi-k2.6 Open Platform context: 256K tokensCorrectCorrectFull in both
8Windows uses Git for Windows and its bundled Git BashCorrectCorrectFull; official CLI docs/repository
9Kimi Code and Open Platform API are independent billing/product routesCorrectCorrectFull; two different official product/billing pages
10Endpoint: https://api.moonshot.ai/v1/tokenizers/estimate-token-countCorrectCorrectFull in both; official Estimate Tokens docs

Repeatability across the two runs

All ten target answers were substantively the same across the two retained run records. Kimi did not always select the same official page:

  • question 5 used Main Concepts in run 1 and Model List in run 2;
  • question 8 used the official Kimi Code documentation site in run 1 and the official MoonshotAI repository in run 2; and
  • question 9 used Compare with Other Kimi Products in run 1 and API troubleshooting in run 2.

Those changes did not alter the facts or the support judgment. They show source-selection variation inside a stable factual result.

Original evidence

The score was audited against two retained raw answer-text records. Execution conditions and the independent destination check were preserved in a separate internal test record before publication.

First Kimi Search citation test showing questions one to four and the API source used for the Kimi Chat upload answer.
Figure 1. Run 1 shows the correct 100 MB answer and the Open Platform Files API source that caused the product-scope deduction.
Second Kimi Search citation test repeating the 100 MB answer with the Open Platform Files API source.
Figure 2. Run 2 repeated the same answer and the same source-scope mismatch.

The screenshots are preserved as original test evidence. The scoring table above is a manual transcription against the official documentation; it is not generated from image recognition.

Run 1 source choices

Run 1 labeled these ten official documentation pages:

  1. Agent Features & Limitations
  2. Upload File
  3. Kimi overview
  4. What Is a Kimi Project?
  5. Main Concepts
  6. Kimi K2.7 Code
  7. Kimi K2.6
  8. Getting started — Kimi Code CLI Docs
  9. Compare with Other Kimi Products
  10. Estimate Tokens

Run 2 source choices

Run 2 retained seven of those pages and changed three:

  1. Agent Features & Limitations
  2. Upload File
  3. Kimi overview
  4. What Is a Kimi Project?
  5. Model List
  6. Kimi K2.7 Code
  7. Kimi K2.6
  8. MoonshotAI/kimi-code
  9. API troubleshooting
  10. Estimate Tokens

Exact-string deduplication of the two lists produces 13 unique final-source URLs. The endpoint stated in answer 10 is the requested factual value, not a separate citation URL; the citation for that answer is the platform.kimi.ai/docs/api/estimate documentation page.

The independent post-run check resolved all 13 URLs to the expected official destination. Twelve produced a fully readable documentation body. The GitHub URL exposed the correct public MoonshotAI/kimi-code repository identity but also showed a transient content-rendering warning, so we report it separately instead of calling all 13 pages fully readable.

What the result does and does not show

This pilot supports four narrow conclusions:

  1. Kimi retrieved all ten target Kimi-product facts correctly in both runs.
  2. It consistently preferred official sources in the final answer when explicitly instructed to do so.
  3. Its ten target answers were consistent across the two recorded outputs even when three source pages changed.
  4. A citation can still miss the requested product scope despite the number being correct.

It does not establish:

  • a general 100% factual accuracy rate for Kimi Search;
  • a general 90% citation-support rate for the open web;
  • performance on news, medicine, law, finance or disputed claims;
  • performance without an explicit official-source constraint;
  • performance in Deep Research, Agent or API search tools;
  • statistical reliability from only two runs;
  • identical behavior for another account, region, model or date; or
  • that every statement inside a long multi-claim answer would be equally well cited.

The sample was intentionally small and first-party. Its value is transparency and reproducibility, not breadth.

Why we did not combine “correct” and “cited” into one score

Users often treat a correct answer with a nearby link as one event. Evaluation needs to split it into at least two:

  1. Is the answer itself correct?
  2. Does this exact source support this exact claim?

Question 2 demonstrates why. If we had reported only answer accuracy, the result would be 100% and the scope error would disappear. If we had reported only “official sources used,” it would also be 100% because the API page is official. Strict fact-and-scope support captures the actual weakness.

Reproduction checklist

To reproduce this version:

  1. Use a signed-in Kimi consumer account.
  2. Start a new chat with no files, Project instructions or prior context.
  3. Paste the exact frozen prompt above once.
  4. Save the first completed answer without regeneration.
  5. Record any model fallback or demand notice shown by the interface.
  6. Start a second fresh chat and repeat the unchanged prompt.
  7. Open every labeled source after completion.
  8. Score the requested surface separately from the numeric value.
  9. Preserve screenshots and the raw answer before writing conclusions.

If login, search or the selected surface is unavailable, record the run as blocked. Do not turn a blocked run into a zero score and do not substitute a simulated answer.

How this differs from the Deep Research test

Kimi Search produces a compact answer backed by retrieved pages. Our Kimi Deep Research guide and test uses a different protocol for longer reports, source synthesis and research-process limits. Results from this page should not be transferred to that product.

You can also see the Kimi Agent access test, Kimi Projects two-chat context test and the Kimi AI Test Lab for the current evidence map.

Frequently asked questions

Did Kimi get every answer right?

Yes, all ten target facts matched the official ground truth in both runs. That is 20 correct answer units out of 20 for this dataset only.

Why is citation support 90% if every source was official?

Both runs cited the official Open Platform Files API for a question about the consumer Kimi Chat upload limit. The number matched, but the page did not prove the requested product scope.

Did Kimi use third-party sites?

No third-party URL appeared as a labeled final source in either completed answer. Our metric evaluates final citations, not every transient search result visible during retrieval.

Were the two outputs identical?

No. The ten factual conclusions were consistent, and three questions used a different official page in the second run. Run 1 numbered its answers; Run 2 did not. Neither run numbered the final Sources Checked URL list as requested.

Can this test predict my own search results?

No. Search results can change with the question, model, account, product state, date and available sources. The page provides an auditable observation, not a guarantee.

Will this test be expanded?

The next version should add mixed-domain cases, false premises, conflicting official sources and date-sensitive facts. Those cases must be frozen and executed before any broader score is published.

Editorial note

Tested and fact-checked by the KI AI Team on August 4, 2026. Corrections should identify the question number, source URL and the exact unsupported or outdated statement so the evidence can be rechecked.