The Kimi AI Test Lab is where we publish reproducible tests of Kimi products, models and tools. Every completed report should let a reader see the exact question, dated environment, scoring rule, raw observations, failures and limits behind the conclusion.
This is not a leaderboard or a collection of polished demos. A successful example can show that one task worked once. It cannot prove that the product is reliable for every user, topic or date. We preserve that distinction throughout the lab.
Batch 1 baseline, recorded August 4, 2026: One controlled Kimi Projects test completed across two chats, and two fresh-chat passes of a 10-question Kimi Search citation pilot completed. Across the Search pilot, 20/20 factual answers matched official ground truth and 18/20 citations fully matched both the fact and requested product scope. All 13 unique citation URLs resolved to their expected official destinations; 12 exposed readable documentation, while the GitHub repository exposed the correct repository identity alongside a transient rendering warning. No final citation URL was fabricated. The same Q2 scope miss appeared in both passes: an official Files API page was used to support a Kimi Chat limit. The Deep Research attempt and a separate submission from Kimi’s official
/agentroute were blocked by a priority-queue gate before task execution. Kimi Code installation, version and failing-fixture baseline were verified, but the authenticated OAuth agent run was blocked when membership benefits could not be verified. Every statement in this baseline is scoped to that dated account and environment; the later dated updates below supersede any changed access or result state.
Batch 3 update, reviewed August 5, 2026: Six additional reports now expose their dated setup, fixture and limits. Websites, Sheets, Slides and Docs were blocked before generation or review produced an output. Kimi Work stopped before a first model token or Agent tool action; all four mounted source-file hashes still matched afterward, which verifies non-mutation in that blocked attempt but does not test the permission safeguard. The Skills setup produced a downloadable
SKILL.mdin an ordinary Agent task, but it did not install the candidate in Custom Skills; the real/skill-creatorattempt was then blocked, so the five-task comparison remains not run. None of these access states is scored as product quality.
Batch 4 update, reviewed August 5, 2026: The K3 record advanced from an access-only protocol to a completed single-pass API pilot under frozen protocol v1.1. All 15 calls completed: tool calls 8/8, code tests 5/5, vision fields 6/6 and long-context fields 20/20 across five requests from approximately 32K through 950K prompt tokens. The usage-derived cost was $5.687397; a read-only balance check 13 minutes 41 seconds later showed a $5.68739 deduction. Every case ran once, so this is not a reliability benchmark or universal K3 score. At that August 5 review, Agent Swarm, Claw, WebBridge, Scheduled Tasks and Memory retained separate access or not-run states; later dated updates supersede those historical labels.
Batch 5 update, reviewed August 5, 2026: Three additional records cover Presets, Plugins and Vision. The Presets entry point was absent from the checked signed-in web interface, so no Preset was created; 0/10 planned messages and 0/5 pairs ran. The World Bank plugin was installed and selectable without OAuth, but the only send attempt met a subscription-priority gate before chat creation or plugin invocation; no output, accuracy score, credit use or cost was measured. One Vision submission completed across six synthetic PNG images and one six-second WebM video, passing 35/36 frozen atomic checks. The completed response displayed a switch to K2.6 Instant, so this result is not attributed to K3 and is not a product-wide accuracy claim.
Batch 7 Memory update, reviewed August 6, 2026: The signed-in consumer Memory test completed under namespace
KIAG-MEM-20260806-V1. Five separately saved synthetic facts were recalled exactly in five individual fresh chats and one bundled fresh chat; a targeted update returned onlyThursday, a deleted checksum returnedNOT STORED, and five never-saved fields all returnedNOT STORED. The four remaining test entries were removed individually and the final namespace check returnedNO. The frozen rubric scored 100/100. Every scored fresh response displayed a high-demand switch to K2.6 Instant, and the visible cleanup result is not a claim of backend erasure.
Batch 7 Scheduled Tasks update, reviewed August 6, 2026: One synthetic Daily cloud task was created under a frozen exact-token protocol. Creation, edit, pause, resume and cleanup persisted, and no run appeared during the paused observation window. The manual
Run once nowcontrol returned Kimi’s service-busy error. For the only authoritative scheduled attempt, the card advanced to tomorrow at 07:30, but no new history row, result conversation or notification appeared through 14:35:14.296Z, 5 minutes 14.296 seconds after its due time. The event is a bounded failure, not a reliability rate; the test title was deleted and the visible task-card count returned to zero.
Kimi AI Guide is independent and is not affiliated with Moonshot AI.
Test index
| Test | Reader question | Protocol | Latest run | Evidence status |
|---|---|---|---|---|
| Kimi Projects: Files, Instructions, Memory and Context | Did a controlled fixture, Project instructions, filename and marker carry into two separate chats? | Controlled two-chat run | Completed August 4, 2026 | Across both chats, 4/4 requested values were exact; both responses followed the table instruction and included the filename and marker; about 29.9s and 28.7s using K2.6 Instant fallback |
| Kimi Search Citation Accuracy Test | Do short Kimi Search answers match official facts and cite the correct product scope? | Two fresh-chat passes of the same 10-question prompt | Passes 1 and 2 completed August 4, 2026 using K2.6 Instant fallback | Combined: 20/20 factual answers matched official ground truth; 18/20 citations fully matched fact and product scope; 13/13 unique URLs resolved to the expected official destination, with 12 readable documentation pages and one GitHub rendering warning; zero fabricated final URLs |
| Kimi Deep Research: How It Works, Access Test and Limits | Can a complete research report cover the brief and support its claims with appropriate primary sources? | v1.0 planned | Access attempt August 4, 2026 | Logged in; Deep Research and K3 High selected; priority-queue message blocked execution before research began |
| Kimi Agent: Capabilities, Limits and a Live Access Check | Can Kimi Agent complete a bounded real task and deliver the requested artifact? | Planned | Access attempt August 4, 2026 | Priority-queue message blocked task execution; no Agent performance result |
| Kimi Code CLI: Setup, Models and Login Test Results | Can a new user install Kimi Code and complete a small repository task under recorded conditions? | Controlled local fixture | Setup attempt August 4, 2026 | CLI 0.32.0 and failing baseline verified; OAuth agent run blocked by membership-benefit verification before code editing |
| Kimi Websites: Builder Guide, Limits and Access Test | Can Kimi Websites build and export a preview-only calculator that passes 12 fixed functional, mobile and accessibility checks? | Twelve-condition calculator fixture | Access attempt August 4, 2026 | High-demand message appeared before generation; no site, preview, code or export existed, and none of the 12 conditions was scored |
| Kimi Work: Desktop Agent, Permissions and Safe-Folder Test | Can Kimi Work inspect a disposable folder and create only one authorized output while preserving the source files? | Safe-folder fixture plus before/after hash manifests | Access attempt August 4, 2026 | Version and signed executable recorded; turn stopped before a first token, permission prompt or tool call; 4/4 source hashes matched afterward |
| Kimi Sheets: Formulas, Charts and Accuracy Test | Can Kimi Sheets preserve raw data and produce correct formulas, summaries, charts and an editable XLSX from a known CSV? | kimi-sheets-sales-quality-v1 | Access attempt August 4, 2026 | CSV reached the entry UI; high-demand alert appeared before workbook generation, so formula, chart and export checks remain not testable |
| Kimi Slides: K3 Workflow, Exports and Test Status | Can Kimi preserve a fixed eight-slide source, avoid overflow and export an editable PPTX? | KS-SLIDES-8-V1 | Access attempt August 4, 2026 | Source uploaded; high-demand alert appeared before generation, so no preview, deck, download or score exists |
| Kimi Docs: Word Review, Revisions and PDF Test | Can Kimi detect 16 known document defects while preserving correct content, tables, links and editable exports? | kimi-docs-operations-memo-v1 | Access attempt August 4, 2026 | DOCX completed entry-stage analysis; high-demand alert appeared before review, comments, revisions or export began |
| Kimi Skills: Creation, Usage and Consistency Test | Does an installed Custom Skill improve consistency across five fixed tasks compared with the same SOP used as an ordinary prompt? | KS-SKILLS-COMPARE-V1 | Setup attempts August 4, 2026 | Ordinary Agent returned a SKILL.md but did not install a Custom Skill; the embedded creator was blocked, leaving 0/5 paired tasks run |
| Kimi K3 Real-World Test: Code, Context, Tools and Cost | Can K3 repair code, produce exact tool calls, read visual values and retrieve long-context facts under preregistered rules? | Four-track single-pass pilot v1.1 | 15 API calls completed August 5, 2026 | 15/15 calls completed with kimi-k3: tools 8/8, code 5/5 tests, vision 6/6 fields and context 20/20 fields from approximately 32K through 950K prompt tokens; usage-derived cost $5.687397 and later balance deduction $5.68739 |
| Kimi Agent Swarm: How It Works and Test Status | Does Swarm improve completeness, evidence quality or time on a fixed 40-source task compared with Standard Agent? | Frozen Standard-versus-Swarm ledger | Access checked August 5, 2026 | Signed-in entry point observed; executable membership or credits were not verified and neither task was submitted |
| Kimi Claw: Setup, Memory, Files and Automation Test | Can an authorized Claw instance use Memory, files, a Skill and HEARTBEAT without leaving the fixed scope? | Synthetic workspace and fixed cases | Controls checked August 5, 2026 | Controls observed, but no usable authorized instance was confirmed; nothing was deployed, linked, scheduled or changed |
| Kimi WebBridge: Setup, Permissions and Safe Browser Test | Can WebBridge complete twelve fixed actions on a localhost fixture without leaving the allowed browser scope? | Twelve-check localhost protocol | Preparation checked August 5, 2026 | WebBridge was not installed or connected; no browser action ran and fixture values are not results |
| Kimi Scheduled Tasks: Timing, Notifications and Missed Runs | Does a synthetic scheduled prompt preserve its configuration, pause safely and produce an observable result? | One-task controlled cloud functional pilot | Completed August 6, 2026 | Creation, edit, pause, resume and cleanup persisted; manual Run once returned a service-busy error; the authoritative scheduled attempt advanced to tomorrow but exposed no new history, result conversation or notification through +5m 14.296s |
| Kimi Memory: What Persists Across Chats? | Can saved information be recalled, corrected and deleted without introducing unsupported memories? | KIAG-MEM-20260806-V1 controlled save-recall-update-delete run | Completed August 6, 2026; fresh responses routed to K2.6 Instant | 100/100: 5/5 individual recalls and bundled recall exact; update and visible deletion passed; five never-stored fields rejected; targeted cleanup completed |
| Kimi Presets: Setup, Limits and a Controlled Consistency Test | Does inserting a saved Preset preserve the frozen input and produce consistent outputs across five paired tasks versus the same text pasted manually? | KP-PRESETS-PAIR-V1 | Access checked August 5, 2026 | No Presets entry appeared in the composer + menu, slash menu or My Kimi navigation; no Preset was created, 0/10 messages and 0/5 pairs ran, and no consistency score exists |
| Kimi Plugins: Setup, Permissions, Credits and a Safe Test | Can the World Bank plugin return pre-frozen public population values without OAuth, write access or an unapproved charge? | KP-PLUGINS-WB-V1 | Invocation attempt August 5, 2026 | Plugin installed and selectable without OAuth; a subscription-priority gate appeared before chat creation or plugin invocation, leaving 0 accepted chats, 0 invocations and no output, accuracy score, measured credit use or cost |
| Kimi Vision: Images, OCR, Charts and Video Test | Can Kimi extract exact fields from synthetic OCR, table, chart, interface, scene and short-video fixtures in one frozen submission? | KVISION-SYNTHETIC-V1 | One run completed August 5, 2026; completed response routed to K2.6 Instant | 35/36 atomic checks passed; the only miss returned null for table.missing_cell where the frozen expected answer was South Q3; one run is not a general reliability or K3 score |
Blocked and planned do not mean a score of zero. The Projects observations come from one controlled two-chat run. The Search figures combine two fresh chats given the same prompt; they document those 20 answers and citation checks, not general reliability. The K3 figures come from one pass per case or context band and do not establish repeatability. Presets and Plugins stopped before any output that could be scored. The consumer-chat Vision figure comes from one frozen submission routed to K2.6 Instant and is separate from the K3 API vision pilot; neither should be generalized to other inputs. The Memory score belongs only to its five-entry synthetic fixture and one dated account run; visible cleanup does not establish backend erasure. The Scheduled Tasks finding is one bounded event-level failure, not an uptime or delivery-rate estimate.
The Batch 3 fixtures validate the test inputs and acceptance rules only. Where generation did not begin, expected values remain the oracle and must not be presented as Kimi output.
What the completed observations mean
- Projects: In the tested Project, the two chats together reproduced all four requested fixture values. Each chat returned its two requested values exactly, followed the required table format, identified the source filename and ended with the required marker. The first chat took about 29.9 seconds and the second about 28.7 seconds. The interface used a K2.6 Instant fallback. This verifies those two outputs only; it does not prove that every Project file is always retrieved or that every model behaves the same way.
- Search citation pilot: Two fresh chats received the same 10-question prompt while the interface displayed K2.6 Instant fallback. Each pass returned 10/10 factual answers matching the official ground truth and 9/10 citations that fully supported both the fact and requested product scope. In both passes, Question 2’s numeric answer was correct, but the cited official page documented the Files API, not Kimi Chat. Combined, the two passes produced 20/20 correct facts and 18/20 full-scope citations. All final citations were official, and all 13 unique URLs resolved to the expected official destination. Twelve exposed readable documentation; the GitHub repository exposed the correct repository identity alongside a transient rendering warning. No final citation URL was fabricated. Run 2 omitted answer numbering, and neither run numbered its final
Sources Checkedlist; those format misses were recorded separately from fact and citation scoring. Repeating the same Q2 scope error is evidence of that issue in these two runs, not a general reliability estimate. - Kimi K3 API pilot: Fifteen single-pass calls under protocol v1.1 all completed with the returned model ID
kimi-k3. The eight tool cases passed 8/8, the code patch passed all 5/5 fixed tests within the allowed file scope, the image response matched 6/6 numeric fields, and the five long-context requests matched 20/20 fields from approximately 32K through 950K prompt tokens. The usage-derived cost was $5.687397; a later balance check showed a $5.68739 deduction. One successful pass per case or band does not establish repeatability, broad coding ability, general vision accuracy or guaranteed near-limit recall. - Kimi
/agentroute and Deep Research: The Deep Research interface was selected for one attempt. A separate attempt was submitted from Kimi’s recorded official/agentroute with K3 High selected; its screenshot does not show a separate Agent-mode badge. A priority-queue subscription message appeared before either task began. These are access observations, not performance failures. - Kimi Code: The exact package
@moonshot-ai/[email protected]was installed, the CLI reported0.32.0, and the four-test fixture baseline produced one pass and three expected failures with exit code1. OAuth then failed at membership-benefit verification before model selection, editing or the final test run. - Presets and Plugins: The Presets comparison produced no messages because the checked account exposed no Presets entry point. The World Bank plugin was installed and selectable, but the only send attempt stopped at a subscription-priority gate before chat creation or invocation. Neither record contains an output-quality score.
- Vision: One consumer-chat submission used six synthetic PNG images and one six-second WebM video. The completed response passed 35/36 frozen atomic checks; only
table.missing_celldiffered. The response identified a completed-route switch to K2.6 Instant, so the observation is attributed to that route rather than K3 and says nothing about repeatability beyond this fixture. - Memory: Five separately saved synthetic facts were recalled exactly in five individual fresh chats and one bundled fresh-chat table. A targeted update returned only
Thursday; a deleted checksum and five never-saved fields returnedNOT STORED. Four remaining tagged entries were then removed and a read-only namespace check returnedNO. The run scored 100/100 under its frozen rubric, with every scored fresh response displaying a K2.6 Instant high-demand switch. This is a small functional run, not a retention-duration, reliability or backend-erasure result.
Sample evidence from the lab




Lab standard at a glance
Every completed test must have a frozen question, dated environment, predeclared scoring rule, retained observations and visible limitations. Blocked access is recorded separately and is never converted into a zero or a simulated result. Retries and regenerations are logged rather than selected after the outcome is known.
The full rules for source priority, evidence levels, corrections, privacy and retesting live on our testing methodology and Sources and Corrections pages. Individual reports add their exact prompt, fixture, scoring units, run conditions and product-specific limits.
Current run record
| Test | Dated outcome | Evidence exposed on the public report | Boundary |
|---|---|---|---|
| Projects context | Two chats completed August 4, 2026; each returned its 2/2 requested fixture values exactly | Two original screenshots, fixture values, timings and scoring table | One synthetic text file; K2.6 Instant fallback |
| Search citations | Two 10-question passes completed August 4, 2026; 20/20 facts and 18/20 full-scope citations | Frozen prompt, question-level results, two screenshots and source-resolution record | First-party documentation only; format misses recorded separately |
| Deep Research | Access blocked before execution on August 4, 2026 | Original queue screenshot and planned scoring protocol | No report or performance result |
| Kimi Agent route | Access blocked before execution on August 4, 2026 | Original queue screenshot and recorded /agent route | K3 High selected; no output or performance result |
| Kimi Code | CLI 0.32.0 and four-test baseline verified; OAuth blocked before agent execution | Installer screenshot, baseline summary and redacted error | No authenticated model, edit or final test |
| Kimi Websites | Blocked before generation on August 4, 2026 | Fixed prompt, 12-condition oracle, entry screenshot and high-demand screenshot | No site, preview, code, export or public deployment |
| Kimi Work | Blocked before a first token or Agent tool action on August 4, 2026; 4/4 source hashes remained unchanged | App metadata, safe fixture, manifests, logs and screenshots | Non-mutation verified for a non-executing turn; permission enforcement and analysis were not tested |
| Kimi Sheets | CSV accepted by the entry UI; blocked before workbook generation on August 4, 2026 | Synthetic CSV, ground truth, 13/13 fixture checks, scoring file and preview image | No formulas, chart, XLSX or accuracy score |
| Kimi Slides | Source accepted; blocked before deck generation on August 4, 2026 | Fixed eight-slide source, ground truth and blank scoring record | No preview, PPTX, layout, editability or factual score |
| Kimi Docs | DOCX accepted and analyzed by the entry UI; blocked before review on August 4, 2026 | Synthetic 16-error DOCX, ground truth and structural/accessibility checks | No corrected DOCX, comments, revisions, PDF or render result |
| Kimi Skills | Ordinary Agent produced a reusable file, but no Custom Skill was installed; embedded creator was blocked on August 4, 2026 | SOP, validated DOCX conversion, five tasks, ground truth and blank scoring record | Creation partial; 0/5 paired comparisons and no consistency score |
| Kimi K3 | Single-pass API pilot completed August 5, 2026; 15/15 calls completed | Protocol v1.1, redacted requests and responses, per-call scores, generated corpora, code diff and tests, metrics, ledger, hashes and independent audit | Tools 8/8, code 5/5, vision 6/6 and context 20/20 at approximately 32K–950K; usage-derived cost $5.687397, later balance deduction $5.68739; one pass only |
| Agent Swarm | Signed-in entry point observed August 5, 2026; no task submitted | Frozen 40-source ledger and Standard-versus-Swarm scoring protocol | Membership or credits not verified; no output, timing, consumption or comparison result |
| Kimi Claw | Controls observed August 5, 2026; no usable authorized instance confirmed | Synthetic workspace, fixed cases and stop conditions | No deployment, channel, Memory, file, Skill, schedule or reliability result |
| Kimi WebBridge | Protocol prepared August 5, 2026; extension and bridge not connected | Localhost fixture, expected-results file and twelve checks | No browser action, screenshot, extraction, retry, timing or privacy-path result |
| Scheduled Tasks | One-task cloud functional pilot completed August 6, 2026; scheduled result not observed and cleanup verified | Frozen exact-token prompt, UTC/local clock record, event ledger and three original screenshots | Manual control failed with a service-busy error; edit, pause and resume states persisted; authoritative due time exposed no new history, result or notification through +5m 14.296s; one event is not a reliability rate |
| Kimi Memory | Controlled run completed August 6, 2026; 100/100 | Exact prompts, synthetic fixture, row-level ledger and three original screenshots | One account and five entries; K2.6 Instant routing; timing not measured; visible cleanup is not backend-erasure proof |
| Kimi Presets | Entry point absent in the checked signed-in web interface August 5, 2026; no Preset created and 0/10 messages or 0/5 pairs ran | KP-PRESETS-PAIR-V1, fixed Preset text, five frozen tasks, oracles, run order and blank scoring rows | Account-, locale- and surface-specific access observation; no input-fidelity, output-consistency or cleanup result |
| Kimi Plugins | World Bank plugin installed and selectable without OAuth; subscription-priority gate stopped the only send attempt before chat creation or invocation August 5, 2026 | KP-PLUGINS-WB-V1, exact prompt, independent World Bank oracle, permission and side-effect checks | 0 accepted chats, 0 invocations and no plugin output, accuracy, latency, measured credit use or cost result |
| Kimi Vision | One submission completed August 5, 2026; 35/36 atomic checks passed after the response displayed a K2.6 Instant route switch | KVISION-SYNTHETIC-V1, six synthetic PNGs, one six-second WebM, frozen ground truth, exact prompt, raw JSON and scoring record | One run and one fixture set; not a K3 score, product-wide accuracy claim or repeated reliability measurement |
Evidence availability
The public reports expose the evidence needed to understand their claims: dated conditions, exact or safely normalized prompts, result tables, official source links and publication-safe screenshots. Working logs, raw response text and fixtures are retained internally when redistribution, privacy or credential hygiene makes a public bundle inappropriate.
No downloadable archive or CSV is currently offered from this hub. An evidence filename in a table is an internal record identifier, not a public download link. If a public bundle is added later, it will state its contents, redactions, protocol version, file format and checksum.
How to read a result
Before applying a result to your own work, check the run date, product surface, visible model or fallback, account/access conditions, sample size, scoring rule and stated exclusions. A correct result on a small fixed fixture does not establish general reliability, and a blocked attempt does not establish poor quality.
Corrections and independent reproductions
To challenge a row or reproduce a test, email [email protected] with the test URL, disputed statement, primary evidence, product/model labels, date and a safely redacted output. Material corrections follow our Sources and Corrections policy; historical runs remain visible when a later product version behaves differently.
Frequently asked questions
Are these official Kimi benchmarks?
No. The Test Lab is an independent project by Kimi AI Guide. Vendor benchmarks and product claims remain labeled as official claims until independently reproduced.
Do blocked runs receive a zero?
No. A run stopped by login, entitlement, queue, credits or safety conditions is reported as blocked and excluded from quality scoring.
Can readers rerun the tests?
Yes when the report publishes a safe prompt or fixture. Product updates, model routing and account conditions can change the outcome, so a reproduction must record its own environment and date.
Why is there no single Kimi score?
Search citation support, file-context recall, coding edits and artifact creation are different tasks. We publish component results rather than hiding those differences inside one number.
Official product references
- Kimi Agent overview
- Kimi Projects
- Kimi Search overview
- Kimi Deep Research
- Kimi Code documentation
- Kimi Open Platform documentation
- Kimi Websites overview
- Kimi Work overview
- Kimi Docs and Sheets overview
- Kimi Slides overview
- Kimi Skills overview
- Kimi Presets documentation
- Kimi Plugins documentation
- Kimi multimodal and Agentic Chat overview
