Kimi AI Guide uses a documented process to distinguish official claims, interface observations and independently measured results. This page explains that process so readers can judge the strength and limits of our evidence.
Our core rule is simple: if a test was not completed, we do not publish a result for it. When access, login, credit, software or another prerequisite is unavailable, we stop and label the coverage accordingly.
Evidence levels used on this site
1. Officially documented
The claim appears in current Kimi or Moonshot AI documentation, a product page, help article, model report, repository or legal page. We link to the exact source and record the verification date.
This confirms what the vendor publishes. It does not prove that a feature works for every account, region or workload.
2. Interface verified
We observed the relevant page or control in an authenticated or public interface. We record the page, date, account access level and important state without exposing private account data.
Seeing a button or model name does not prove successful execution. Interface availability and live functionality are reported separately.
3. Live tested
We completed a defined task and retained sufficient evidence to reproduce or audit it. The report states the prompt or procedure, settings, repetitions, output and limitations.
4. Measured
We collected a numerical result such as latency, token use, cost, success rate or file-processing time under a stated environment. Measurements include units and sample size. They are not generalized beyond the tested conditions.
5. Unverified or conflicting
The claim cannot be confirmed, requires unavailable access, or conflicts with another credible source. We explain the problem and do not turn it into a definite conclusion.
Documentation verification protocol
For prices, models, limits and product rules:
- Open the official English documentation or product page.
- Record the direct URL and date.
- Identify the product: Membership, Kimi Code, Open Platform API or Business.
- Capture the unit and qualifier, such as “per 1M tokens,” “annual total” or “excluding tax.”
- Check a second relevant official page when available.
- Note any conflict, stale wording or region-dependent value.
- Prefer the most specific current source, while disclosing unresolved disagreement.
We do not use a Google search snippet as the final source for a price.
Product and interface tests
Before a product test we define:
- the reader question;
- the required account or plan;
- the exact starting state;
- the task steps;
- what counts as success, partial success or failure;
- the evidence to retain; and
- the stop conditions.
We avoid changing or deleting user data when a read-only check can answer the question. If a test could incur charges, publish content, send messages or affect an external account, the scope and cost limit must be set before execution.
Login and access limitations
If the required login is unavailable, we do not simulate the result. If a feature is visible but locked behind a higher plan, we report that access state rather than claiming the feature was tested.
Plan, region and staged rollout can affect availability. Every access-dependent result should state what account level was used without exposing identifying account details.
API test protocol
A reproducible Kimi API test should record:
| Field | What we record |
|---|---|
| Date and time | UTC timestamp or stated time zone |
| Endpoint | For example, /v1/chat/completions |
| Model | Exact API model ID |
| SDK / client | Name and version, or cURL version |
| Parameters | Temperature, reasoning/thinking mode, output limit, tools and streaming state as relevant |
| Input | Exact prompt or a downloadable non-sensitive fixture |
| Repetitions | Number of independent runs |
| Response | Status, finish reason and redacted output |
| Usage | Input, cached input, output and total tokens when returned |
| Cost | Formula, official rate date and calculated or invoiced amount |
| Errors | Status code, retry behavior and unresolved failure |
API keys are never included in logs, screenshots or downloadable files. A leaked key is revoked, not redacted and reused.
A test is not a benchmark by default
A single prompt can show whether an example worked. It cannot establish general model quality. A comparative benchmark needs:
- a defined task set;
- the same inputs and scoring rules for every model;
- enough repetitions for nondeterministic outputs;
- version and parameter control;
- blinded or objective scoring where practical; and
- disclosure of excluded or failed runs.
Vendor benchmark results are labeled as vendor-reported unless we reproduce them.
Nondeterminism and repeatability
AI outputs can differ across runs even with the same prompt. We therefore:
- avoid treating one response as a universal capability;
- preserve exact prompts and settings;
- repeat tests when the claim concerns reliability;
- report variation and failures, not only the best output; and
- avoid “always” and “never” unless the evidence supports them.
When a model is updated behind the same product name, older results remain historical observations, not a current guarantee.
Latency measurements
When latency matters, we distinguish:
- time to first token;
- total completion time;
- output token count;
- approximate output tokens per second; and
- network and client overhead.
We record region, connection type, streaming state, model, prompt size and sample count. We do not compare latency measured on different days or networks without a warning.
Pricing calculations
Our API cost examples use published rates and explicit arithmetic. They are labeled calculated unless checked against an actual usage record or invoice.
A complete calculation can include:
- cache-hit input tokens;
- cache-miss input tokens;
- output tokens;
- tool invocation fees;
- additional search-result tokens;
- batch discounts where applicable; and
- tax, if known.
If two official pages show different rates, we publish the conflict and choose a source hierarchy rather than averaging the numbers.
Screenshots and original evidence
Screenshots should be captured from the relevant page during the stated check. Before publication we:
- remove or obscure email addresses, account IDs, payment information, API keys and unrelated private content;
- crop only when the surrounding interface is not needed to interpret the evidence;
- keep enough context to show what the screenshot represents;
- add descriptive alternative text and a factual caption;
- record the capture date; and
- avoid modifying the evidence in a way that changes its meaning.
On this site, evidence images should normally be inserted at full file size, displayed at 75% width, centered and without a click-through link.
Test files and downloadable evidence
We provide a prompt file, input document, data table or result log when it materially improves reproducibility and can be shared safely. Downloadable evidence is not added merely to make a page appear more technical.
Files must not contain:
- secrets or private account details;
- copyrighted material we lack permission to redistribute;
- personal data unrelated to the test; or
- generated values presented as measured observations.
Updating old results
High-change claims receive a visible verification date. When a product changes:
- documentation-only facts are rechecked against the current source;
- a live-test result keeps its original test date;
- a new test is run before calling the old result current;
- material changes are explained; and
- URLs and publication dates are preserved when an existing page is updated.
We do not erase an old test date and replace it with “updated today” unless the test was actually repeated.
Standard test report template
Every substantial test should answer:
- Question: What are we trying to learn?
- Date and access: When and with which product/plan?
- Environment: Device, software, SDK, region or network details that matter.
- Method: Exact steps, prompt, inputs and settings.
- Success criteria: How is the result judged?
- Results: All relevant observations, including failures.
- Evidence: Screenshots, logs or downloadable fixtures, safely redacted.
- Limitations: What this test cannot establish.
- Sources: Official material used to interpret the result.
- Reproduction notes: What another reader needs to repeat it.
Corrections and challenges
Readers are invited to challenge our method or reproduce a result. Send the page URL, disputed result, your environment, evidence and the date to [email protected].
We review corrections under our Editorial Policy and document material changes on the affected page.
