How to Use Kimi AI with PDFs and Long Documents

Kimi can work with PDFs, Word files, spreadsheets, presentations, images, text and video according to its current official Help Center. Uploading a file, however, does not prove that every page was read correctly or that a summary is complete.

The reliable approach is to treat document analysis as an auditable workflow: inspect the file, define coverage, request page-level evidence and verify the answer against the original. This guide gives you that workflow without claiming unperformed product tests.

Independent guide: Kimi AI Guide is not affiliated with Moonshot AI. Product limits below are attributed to official Kimi documentation checked on August 4, 2026 and can change.

Current documented file support and limits

Kimi’s official overview currently lists support for PDF, Word, Excel, PowerPoint, images, TXT and video, with up to 100 MB per file and up to 50 files per session. A separate troubleshooting article says that an oversized file should be split and uploaded in logical batches.

ItemCurrent official descriptionHow to interpret it safely
Input formatsPDF, Word, Excel, PPT, images, TXT and videoFormat support does not guarantee accurate extraction from every layout or scan
Per-file sizeUp to 100 MBA smaller, clean file can still work better than a maximum-size file
Files per sessionUp to 50Many files increase naming, coverage and context-management risk
K2.6 conversation contextOfficial troubleshooting describes about 128K tokens and gives an approximate word comparisonToken-to-word ratios vary; plan by sections, not the marketing-style word equivalent
K3 long contextOfficial guidance describes a 1-million-token context; an Extra Long conversation option is tied to top-tier membership in the troubleshooting articleAvailability and usable capacity depend on account, mode, file extraction and other conversation content
Editable document outputOfficial guidance assigns end-to-end editable Office output to K3/Agent workflowsPlain Chat and K2.6 may provide text rather than a downloadable Office file

These are consumer-product ceilings and documented capabilities, not an independent accuracy score. Do not reuse a Chat context figure as an API specification; the developer platform publishes separate model limits. A 100 MB file that uploads successfully can still contain unreadable scans, missing fonts, rotated pages or tables that need manual checking.

Before uploading a long text file, use the Kimi Context & File Fit Checker for a cautious planning estimate. It counts pasted or locally selected plain text in your browser and compares an estimated token range with documented Chat, API and Kimi Code windows. It does not upload the file, provide an exact token count or guarantee that Kimi will parse the complete document.

The eight-step document workflow

1. Define the decision or deliverable

“Summarize this PDF” is often too vague. Choose a concrete output:

  • an executive summary for a named audience;
  • a table of contractual obligations;
  • a list of findings and supporting pages;
  • a comparison across document versions;
  • a data-extraction sheet with validation rules; or
  • a question-answer set limited to the uploaded material.

Write the acceptance criteria before upload. For example:

Deliverable: A two-page briefing for a non-technical director.
Coverage: Every numbered section and appendix in the report.
Evidence: Page number and short supporting passage for each factual claim.
Exclusions: No outside sources and no unstated assumptions.
Uncertainty: Mark unreadable or contradictory passages for manual review.

2. Inspect and classify the file

Open the original outside Kimi and record:

FieldWhat to record
FilenameA short unique name, such as annual-report-2025.pdf
Document titleThe title printed inside the document
Version/datePublication or revision date
LengthPage count or sheet/slide count
Text typeSelectable text, image scan or mixed
StructureContents page, numbered sections, appendices
RiskPersonal, confidential, licensed or regulated material
Expected outputSummary, extraction, comparison or transformation

If the file is a scan, test whether you can select and copy a sentence. OCR may misread columns, small print, handwriting, mathematical notation and tables. Plan extra visual checks.

3. Remove unnecessary sensitive data

Redact information the task does not require. Typical candidates include:

  • identity and passport numbers;
  • account, payment and authentication details;
  • health information;
  • private addresses and signatures;
  • legal-privilege or confidential business material; and
  • API keys, passwords and one-time codes.

Check whether you have permission to upload the document. A file being available to you does not automatically grant the right to send it to an AI service.

Kimi’s current consumer data-usage article says suitably processed consumer content may be used for model training and documents an account-level opt-out route. It also says per-file opt-out is not currently available. Review the current policy for your product before uploading sensitive or proprietary material.

4. Build a document coverage ledger

The coverage ledger is a simple table that makes omissions visible. Create it from the document’s table of contents or page ranges before asking for a final summary.

Section IDTitlePagesMust answerStatusEvidence checked
S1Executive summary1–3Main conclusionsPendingNo
S2Methodology4–12Sample, dates, limitationsPendingNo
S3Results13–38Findings and key figuresPendingNo
A1Appendix A39–45Definitions and exceptionsPendingNo

Ask Kimi to fill the Status column as Covered, Partial, Unreadable or Not found. Verify the page references yourself. A completed-looking table is not evidence until it matches the file.

For a very long document, work section by section, then synthesize from reviewed section notes rather than repeatedly uploading or summarizing the entire file.

5. Upload with a file manifest

When more than one file is involved, tell Kimi exactly what each file represents:

File manifest
- POLICY-2025.pdf: policy currently in force; 42 pages.
- POLICY-2024.pdf: previous version; 39 pages.
- DEFINITIONS.xlsx: approved term definitions; sheet “Terms” controls.

Task
Compare POLICY-2025 with POLICY-2024. Use DEFINITIONS.xlsx only to interpret
defined terms. Do not use web sources. For every change, cite both page numbers.
If a page is unreadable, list it rather than guessing.

Avoid names such as document.pdf, document-final.pdf and document-final2.pdf. Ambiguous filenames create attribution errors that good prompting cannot always repair.

If the same source files and instructions must remain available across separate chats, see the Kimi Projects guide. Its dated test used one original synthetic fixture across two new chats; it does not establish general file-parsing accuracy.

6. Request evidence before interpretation

For important documents, split extraction and analysis into two passes.

Pass A — extraction

Create an evidence table only. For each question below, give:
Question | Answer stated in the document | Page | Section | Short passage | Status

Status must be Direct, Inferred, Contradicted, Not found or Unreadable.
Do not provide recommendations yet.

Pass B — analysis

Using only the evidence table entries I confirm, write the analysis.
Keep direct findings separate from inferences. Do not fill gaps with general knowledge.

This prevents a polished recommendation from hiding weak extraction.

7. Run spot checks and adversarial checks

Choose pages that are easy and difficult to parse:

  • one ordinary text page;
  • one table or chart;
  • one footnote-heavy page;
  • one scanned or image-rich page; and
  • one appendix or exception section.

For each page, compare names, dates, units, negative signs, footnotes and qualifiers with the original.

Then ask questions designed to expose overreach:

Which requested answer is least supported by the document?
List conclusions that would change if the appendix controlled over the main text.
Find statements where “may,” “must,” “should” or “will” changes the obligation.

Do not accept Kimi’s self-review as the final validation. It helps identify where you should look.

8. Export a handoff with provenance

End with a compact record that another person can audit:

Create a handoff containing:
1. file manifest and versions;
2. task and exclusions;
3. verified findings with page references;
4. unresolved or unreadable sections;
5. calculations and assumptions;
6. human checks still required;
7. date and model used.

Keep the original files and your verified ledger. A generated document without its provenance is harder to correct later.

Long PDFs: when and how to split

Split a document when:

  • the upload exceeds a documented limit;
  • the conversation reports that the context is too long;
  • the file contains clearly independent volumes or appendices;
  • OCR quality varies sharply by section; or
  • you need separate reviewers for different chapters.

Split by logical boundaries, not arbitrary equal page counts. Keep the front matter, definitions and table of contents available to every relevant batch because they can control interpretation.

Use stable names:

REPORT-2026-00-front-matter.pdf
REPORT-2026-01-methodology.pdf
REPORT-2026-02-results.pdf
REPORT-2026-03-appendices.pdf

Then maintain a synthesis table:

BatchPagesReviewedKey findingsCross-references unresolved
00i–xiiYes/No
011–45Yes/No
0246–130Yes/No
03131–190Yes/No

Useful prompt recipes

Citation-grounded summary

Summarize this document for [audience] in [length].
Cover every section in the supplied coverage ledger.
After each factual sentence, include the document page in parentheses.
Use only this file. Mark missing, contradictory or unreadable evidence explicitly.
End with a table of sections and coverage status.

Compare two versions

Compare OLD.pdf and NEW.pdf.
Report additions, removals and materially changed wording.
For each change, quote no more than the minimum needed and cite the page in both files.
Separate semantic changes from formatting-only changes.
Do not infer legal effect; add a “requires expert review” column.

Extract a table from a PDF

Extract [fields] from pages [range] into a table.
Preserve original units and spelling.
Add Source page, Source row/label and Confidence columns.
Use “unreadable” rather than estimating a value.
After extraction, list duplicate IDs and validation failures.

Analyze a research paper

Create an evidence map with:
Research question | Dataset/sample | Method | Main result | Effect size | Limitation | Page
Keep author claims separate from your interpretation.
Do not generalize beyond the population and conditions in the paper.
List any result that cannot be matched to a table, figure or passage.

Files that need extra caution

Scanned PDFs

OCR can lose characters, reading order and table structure. Verify sample pages visually, and never assume a blank extraction means a blank page.

Spreadsheets

Check formulas, hidden sheets, filters, merged cells, units and date formats. A visible value can differ from the stored formula or underlying precision.

Contracts and policies

Definitions, exceptions, appendices and amendment order can control meaning. Use Kimi to organize review, not to replace qualified legal advice.

Financial documents

Reconcile subtotals, currencies, periods, signs and restatements. Do not rely on a generated total without a separate calculation.

Medical or personal records

These are highly sensitive and high stakes. Use only an approved environment and appropriate professional review. Redaction may not be sufficient when the remaining facts can re-identify a person.

Troubleshooting

Kimi says the file is too large

Compress images only if readability remains acceptable, remove irrelevant attachments or split at section boundaries. Keep a manifest and coverage ledger across the parts.

The summary omits an appendix

Add every appendix to the ledger and request its status explicitly. Ask for “Not found” instead of allowing a silent omission.

Page citations do not match

Check whether the PDF’s printed page numbers differ from the viewer’s page index. Define one system in the prompt: for example, “Use the printed page number; if absent, use PDF page index and label it PDF page.”

Tables are distorted

Ask for a direct row-by-row extraction with source coordinates before any analysis. Verify totals and a sample of cells manually. If the layout remains ambiguous, export the original table to a structured format through an approved method.

The chat cannot create a downloadable Word or Excel file

Kimi’s official troubleshooting directs Word and Excel generation to Agent mode and presentation output to Kimi Slides rather than the ordinary chat window. Current K3 documentation describes editable file generation, but the entry point and credit use can depend on the live interface and plan.

The conversation hits its length limit

Create and verify a handoff, then start a new chat. Carry forward only the manifest, reviewed evidence table, constraints and unresolved questions—not the entire unverified transcript.

A reproducible test plan for future updates

This page does not claim that the following test has been run. It is the method we recommend before publishing an independent accuracy result:

  1. Create or obtain a non-sensitive test PDF with known ground truth.
  2. Include ordinary text, a multi-column page, a table, a footnote, a scanned page and an appendix.
  3. Publish the file or a checksum and legally shareable excerpts.
  4. Record Kimi product, model, thinking level, plan, region and date.
  5. Use a fixed extraction prompt in a new session.
  6. Score field accuracy, page-citation accuracy, section coverage and unsupported claims.
  7. Repeat the run to reveal variability.
  8. Publish raw results, corrections and limits—not just an overall score.

Until that work is complete, documented capability and independent performance must remain separate.

Frequently asked questions

Can Kimi AI read PDFs?

Yes. PDF is included in Kimi’s current official list of supported file inputs. Parsing quality still depends on the file’s structure, text layer, scans and layout, so verify important passages against the original.

What is the Kimi file-size limit?

The official Kimi overview currently states up to 100 MB per file and up to 50 files per session. These limits were checked on August 4, 2026 and can change.

Can Kimi summarize a very long document?

Kimi documents long-context models and workflows, but practical success depends on the model, account, file extraction and other conversation content. Use a coverage ledger and section-by-section synthesis rather than assuming a successful upload means complete coverage.

Can Kimi compare two PDFs?

Official Kimi Docs guidance describes document comparison. For a reliable result, identify each version, require page references from both files and separate wording changes from conclusions that require expert review.

Is it safe to upload confidential files?

Do not assume so. Check authorization, minimize and redact data, and review the current policy for the exact Kimi product and account type. Consumer, enterprise and API data-use terms are not identical.

Related guides


Official sources for fact-checking

All sources accessed August 4, 2026.

  1. Kimi Help Center: Kimi overview
  2. Kimi Help Center: Getting started with Kimi
  3. Kimi Help Center: Kimi Docs & Kimi Sheets
  4. Kimi Help Center: Kimi chat common issues
  5. Kimi Help Center: Data Usage and Sharing
  6. Kimi Privacy Policy
  7. Kimi Terms of Service