Out of a document

OCR a scanned Chinese PDF on your device

Simplified Chinese scans, read on this device, with a measured error rate.

On this device

Choose the scanned Chinese PDF

Leave the language on its default: the same model reads Chinese, Japanese and the Latin-script languages. The reader downloads once, about seventeen megabytes.

    Opened on this device · never transmitted

    What this tool does, exactly

    The reader on this page is PaddleOCR's — a model built in Chinese first — running in this browser tab rather than on somebody's server. A scanned contract, a form, a mixed Chinese and English page: every page is drawn and read here, and you get the text, the text as Markdown, and the scan made searchable, with the layer read back before the file is offered.

    What was measured
    A printed page of Simplified Chinese with English names and Arabic numerals mixed in, drawn at 200 dpi in a system font. This site's own test read 0.97 % of its characters wrong — under one in a hundred — with every line found, in about 1.4 seconds on one processor thread. Traditional Chinese, vertical text and handwriting were not measured and are not promised.
    Which model, and why it is the default
    PP-OCRv6 from PaddleOCR, a project whose home language is Chinese; one recogniser covers Chinese, Japanese and the Latin scripts, so a bilingual page needs no choice and no second download. Russian, Ukrainian and Belarusian are the one script that needs a second recogniser, and that is a choice in the options.
    Mixed pages
    A Chinese page with an English company name, a date in Arabic numerals or a product code is read as one page: the recogniser holds all three in one character set. Full-width punctuation — ,。:— is kept as it was printed.
    The searchable PDF
    The recognised characters are laid invisibly over the original pages and encoded so that a viewer's search finds them: the site's test reopens the file and looks for the words that were written. Copying from the searchable file gives the recognised text; the page image underneath is the untouched scan.
    Where it stops
    Vertical text is not laid out as columns; the reader finds horizontal lines. Handwritten characters are not read reliably. A coarse scan loses strokes, and a character with a lost stroke is read as its nearest neighbour — the doubtful-lines list under the result names the lines it was least sure of.

    About this tool

    Does it read Traditional Chinese?

    PaddleOCR lists Traditional Chinese among the scripts this recogniser covers; this site has not measured it and so quotes no figure. Try a page and read the doubtful-lines list — it is the honest answer for a script that was not in the test.

    Does it translate?

    No. It reads the characters on the page into text you can search, copy and paste into a translator of your choosing. Translation is a different job and one that would need either a large model or a server; this page has neither.

    Can I search the PDF for Chinese words afterwards?

    Yes. The invisible layer carries the recognised characters with their Unicode mapping. The site's own test reopens the file in the same PDF library and finds the words that were written; other viewers and desktop search read the same text layer, but were not tested here.

    Does the document leave my device?

    No. The pages are drawn and read in this browser tab; the reader's own code and models come from this site once and are cached. The only other request is the page-view counter, which receives the page address and nothing from your file.

    More out of a document tools · All 52 · How they work