Out of a document

OCR a scanned contract without uploading it

A scanned agreement, read on this device.

On this device

Choose the scanned contract

A PDF scanned from paper. The reader downloads once, about seventeen megabytes, and a page takes about a second on a graphics chip. A PDF that already has text is refused by name — it does not need this.

    Opened on this device · never transmitted

    What this tool does, exactly

    A signed agreement is the document people are least willing to hand to a website, and the one they most often need as text: to search it, quote a clause, or find the date buried on page nine. This page reads it here — every page drawn and read in this browser tab — and gives back the text, the text as Markdown with its clause numbering, and the scan itself made searchable. Nothing about it is sent anywhere.

    Why it stays here
    A contract carries names, sums, signatures and often a confidentiality clause of its own; a document under legal privilege or a non-disclosure agreement is exactly the kind you may not be allowed to hand to a third party. There is no server behind this page and no request that carries a file: the network panel shows the reader's own code arriving from this site, and nothing carrying your file leaving. That is a fact about the mechanism, not legal advice about what you may disclose.
    What comes back
    Three files. The text in reading order, for searching and quoting. A Markdown version in which each line is kept as read, so a clause number stays at the start of its line and 7.2 is still findable as 7.2. And the scan with an invisible text layer laid over the original pages, which opens in any viewer and can be searched, with the layer read back before the file is offered.
    Signatures, stamps and handwriting
    The reader reads print. A signature is not text and becomes nothing, or a short doubtful line the page lists by name. A stamp over a paragraph was not measured; expect the lines under it to read worse and to appear in the doubtful-lines list. Handwriting was not measured and is not promised — check any handwritten addition against the scan.
    Tables and schedules
    The fast reader gives a table as lines of cells in reading order. The Deep reader behind the Pro door reads a page as a page, so a fee schedule comes back as rows of cells in the Markdown. It takes about twenty seconds a page on a laptop's graphics chip and is free during the preview.
    What it was measured at
    0.23 % of characters wrong on a synthetic 200 dpi scan of a printed English page, 0.71 % in Portuguese, 0.64 % in Russian with the second recogniser, in this site's own tests. The one systematic slip on English print: an old-style zero — the kind that sits low like an o — is read as the letter, which matters in account numbers. Check those.

    About this tool

    Is the searchable PDF still the original scan?

    Yes. The original page images are untouched; the recognised text is laid over them invisibly, at the positions the words were found, and the file is reopened to check that the layer reads back before it is offered. What you see is the scan; what you can search is the reading of it.

    Can I rely on the text for a quotation?

    Rely on it to find the passage, then read the passage in the scan. The reader was measured at about two characters wrong in a thousand on clean print; a quoted clause should come from the page image, not from the recognised text. The doubtful-lines list under the result names the lines it was least sure of.

    Does the contract, or any part of it, leave my device?

    No. The pages are drawn and read in this browser tab. The only requests this site makes are for its own code, its page-view counter — which receives the page address and nothing from your file — and, only if you open the Pro door, the Deep reader's model from Hugging Face. The privacy page sets out each.

    What about a contract in Russian, Chinese or Portuguese?

    The default reader covers English and forty-six other Latin-script languages, Chinese and Japanese; Russian, Ukrainian and Belarusian are a choice in the options that fetches a second recogniser. There is a separate page for scanned Chinese documents with its own measured figure.

    More out of a document tools · All 52 · How they work