Scanned PDFs

Chat with a scanned PDF, read by OCR

A photocopied handout or a photographed page has no text underneath to search. Pages like that are read with OCR first, then answered like any other document, with the page named each time.

What to ask a scanned document

What is on page 12?Which pages mention photosynthesis?What year is this past paper from?How many marks is question 3 worth?What definition is given for enthalpy?Where does the marking scheme start?Is the conclusion part of this scan?List the section headings hereWhich questions are compulsory?What does this handout say about titration?

A scanned PDF is a picture of a page rather than text. A photocopied handout, a chapter copied on a library machine, an old article that only ever existed on paper, a past paper somebody photographed on their phone: they all open normally and look entirely ordinary, but there is nothing underneath for software to read. The test takes two seconds. Drag your cursor across a sentence, and if individual words highlight, the file carries a real text layer. If a single flat box covers the page instead, or your reader's find command cannot locate a word you are looking straight at, the page is an image. Anything built to read a text layer finds an empty file there, which is why a scan so often produces either a flat refusal or, worse, an answer with nothing behind it.

Each page is checked as it uploads, and where a page has no text of its own it is read by optical character recognition instead: the image is examined and turned into text for the assistant to work from. A file that is half typed and half photocopied is handled page by page, so only the pages that need recognition get it. What comes out feeds the answers, not the viewer, which carries on showing you the scan exactly as it is. Recognition is not transcription, and it helps to hold that distinction: it is reliable enough to ask what a page covers, what it defines and what it claims, and not reliable enough to lift a quotation from. Answers still name their page, so the original image is always one click away.

What you get, on every document

Answers that cite the page
Replies name the pages they came from, as chips that scroll the viewer straight there so you can read the source yourself.
Grounded in your document
It answers from the file you uploaded and says so plainly when your question is not covered, rather than inventing something plausible.
PDFs up to 25 MB
One document per chat, shown beside the conversation with real text selection — highlight a passage and ask about that exact part.
Scans read automatically
Pages with no selectable text are read with OCR instead, page by page, so a scanned or photographed document still works.
Six interface languages
English, Bangla, Spanish, Indonesian, Arabic and German — and replies come back in whichever language you ask in.
Turn it into practice
The same document can become practice multiple-choice questions, spaced-repetition review and mock exams once you have read it.

Scanned PDFs FAQ

Do I need to switch OCR on?

No. Every page is checked as the file uploads, and recognition runs only on the pages that have no text of their own. A document that mixes typed pages with photocopied ones is handled correctly without you sorting them first, so you do not have to establish in advance which kind of file you are holding.

How accurate is the recognition?

On one test scan we measured roughly 83% of words read correctly. That is a single measurement on a single document rather than a figure to rely on, and your scan may land either side of it. As a working expectation: about five words in six is enough to answer what a page covers or how it defines something, and not enough to trust a sentence quoted back at you word for word. Open the page when the exact wording matters.

Why can't I highlight text on a scanned page?

Because the text you select in the viewer comes from the PDF's own text layer, and a scanned page has none — there is nothing there to select. Recognition produces text for the assistant to answer from; it is not a layer laid over the image. So with a scan you type the question in the chat and use the page reference to get back to the right image, rather than selecting a passage first.

A page came back as if it were empty. What now?

That is the honest failure rather than the dangerous one: where a page cannot be read, you are told the answer is not in the document instead of being handed something invented to cover the gap. Re-scan or re-photograph that page and upload the file again. Either way, check what the answers say against the image on the page before you revise from it.

Put a scan in and see what comes back

Upload the photocopy or the photographed pages and ask what a page covers before you rely on it.