Convert a scanned PDF to EPUB

Last updated 10 October 2026

A scanned PDF holds a picture of each page, not text. To make an EPUB from it, a converter first has to read the words off each picture with optical character recognition (OCR). Finepress runs OCR in your browser, on your device, and then builds a reflowable EPUB from the text. Your PDF is not uploaded.

How do I convert a scanned PDF to EPUB?

Open Finepress, add the PDF, and wait while it reads the scanned pages. The book then sits on your shelf, where you can read it or download the EPUB.

  1. Open Finepress. You do not need an account or an install.
  2. Press Add a book and choose the PDF, or drop the file anywhere on the page.
  3. Watch the progress line. A scanned page shows as "Reading scanned page 3 of 120". The first scan you convert also shows "Fetching the OCR engine" while the engine downloads, once.
  4. If the scan is long, Finepress times the first page and tells you how many minutes are left. Stop is on screen the whole time.
  5. When it finishes, press Open to read the book. To get the file, open the book's details on the shelf and press Download EPUB.

Which languages can the OCR read?

Printed English and simplified Chinese. The engine loads both languages, so you do not pick one before you start.

The OCR engine is Tesseract, and with its language data it is about 10.6 MB. Finepress downloads it the first time it meets a scanned page and keeps it in the browser's cache, so later scans do not download it again and can convert offline. On clean synthetic test scans, Finepress reads English at 100% and simplified Chinese at 97% character accuracy. Real scans, with uneven light, skew or show-through from the other side of the paper, do worse.

Which pages go through OCR?

Only pages that have almost no text and are mostly covered by an image. Before it starts, Finepress checks up to ten pages spread through the PDF. If more than half of them already carry text, it uses the PDF's own text for the whole book and runs no OCR.

What if my scan already has a text layer?

Finepress uses that text as it is. Many scanners and PDF tools run their own OCR and hide the text behind the page pictures. Finepress reads that hidden text the way it reads any text PDF, and does not run OCR again.

If the hidden text has mistakes, the EPUB has the same mistakes. Finepress has no setting to ignore a text layer and read the pictures instead.

How do I check the result?

Read the summary Finepress shows when it finishes, then look through the book.

To read the result on an e-reader, see Read a PDF on your Kindle or Read a PDF on a Kobo.

Common questions

Can Calibre convert a scanned PDF to EPUB?

Not on its own. Calibre's manual says image-based documents are not supported, and that for a scan with OCR text behind the page pictures, Calibre uses that text. A scan with no text layer needs OCR from another tool first. See Finepress and Calibre for PDF to EPUB.

How long does OCR take per page?

It depends on the device and the scan. On clean synthetic test scans it takes about half a second a page on an Apple Silicon laptop. Real scans have taken from about a tenth of a second to 18 seconds a page. For a long scan, Finepress shows its own estimate after the first page.

Can it read handwriting?

No. The OCR is meant for printed text, so handwritten pages will not come out correctly.

Can it read scans in other languages?

No. OCR reads English and simplified Chinese only, so a scan in another language, traditional Chinese included, comes out wrong. A PDF in another language that already has a text layer does not need OCR.

Is my PDF uploaded?

No. The PDF and its text are read and converted in your browser. The guide to converting without uploading lists what does leave your device.

Finepress is free to use and needs no account. Your PDF is converted in your browser.

Open Finepress and add your PDF