Text trapped in an image is useless until you retype it: a supplier price list sent as a screenshot, a product label, a scanned invoice, a slide from a webinar. This converter reads the text for you with optical character recognition (OCR), so you can copy it, edit it and paste it where it belongs.
It runs Tesseract, an open source OCR engine, inside your browser. The image is processed on your own device and is never uploaded, so it is safe for invoices and internal documents.
How to extract text from an image
- Upload a JPG, PNG or WebP file, drag it onto the box, or paste a screenshot straight from your clipboard with Ctrl+V or Cmd+V.
- Choose the language of the text. English is the default.
- Start the conversion. The first time you use a language, its data file downloads once, so the first run takes longer.
- Read the result and fix any errors. OCR is rarely perfect on the first pass, especially with numbers and punctuation.
- Copy the text or download it as a TXT file.
Supported languages
| Language | Script | Notes |
|---|---|---|
| English | Latin | The default and usually the most accurate |
| Spanish, French, German, Portuguese, Italian | Latin with accents | Choose the right language so accented letters such as รฑ, รฉ, รผ and รง are read correctly |
| Urdu | Arabic script, right to left | Plain printed text works best; decorative Nastaliq on posters reads less reliably |
| Arabic | Arabic script, right to left | Diacritics and very small text lower accuracy |
| Hindi | Devanagari | Clear print gives the best results |
Pick the language that matches the text. If you run a Spanish page as English, the engine will still find words, but accented letters and some words will come out wrong. For a page that mixes languages, run it with the main language and correct the rest by hand.
Tips for better accuracy
Most OCR errors come from the image, not the engine. Tesseract's own documentation recommends at least 300 dpi for scans, straight text lines and dark text on a light background. In practice:
- Use enough resolution. Small text in a low-resolution screenshot is the most common cause of errors. Zoom in on the page before taking the screenshot, or scan at 300 dpi.
- Crop to the text. Remove toolbars, photos and logos around it. A small margin is fine, but large unrelated areas can turn into stray characters.
- Keep it straight. Photograph pages flat and square on, not at an angle. Tilted lines are harder to separate.
- Get good contrast. Dark text on a light background works best. White text on a dark banner, or text over a photo, often fails. Even lighting without shadows helps with phone photos.
- Convert unusual formats first. iPhone photos saved as HEIC, and other formats, can be turned into JPG or PNG with the image converter.
Tip: For a scanned PDF, convert the pages to images with the PDF to JPG converter, then run each page here.
What OCR cannot read well
| Content | Result | What to do |
|---|---|---|
| Handwriting | Poor. Tesseract is built for printed text | Type it by hand |
| Script, decorative or very bold display fonts | Letters misread or skipped | Crop to plain body text where possible |
| Text over photos or patterns | Missing words and random characters | Crop tighter or find a cleaner source |
| Multi-column layouts and tables | Text is found, but the order or columns may mix | Crop one column or table section at a time |
| Curved or rotated text | Often unreadable | Rotate the image or retake the photo |
| Math symbols and special characters | Often misread | Check and correct by hand |
Always check numbers carefully. A 0 read as O, a 1 read as l, or a missing decimal point in a price or SKU is the kind of error that costs money later.
Privacy: what happens to your image
The OCR engine runs in your browser tab. Your image is read on your device and is not sent to our server or saved. The only download is the language data file, which your browser fetches the first time you use a language and stores so later runs start faster. If you clear your browser's site data, it downloads again the next time. This makes the tool suitable for receipts, invoices and supplier documents you would not want to upload to an unknown service.
Uses for store owners and marketers
- Supplier price lists and spec sheets sent as screenshots or photos, turned into text you can paste into a spreadsheet.
- Product labels and packaging for ingredients, materials, care instructions and barcodes printed as text.
- Receipts and invoices for bookkeeping, when you need the numbers without retyping.
- Slides, infographics and social posts you want to quote or summarize, with credit to the source.
After extraction, the word counter helps you trim text for product descriptions. If you are copying product details from another online store rather than from images, there is a faster route: AM Jarvis Product Importer copies products from public Shopify and WooCommerce stores into your own store with variants, images and prices, and lets you edit and rewrite the text before import.