Text copied from a PDF pastes as gibberish on a Mac: why and the fix
· 6 min read
When text copied from a PDF pastes as gibberish, the PDF is missing the table that says which character each drawn shape stands for. The page looks right because the shapes are right; the copy is wrong because the viewer had to guess the letters. You cannot repair that table from outside the file, so the dependable fix is to ignore the broken text and read the page as a picture, with OCR.
Three other things also stop copying (a scan, text turned into outlines, a copy restriction), and each has a different fix. Work out which one you have first.
What a PDF actually stores
A PDF does not hold sentences. It holds drawing instructions: place glyph number 37 from this font here, then glyph 12 beside it. To turn glyphs back into letters for copying and searching, a viewer needs a table inside the file that maps each glyph number to a character.
When the program that made the PDF embeds a cut-down font with its own numbering and leaves that table out, the viewer guesses. The result is boxes, symbols, or real letters that are simply the wrong ones. Some publishers scramble the table on purpose to discourage copying. Either way the fault is in the file, so it follows the PDF to every Mac.
Work out which problem you have
Open the PDF in Preview and try three things.
- Drag across a line. If individual words highlight, the file has a text layer. If nothing highlights, the page is a picture of text, or the text was converted to shapes.
- Search for a word you can see. Press Command-F and type it. No matches for a word that is plainly on the page means the text layer is missing or wrong.
- Paste into a plain document. Open TextEdit, choose Format, Make Plain Text, and paste. That shows what was really copied.
Then read the result:
- Highlights fine, pastes as nonsense, search finds nothing: a broken character mapping. This is the gibberish case.
- Highlights fine, pastes readable text with odd line breaks: the text is intact and only needs tidying, covered further down.
- Nothing highlights: a scan or outlined text. macOS 12 Monterey and later can select words in a scanned page with Live Text, but that is recognition done on the fly and nothing is saved in the file. Make a scanned PDF searchable on a Mac covers the permanent fix.
- Highlights, but Copy is dimmed in the Edit menu: the file carries a permission restriction. Choose Tools, Show Inspector and click the lock tab to see what is allowed. If the document is yours and you have its password, see Remove the password from a PDF you own.
Quick fix for a paragraph or two
For a short passage, let macOS read the pixels instead of the text layer.
- Zoom the PDF in until the passage fills most of the window. Larger letters are recognized more accurately.
- Press Shift-Command-Control-4 and drag over the passage. Adding Control sends the screenshot to the clipboard instead of saving a file.
- In Preview, choose File, New from Clipboard.
- Move the pointer over the words in the new image. It becomes a text cursor: that is Live Text. Drag to select, then copy.
It does not scale past a page or two, and the PDF stays as broken as it was.
Fix for a whole document
For a long PDF, rebuild it: turn every page into an image, then run OCR over the images so the new file gets a correct text layer.
Turn the pages into images. Preview exports only one page at a time to an image format, so use Automator for anything longer:
- Open Automator and create a new Workflow.
- Add Ask for Finder Items, then Render PDF Pages as Images. Set the format to PNG and the resolution to 300 dots per inch.
- Add Move Finder Items and choose an empty folder as the destination.
- Click Run and pick the PDF. One image per page lands in the folder.
Recognize the text. Finder can combine the images into a PDF (select them, right-click, Quick Actions, Create PDF), and Live Text then lets you select words page by page. It writes no text layer into the file, so searching the whole document still falls short.
For a saved text layer you need OCR that writes its result into the PDF. Basalt has both halves as tools in one window: Images to PDF builds the document from the page images, and Make Searchable runs OCR and writes a new PDF with real text in it. Everything happens on the Mac with no upload, which matters when the document is a contract or a statement. Every tool writes a new file, so your original stays as it was. Basalt is $29 once for up to three Macs, with 24 free hours to try every tool.
Starting from images is deliberate: running OCR straight over the original is less certain with any software, because the file already claims to have text.
Two honest limits. OCR makes its own mistakes, so check figures and names. And a rebuilt PDF is larger than the original, because every page is now a picture.
When the text is right but messy
A line break at the end of every line. PDFs store lines, not paragraphs. Copy a paragraph, then run this in Terminal to join the lines on the clipboard:
pbpaste | tr '\n' ' ' | pbcopy
It turns every line break into a space, so do one paragraph at a time. The other common faults:
- Words split by hyphens. The hyphen at a line end is a real character in the file. After joining lines, search for a hyphen followed by a space and fix each by eye.
- Two columns pasted as one jumble. Hold Option while dragging in Preview to select a rectangular block, one column at a time.
- Missing “fi”, “fl” or “ffi”. The PDF drew those pairs as a single joined glyph and the viewer could not translate it. Find and replace is the cure once you spot the pattern.
Ask for a better file
If the PDF came from a colleague or a client, the cleanest fix is upstream. Ask for the original document, or for a fresh export made directly from the program it was written in. A PDF printed from one viewer into another is a common source of broken mappings.
If you produce PDFs yourself, copy a sentence out of your own export before you send it. Search and screen readers depend on the same table.
Questions
Why does the same PDF copy correctly in another viewer? Viewers guess differently when the mapping is missing. Preview and Safari share Apple’s PDF engine, so they agree with each other; a browser with its own engine may guess better or worse. If another viewer gets it right, copy from there.
Will printing to PDF fix the gibberish? No. Print, then Save as PDF, carries the same glyphs and the same missing table into the new file. Only recognition from the page image produces correct text.
Can I search a PDF like this at all? Not until it has a correct text layer. Spotlight and Preview both search the same broken text you are copying. Once the file is rebuilt with OCR it behaves like any other document, and Search for a word across many PDFs at once covers searching a folder of them.