"Text encoding Unicode (UTF-8) isn't applicable" on a Mac
TextEdit shows “Text encoding Unicode (UTF-8) isn’t applicable” when it tries to read a file as modern Unicode text and finds bytes that cannot be that. There are only two explanations: the file is text saved in an older encoding, often by a Windows program, or it is not a text file at all and only looks like one from its name. One Terminal command tells you which, and each case has a quick fix.
What the message is telling you
A text file is a run of numbers, and an encoding is the table that turns those numbers into letters. Plain English letters and digits use the same numbers in nearly every table, so a file containing nothing else opens anywhere. The tables disagree about everything beyond that: accented letters, curly quotes, currency signs and every non-Latin script.
UTF-8 is the modern standard and the one a Mac assumes. It also has strict rules, so when a file breaks them TextEdit knows for certain it is not UTF-8, and it stops instead of showing you the wrong characters. The rest of the message says as much: the file may have been saved using a different text encoding, or it may not be a text file.
Find out what the file really is
Open Terminal, type file -I and a space, drag the file into the window, and press Return:
file -I ~/Downloads/notes.txt
The reply is one line. Read it like this:
- text/plain; charset=iso-8859-1 or charset=unknown-8bit: real text in an older encoding. For a file that came from a Windows PC, this nearly always means Windows Latin 1.
- charset=utf-16le or utf-16be: Unicode text in a two-byte form, common in exports from Windows programs and spreadsheets.
- application/msword, application/pdf, application/zip, image/jpeg and the like: not text. The file has the wrong extension.
- application/octet-stream; charset=binary: binary data the command does not recognize. It is not text, and the name is your only clue.
file makes an educated guess from the start of the file, so treat the charset as a strong hint, not a certificate.
If it is text, open it with the right encoding
In TextEdit. Choose File, Open, select the file, and click Show Options at the bottom of the dialog if the encoding menu is not visible. Set Plain Text Encoding to a likely candidate and open the file:
- Western (Windows Latin 1) for most files from a Windows PC in English or a Western European language.
- Western (Mac OS Roman) for files written on a Mac many years ago.
- Unicode (UTF-16) when
filereported utf-16.
For other languages the menu has entries for Cyrillic, Japanese, Chinese, Korean, Greek and more; choose Customize Encodings List at the bottom of the menu to show the ones that are hidden. You have the right one when accented letters and quotation marks read correctly. A wrong guess harms nothing as long as you close the window without saving.
To make the fix permanent, hold Option, choose File, Save As, and set Plain Text Encoding to Unicode (UTF-8). Save under a new name so the original stays exactly as it was.
While you are there, check TextEdit, Settings (Preferences on macOS 12), Open and Save. Under Plain Text File Encoding, Opening files should say Automatic. If it has been set to Unicode (UTF-8), TextEdit refuses every file that is anything else.
In Terminal. iconv converts in one line and writes a new file:
iconv -f WINDOWS-1252 -t UTF-8 notes.txt > notes-utf8.txt
Swap in UTF-16 or MACROMAN as the source when that is what you have, and run iconv -l to list every name it accepts. Afterward, open the result and read a few lines that contain accents. iconv cannot know whether you named the right source, and with the older single-byte encodings it converts without complaint either way.
Spreadsheet data. A CSV with the same problem opens in Numbers with scrambled accents instead of an error. Numbers briefly offers an Adjust Settings button after opening a CSV, and the import settings include the text encoding.
If it is not text at all
When file reports a document, image or archive type, forcing the file open as text only produces a wall of symbols. The fix is the name, not the encoding.
- In Finder, select the file and choose File, Duplicate.
- Give the copy the extension that matches what
filereported: .doc for application/msword, .pdf for application/pdf, .zip for application/zip. - Confirm the change when Finder asks, then open the copy normally.
A modern Word document is a zip container inside, so file may call a .docx a zip or give its long Word type name. If it reached you as a Word document, .docx is the extension to try.
Changed a file extension and now it will not open and a file with no extension that will not open go further into matching names to contents.
When it turns out to be a damaged document
Sometimes the file has the right extension and still opens as garbage, both in TextEdit and in the app it belongs to. That is damage, not encoding.
Manatee has a Repair mode for corrupt Word documents, PDFs and video. Drop the file onto its menu bar icon or into its window, or right-click it in Finder and choose Services, Fix with Manatee. It reports one of three outcomes: Fixed, Partial (for example 34 of 41 pages), or Not fixable. It always writes a new copy beside the file and never modifies the original. Its Convert mode also turns an old .doc, an RTF or an HTML file into DOCX, RTF or TXT once the file opens. Everything runs on the Mac with nothing uploaded, which matters when the file is someone’s private notes. Every feature is free for 24 hours, then it is $9 once.
The encoding step is still yours. Do not count on any converter, Manatee included, to guess the encoding of an old plain text file for you: settle it with TextEdit or iconv first, then convert the UTF-8 copy. And data missing from a file that was cut short cannot be brought back by any tool. A Word document that opens full of strange symbols covers the Word side of this in more depth.
Questions
Why did the file open without trouble on the PC that made it? Older Windows programs save text in the regional encoding of that PC and read it back the same way, so nothing looks wrong there. The mismatch only appears on a system that assumes UTF-8.
I see é where é should be. Is that the same problem? It is the mirror image: a UTF-8 file being read as Windows Latin 1. Reopen it through File, Open with Plain Text Encoding set to Unicode (UTF-8).
Does this affect subtitle files too? Yes. An SRT is plain text, and one saved in an older encoding shows the same scrambled accents in a video player. Subtitles that show strange characters on a Mac walks through that case.