
Choosing a text encoding is easier when you compare the same specimen instead of memorizing encoding names. Preserve the source bytes, save copies in candidate formats, inspect warnings for characters that cannot be represented, and reopen each copy. The destination’s documented requirement should decide the final format.
Build one specimen that exposes differences
Create a four-line file. Use ASCII on the first line, common Japanese text on the second, characters and an emoji that may not be representable in older encodings on the third, and punctuation around a line break on the fourth. This is a diagnostic copy, not the only manuscript file.
Keep encoding-source.txt untouched. Name test outputs by format, such as encoding-utf8.txt, encoding-sjis.txt, and encoding-eucjp.txt. Record the chosen format, any warning, the affected character, and the result after reopening. A file looking correct before saving does not prove that the target encoding can preserve it.
Separate UTF-8 from UTF-8 with BOM
UTF-8 can represent the Unicode characters used by modern multilingual manuscripts. A UTF-8 BOM file begins with the bytes EF BB BF; a no-BOM file does not. The visible text can be identical, while a receiving program may require one variant.
A BOM is not a universal repair for garbled text. Follow the pipeline specification or the established project convention. If BOM status is the only question, inspect the leading bytes and the editor’s encoding indication rather than changing the text.
Do not treat Japanese legacy encodings as interchangeable
Shift_JIS, EUC-JP, and ISO-2022-JP can all carry Japanese text, but they do not assign the same byte sequences and do not necessarily represent the same set of characters. A legacy system or a business partner may require one specific format. “Japanese encoding” is not precise enough for an exchange requirement.
Never overwrite the only copy while guessing the source encoding. Preserve the original bytes, open a duplicate, and compare every specimen line. If the sender can provide the encoding specification, prefer that evidence over detection alone.
Repeatedly saving misdecoded text can destroy information. When the text is already garbled, stop conversion, preserve the source, and identify the original encoding before attempting repair.
Understand UTF-16 byte order and the Latin-1 fallback
UTF-16LE and UTF-16BE encode Unicode code units in different byte orders. A detector may use a byte order mark or other evidence to distinguish them. Do not choose between them from visible content alone.
Latin-1 represents a limited range of Western characters and is not a suitable target for a Japanese manuscript. A detector may use Latin-1 as a fallback because every byte can be mapped, but a file opening without an error does not mean its text has been interpreted correctly. Compare it with the known source content.
Inspect unrepresentable characters before saving
When the target encoding cannot represent a character, identify every occurrence. The editorial choices include keeping UTF-8, replacing the character with an approved alternative, removing it when appropriate, or negotiating a different delivery format. Silently discarding or substituting the character is not a safe default.
If the editor highlights unrepresentable characters, navigate through all of them and record the decision. Check whether automatic saving pauses while the conflict exists. Otherwise, the application could write an unintended conversion before the writer has chosen a response.
A warning is evidence that a decision is required; it does not determine the correct replacement. A person must judge names, quotations, mathematical symbols, and emoji in context.
Reopen every saved copy
Close and reopen the copy after saving. Verify the encoding indication and compare all four lines with the source. Look for replacement symbols, missing text, altered punctuation, or unexpected control characters. Include line endings in the comparison if the receiving system has a line-ending requirement.
Record both success and exclusions. For example, a Shift_JIS delivery may pass the first two lines but be rejected because the approved manuscript includes an unrepresentable character on the third. That is a useful result, not a reason to force the save.
The documented Rune Studio scope
Current documentation for the Mac version of Rune Studio lists UTF-8, UTF-8 with BOM, Shift_JIS, EUC-JP, ISO-2022-JP, UTF-16LE, and UTF-16BE in its encoding detection and reading scope, with Latin-1 as a fallback. It describes checking for a BOM before other detection.
The documentation also describes marking characters that cannot be represented in the current save encoding in red, warning before saving, and pausing automatic saving while such characters remain. These are documented capability claims, not proof that every character can be converted to every encoding or that damaged text can be reconstructed automatically.

Decide with a delivery matrix
Use columns for destination requirement, original encoding, required characters, candidate format, warning result, and reopened result. If the destination requires UTF-8, there is no benefit in converting to a legacy encoding merely because it can hold a simple sample. If an existing Shift_JIS system must be maintained, test every new character before delivery.
The safe encoding workflow is preserve, detect, compare, save a copy, inspect warnings, reopen, and verify. To review the currently documented encoding and save-warning scope, see the Rune Studio product page.