Open UTF-16LE and UTF-16BE Text Without Guessing

An abstract editorial 3D still life illustrating Open UTF-16LE and UTF-16BE Text Without Guessing

To open UTF-16LE or UTF-16BE safely, combine three clues: the byte-order mark, the editor’s detected encoding, and a sentence whose correct form you already know. LE and BE store the bytes of each code unit in opposite orders. Guessing from the filename or from a partly readable screen is not enough.

What little-endian and big-endian mean

UTF-16 commonly represents text in two-byte code units. Little-endian stores the lower-order byte first; big-endian stores the higher-order byte first. A byte-order mark can identify the order: FF FE suggests LE and FE FF suggests BE.

The mark is evidence, not ordinary manuscript text. If a comparison counts it as a visible character, the first-position result may appear different even when the body is correct. Separate the mark from the text before comparing lengths or first characters.

Build a forty-line reference

Use the same known Japanese passage saved as one LE file and one BE file. Place recognizable names or punctuation near the beginning, middle, and end. A candidate passes only when all of those anchors display correctly.

Representative checks detected UTF-16LE and UTF-16BE separately. After the byte-order mark was excluded from the text comparison, the known passages aligned. That is evidence for the sample route, not a promise that every truncated file or unusual extension can be recovered.

Open candidates without writing them

In Rune Studio, open a duplicate and read the tab’s encoding indicator. If detection is uncertain, try LE and BE interpretations and compare the same anchors each time. Do not save while testing. A screen full of NUL-like gaps usually means the file is being decoded incorrectly, not that the gaps should be deleted.

When one interpretation matches every anchor, record it. If your workflow needs UTF-8, save a separately named UTF-8 copy. Close and reopen that result, then check the anchors, line count, and final sentence. Keep the original UTF-16 file unchanged.

Stop when the evidence does not converge

If neither LE nor BE restores the known passage, the byte-order mark may be missing, the file may be truncated, or several encodings may have been combined. Switching settings repeatedly does not create new evidence. Return to another source copy or use a byte-level recovery tool.

Rune Studio is useful for normal LE/BE identification and a controlled move to UTF-8. It is not a repair guarantee for arbitrary corrupted byte streams. Review current encoding support on the Rune Studio product page, and test the full reopen cycle on a short duplicate before converting an archive.