Check Text Encoding with BOM and Automatic Detection

A text editor’s encoding label answers only one question: how the file was interpreted when it opened. It does not, by itself, tell you whether a BOM exists or which encoding a future Save As operation will use. Keep those as three separate fields—BOM presence, open-time detection, and save-time choice—or a clean-looking document can still leave with the wrong delivery settings.

Rune Studio was used to inspect three copies of the same 48-line Japanese text: UTF-8 without BOM, UTF-8 with BOM, and Shift_JIS. The first two fields were observed. The save-time choice was not exercised, so the procedure below marks it as a check the reader must perform rather than a confirmed result.

A three-column intake sheet

Suppose a translator receives these files:

File BOM in source Detected when opened Chosen when saved
utf8-plain.txt no UTF-8 not tested
utf8-bom.txt yes UTF-8 with BOM not tested
legacy-sjis.txt no Shift_JIS not tested

The second column belongs to the original bytes. The third is the editor’s opening decision. The fourth belongs to a new output file and stays blank until the user chooses an encoding in the save dialog.

Add three content checkpoints as well. For the 48-line sample, they might be line 1 “Chapter One,” a Japanese name on line 24, and the final sentence on line 48. Correct checkpoints show that decoding produced readable text; they do not prove whether a BOM is present.

Inspecting an incoming file without overwriting it

Preserve the received bytes

Copy delivery/legacy-sjis.txt to work/encoding-intake/legacy-sjis.txt. Record the source name, size, and known BOM state before opening it. If the file is evidence of what a client sent, never use it as the conversion destination.

Read the detected encoding

Open each of the three samples separately in Rune Studio and write down the encoding shown for that file. In the observed run, the plain UTF-8 sample was detected as UTF-8, the BOM sample as UTF-8 with BOM, and the Shift_JIS sample as Shift_JIS.

Check the first, middle, and last content checkpoints. A mismatch here suggests a decoding problem; it is not a reason to overwrite the source with whichever encoding happens to display something plausible.

Choose a destination deliberately

If conversion is required, use a new name such as out/chapter-01-utf8.txt. In the Save As step, read the encoding selector and record the value actually chosen. Do not fill this field by copying the open-time label.

This article does not claim a measured Rune Studio save-time result. The safe rule is therefore explicit: the reader must inspect the save choice in the current app and preserve it in the handoff record.

Reopen the output

Close the new file and reopen it from out/. Compare its detected encoding, BOM state, and the same three content checkpoints with the intended output specification. If the client requested UTF-8 without BOM but the reopened file reports UTF-8 with BOM, the prose can look identical while the delivery still fails.

Why “the text looks fine” is insufficient

BOM and encoding are metadata-level facts that ordinary reading does not reveal. Two UTF-8 files may display identical Japanese text while differing at the start of the byte stream. Conversely, an editor may correctly detect a Shift_JIS source, but a later save operation may target UTF-8 because the user deliberately chose it.

Rune Studio’s published feature information describes detection for major Japanese encodings and surfacing characters that cannot be converted. That does not promise recovery of bytes already damaged by an earlier bad conversion, and it does not choose the correct delivery encoding on the user’s behalf.

The observed boundary

The controlled three-file comparison established distinct detection for UTF-8 without BOM, UTF-8 with BOM, and Shift_JIS, while recording BOM presence separately. It did not confirm the save-time selector or a saved output.

For a real job, complete one row end to end: source BOM, open-time detection, three visible checkpoints, chosen save encoding, and reopened result. Readers who need Japanese encoding checks can review Rune Studio’s current supported formats on the product page before converting a client copy.