Detect UTF-8 BOM Presence in a Text Editor Without Guessing

An abstract editorial 3D still life illustrating Detect UTF-8 BOM Presence in a Text Editor Without Guessing

The decision in this article is not the generic question of how to recognize a BOM. It is whether the leading-byte evidence and the text editor’s detection can be accepted as one consistent result for the same file. If the prefix is EF BB BF and the detector reports UTF-8 (BOM), accept BOM present. If that prefix is absent and the detector reports UTF-8, accept BOM absent. If the two observations disagree, leave the state unresolved and return to the pre-save copy.

In the tested 40-line pair, the no-BOM file began with decimal bytes 48, 49, and 232—hexadecimal 30 31 E8—and was detected as UTF-8. The BOM file began with 239, 187, and 191—EF BB BF—and was detected as UTF-8 (BOM). Both contained 520 characters. Because the two observations agreed for each file, this bounded result is passed. It does not decide which BOM policy a receiver requires or whether an external service accepts either file.

The central question is whether the evidence agrees

Identical visible text and identical character counts cannot decide BOM state. This workflow treats the leading bytes as one observation and the editor’s detection as a second observation.

Leading bytes Editor detection Decision here
EF BB BF UTF-8 (BOM) Accept BOM present
Not EF BB BF UTF-8 Accept BOM absent
EF BB BF UTF-8 or another label Conflict; do not save
Not EF BB BF UTF-8 (BOM) Conflict; do not save

The table is not a rule for preferring one tool over another. It creates a stop condition. A conflict must not be rounded into “probably BOM” or “the text looks readable, so it is fine.”

The 40-line, 520-character pair produced consistent evidence

The controlled material contained the same Japanese text in two files, one saved as UTF-8 without a BOM and one with a BOM.

Check No BOM With BOM
Logical lines 40 40
First three bytes, decimal 48, 49, 232 239, 187, 191
First three bytes, hexadecimal 30 31 E8 EF BB BF
Detected encoding UTF-8 UTF-8 (BOM)
Character count 520 520

For the BOM file, the marker and detector agreed. For the no-BOM file, the marker was absent and the detector returned UTF-8. The no-BOM prefix 30 31 E8 is not a universal signature; it belongs to this specimen’s opening text.

The matching 520-character totals are supporting evidence that the marker was not counted as manuscript text in this specimen. A count of 520 cannot identify BOM state by itself.

Stop before saving when observations conflict

Use the following order to preserve the identity of the file being judged:

  1. Record the filename and an identifier such as SHA-256, then inspect only that copy.
  2. Read its first three bytes and record whether they equal EF BB BF.
  3. Open the same copy and record the editor’s detected encoding in a separate field.
  4. Accept BOM present or absent only when the pair matches the decision table.
  5. If the pair conflicts, do not save. Reinspect the unchanged copy with another byte-level method.

This article deliberately stops before rewriting. STUDIO-536 owns the save workflow that selects BOM and line endings. STUDIO-171 owns the general two-specimen identification and copy-conversion exercise.

Rune Studio supplies the second observation

Current documentation for the Mac version of Rune Studio states that it supports UTF-8 with and without a BOM and checks for a BOM before continuing through automatic encoding detection. This article uses that detected label as the second observation.

The supplied documentation does not establish a Rune Studio interface that exposes raw leading bytes. The byte prefix therefore remains an external byte-level observation, while UTF-8 and UTF-8 (BOM) remain product detection results. The two claims must not be merged into an invented raw-byte display.

A separate common Stage 4 run inspected one short text in five encodings. UTF-8, Shift_JIS, and EUC-JP each returned 39 characters; UTF-16LE and UTF-16BE each returned 40 characters, and the working tab read back UTF-8. That run supports broader encoding inspection but does not replace this article’s 40-line pair.

The role differs from STUDIO-171

STUDIO-171 explains the general identification exercise: prepare BOM and no-BOM specimens, compare visible text, inspect leading bytes, read the editor label, and test conversion on copies.

This article instead asks whether a measured file has internally consistent evidence. Its action is to accept agreement or stop on conflict. The reader does not finish by adding or removing a BOM; the reader finishes with a recorded decision or a preserved unresolved copy.

Save policy and receiver acceptance remain separate

Neither BOM state is universally correct. Follow the receiver’s documented requirement. If no rule is available, preserve the current state while keeping the detection record.

External-service acceptance, complete recovery of damaged data, OCR, and binary repair were not tested. The run also did not display deliberately mis-decoded text or overwrite a file from that state.

Conclusion: accept only a matching evidence pair

For a UTF-8 BOM decision, pair the leading bytes with the detector result from the same file. Accept EF BB BF plus UTF-8 (BOM), or a non-BOM prefix plus UTF-8. If the observations conflict, do not choose by majority and do not save; return to the unchanged copy and inspection method.

In this 40-line specimen, 30 31 E8 plus UTF-8 and EF BB BF plus UTF-8 (BOM) were consistent, and both files contained 520 characters. The passed status covers that consistency decision only. BOM policy, rewriting, and downstream acceptance require separate evidence. The current Mac feature scope is available on the Rune Studio product page.