
When you open an old manuscript, you do not choose the encoding — the editor guesses it. That guess is usually right and not always right. So when a file opens garbled, the first thing to look at is not the text but what encoding was detected.
This article covers how detection proceeds and what happens when it misses. Converting to UTF-8, and repairing garbled text, are covered separately.
The encoding is not written in the file
It surprises people: a plain text file contains no statement that it is Shift_JIS. It contains a sequence of numbers and nothing else.
So the reader infers the encoding from the pattern of those numbers — “this sequence is plausible as Shift_JIS”, “it is implausible as UTF-8”.
There is one exception. A small marker at the start of a file, a byte-order mark, identifies the encoding outright. But it is frequently absent, and for Japanese text files its absence is the norm.
Choosing manually instead
The naive response to a miss is reopening with an encoding you choose: try Shift_JIS, UTF-8, EUC-JP in turn and keep whichever reads correctly.
“It reads correctly” is not proof
The problem is that a miss can still read correctly.
How it fails depends on which pair is confused. Read a Shift_JIS file as UTF-8 and almost everything turns into unreadable symbols, which is obvious enough. The awkward case is when detection is correct and only certain symbols arrive as different characters. The prose reads normally, rereading raises no flag — and then you save.
Saving can be the point of no return. Write characters that were displayed under a wrong interpretation back out in another encoding, and the original numbers are gone.
Some files are harder to detect than others: short files, files that are mostly alphanumerics, files dense with symbols. Less evidence means a weaker guess.
rune Studio reports what it detected
rune Studio detects the encoding when reading a file. It looks first for a byte-order mark, then uses system detection, then tries the major Japanese encodings in turn.
The encodings it handles are UTF-8, UTF-8 with a byte-order mark, Shift_JIS, EUC-JP, ISO-2022-JP, both byte orders of UTF-16, and a final fallback.
I verified this by driving the development build from the command line. Reading a file saved as Shift_JIS reported Shift_JIS, along with the line-ending style, character count, and line count. A UTF-8 file reported UTF-8.
A result rather than a guess.
A correct detection is not a guarantee of identity
One practical finding. I wrote a wave dash into a Shift_JIS file, saved it, and read it back — it returned as a full-width tilde, a visually similar but different character.
That is not a detection failure; it is a long-known mapping issue between encodings. But it is worth knowing that “detected correctly” does not mean “identical to what was written”. In symbol-heavy files, compare the symbols after reopening.
The detected encoding is visible on screen
The detected encoding and line-ending style are shown at the bottom of the window, and can be switched. Making a habit of glancing there on open catches problems before they become problems.
Saving halts while a character cannot be represented
Separately from detection, saving halts when the chosen encoding cannot represent a character in the file. I confirmed this: no file was written, and the result reported one unrepresentable character.
To be precise about scope: this refusal to write is what I confirmed when saving from the command line. For saving from the application window, the documentation describes a warning at save time, automatic saving being paused, and an on-screen notice. What actually happens when saving from the window was not verified in this pass. What was verified is the command-line path only.
One of the routes to a silently damaged file is closed.
Who this suits, and who is fine without it
It suits people handling old manuscripts or files received from others. Even a document you started in UTF-8 picks up other encodings through hand-offs.
Stay in UTF-8 throughout and detection rarely enters your awareness. For a manuscript you are starting now, it is hardly ever a problem.
Scope note: what was verified is that the detected encoding is reported, and that one symbol returned as a different character. Detection cannot be claimed to be correct for every file.
Summary
- A text file carries no encoding information, so the reader infers it
- A wrong guess can still read correctly, and saving in that state is not recoverable
- rune Studio checks for a byte-order mark first, then tries the major Japanese encodings, and reports the result
- Even a correct detection can return a symbol as a different character, so compare symbols in symbol-heavy files
For a first step, open a manuscript you are working with and check what encoding is reported. If it is not UTF-8, compare the symbols too.
See the current product scope on the Rune Studio product page.