
Converting a Shift_JIS or other non-UTF-8 manuscript to UTF-8 has three conceptual stages: decode the source bytes correctly, encode the decoded text as UTF-8 in a separate file, and compare the complete decoded text before and after. These are general requirements for an encoding conversion, not instructions for a particular interface.
This article explains what each stage must establish. The rune Studio tests completed for it cover reported encoding detection and a refused save in the reverse direction, from UTF-8 to Shift_JIS. They do not include an end-to-end Shift_JIS-to-UTF-8 save, reopen, and whole-text match.
Conversion is a read paired with a write
Changing an encoding means reading with one encoding and writing back with another. Reading and writing are separate processes.
The order matters. Write while the read was wrong and the wrong characters go into the new file. Which is why the check at the "open" step cannot be skipped.
What to require from a bulk conversion tool
When selecting a tool to convert a folder, check whether it can report the source encoding for each file, hold the destination at UTF-8, preserve the originals, and identify failures. Processing many files does not by itself establish that every source was decoded correctly.
Whether rune Studio has a bulk conversion feature was not tested for this article.
What happens when the three stages are confused
If decoding is wrong, the later UTF-8 output receives the wrong characters. If that wrongly decoded text overwrites the sole source and no backup remains, the original bytes are no longer available for another decoding attempt. Keeping the source or a backup preserves that recovery path.
If the input and output encodings remain the same, no conversion to UTF-8 has been established. A changed file timestamp and an output encoded as UTF-8 are different facts.
Without a check after writing, content preservation is unknown. A file being created and its complete decoded text matching the source are also different facts.
Stage one: decode the source correctly
rune Studio detects the encoding on read and reports it. I verified this by driving the development build from the command line. A file saved as Shift_JIS reported Shift_JIS; a UTF-8 file reported UTF-8; each came with the line-ending style, character count, and line count.
The source-decoding stage has three points to inspect.
- Does the reported encoding match what you believed about this file?
- Does the prose read correctly?
- Are the symbols intact?
The third is the one people skip. Prose can read correctly while only symbols differ — a wave dash written into a Shift_JIS file came back as a full-width tilde in my test.
Stage two: encode the decoded text as UTF-8
The second requirement is to write the correctly decoded text as UTF-8 bytes. Guessing the source encoding and choosing the destination encoding are different decisions. A suitable tool must let you specify UTF-8 as the output and retain a result separate from the source.
This is a requirement for selecting a tool. It is not a claim that rune Studio completed a Shift_JIS-to-UTF-8 save for this article.
Stage three: verify the result separately
A general verification condition is that the result can be decoded as UTF-8 and compared in full with the source decoded under its original encoding. A file being created is not evidence that the complete text matches.
The rune Studio testing for this article did not complete this Shift_JIS-to-UTF-8 save, reopen, and whole-text match. This stage is therefore not presented as a completed product procedure or result.
A refusal example in the other direction
In a development-build test, a UTF-8 manuscript containing three Japanese name variants was asked to save as Shift_JIS. It produced no file, and a report that one character could not be represented.
This establishes that, under the tested condition, rune Studio reported the position and character and did not write the Shift_JIS file.
This measurement is the reverse direction, from UTF-8 to Shift_JIS. It does not establish completion of a Shift_JIS-to-UTF-8 conversion.
Keep the originals and compare the whole text
When several manuscripts are involved, retaining each original and a one-to-one mapping to each result is a prerequisite for comparison.
Content preservation calls for a whole-text comparison, not the count of one phrase. Decode the source as Shift_JIS and the result as UTF-8, then compare the complete text. A match across the complete decoded text is the relevant condition.
A workspace-wide search can still help you locate the files or passages that need attention. It cannot prove the conversion by itself: equal counts can hide a deletion in one place and an addition in another, and text outside the query can change.
Who this suits, and who is fine without it
This three-stage model suits anyone holding several manuscripts in mixed encodings — collaboration, reused old material, files received from outside. It keeps source decoding, UTF-8 output, and content comparison as separate decisions.
If you have written only in UTF-8 from the start, no conversion is needed.
Scope note: what was verified is that detection is reported, that the reverse-direction Shift_JIS save was refused, and that workspace-wide search works. A bulk conversion feature and the complete Shift_JIS-to-UTF-8 save, reopen, and whole-text match were not tested here.
Summary
- UTF-8 conversion has three stages: decode the source correctly, encode the decoded text as UTF-8, and compare the complete decoded text
- Source decoding, UTF-8 output, and whole-text equality are separate conditions
- Choose a tool that can preserve the source and keep each result paired with it
- rune Studio testing established encoding detection and a reverse-direction refusal, not a completed Shift_JIS-to-UTF-8 result
When selecting a tool, establish whether it can report the source encoding, write a separate UTF-8 result, reopen that result as UTF-8, and compare the complete decoded text. Those four capabilities are what allow the three stages in this article to be verified end to end.
See the current product scope on the Rune Studio product page.