
When converting a Shift_JIS manuscript to UTF-8, resolve unrepresentable characters first, then reopen a separately named UTF-8 copy and reconcile the complete text, character count, and line endings. This specimen was requested as 1,400 characters but measured 1,399 with CRLF endings. 😀 at index 16 and 🚀 at index 17 were recorded as unrepresentable; approved substitutions reduced the count to zero, and the reopened UTF-8 and Shift_JIS copies matched. If a check fails, return to the zero-finding repaired working copy rather than the source or a partially accepted output.
The article-specific status is partial. The two repairs and matching reopened text were verified, but the starting count was short by one. The anticipated private-use candidate was not identified in the result, and the CLI did not verify red highlighting, the pre-save warning, or autosave stop and resume.
Treat UTF-8 conversion as three separate gates
Changing an encoding selector to UTF-8 is not the entire conversion:
- The Shift_JIS source must be decoded correctly.
- Every recorded unrepresentable character must receive an approved decision, with zero remaining findings.
- A separately named UTF-8 output must reopen with the repaired text, expected count, and intended line ending.
Saving after a bad first gate can preserve the wrong decoded text in a new encoding. Compare with the source before creating the output.
The requested 1,400 characters measured 1,399
The planned fixture was Shift_JIS-derived text with CRLF endings, two emoji, and one private-use candidate. The title-specific evidence reported:
| Field | Measured value |
|---|---|
| Requested characters | 1,400 |
| Observed source characters | 1,399 |
| Unrepresentable before repair | 2 |
| Positions | 😀 at 16; 🚀 at 17 |
| Unrepresentable after approved substitutions | 0 |
| Repaired UTF-8 copy | Reopened as UTF-8 |
| Comparison Shift_JIS copy | Reopened as Shift_JIS |
| Source line ending | CR+LF |
| Repaired text | Matched |
Do not round 1,399 up to the requested value. The available record does not establish the cause of the difference, so 1,399 becomes the observed baseline for this run. The private-use candidate is also absent from result_values and remains unverified.
A product-independent conversion workflow
- Duplicate the Shift_JIS source and record its SHA-256, observed 1,399 characters, CRLF, and representative sentences.
- Decode the copy as Shift_JIS and list unrepresentable characters with positions.
- Reconcile indexes 16 and 17 with the source and apply only approved substitutions.
- Repeat the check and require zero remaining findings.
- Save the zero-finding copy under a separate UTF-8 name; retain a separately named Shift_JIS comparison copy when needed.
- Close and reopen both files, then compare the full text, observed count, CRLF, and representative sentences.
The authoritative source is never overwritten. An unsuccessful output remains disposable.
Why zero findings still produce a partial result
The measured two findings did reach zero, and both encodings reopened with matching repaired text. That is strong evidence for the recorded repair path.
It does not resolve the one-character discrepancy between requested and observed length. Nor does it classify the anticipated private-use candidate. The evidence cannot tell whether that candidate was representable, absent from the specimen, or omitted from the recorded scan.
Separate Rune Studio’s Stage 3 scope from Stage 4 evidence
Current Rune Studio documentation states that the Mac app handles UTF-8, UTF-8 with BOM, Shift_JIS, and other major encodings; highlights characters the selected encoding cannot represent; pauses autosave while such characters remain; warns on manual save; and lets users change encoding and line endings from the status bar. Those are Stage 3 product capabilities.
The title-specific Stage 4 run verified two positions in the 1,399-character specimen, zero after approved substitutions, reopened UTF-8 and Shift_JIS copies, CRLF, and matching text. It did not verify the visual warning or interaction.
A separate common Stage 4 run inspected one short text as UTF-8, Shift_JIS, and EUC-JP with 39 characters and as UTF-16LE and UTF-16BE with 40, all using LF. The working tab read back UTF-8. That supports multi-encoding inspection but is not the conversion evidence for the 1,399-character specimen.
This is an operation guide, not an editor-selection guide
A Mac UTF-8 editor-selection article asks whether a tool exposes source encoding, line endings, counts, and problem positions well enough to audit a migration. This article assumes a method has been chosen and follows the concrete repair through a separately named UTF-8 output.
Price, feature breadth, and comparative product recommendations are outside this decision. The conversion gate is zero recorded findings plus matching reopened text.
Rollback points and open limits
- If the source encoding cannot be established, stop before saving and return to Shift_JIS decoding.
- If a detected position does not match the source, return to the position list without substituting.
- If findings remain after repair, do not create the UTF-8 output.
- If reopened text differs, discard both outputs and return to the zero-finding working copy.
- If the observed 1,399 count changes, identify the added or missing character before saving again.
Visual pre-save warnings, red highlighting, deliberate mis-decoding, autosave stop/resume, and acceptance by an external application or submission system were not tested. Complete recovery of damaged data, OCR, and binary repair are outside scope.
Conclusion: use the observed 1,399-character baseline
This conversion started with a requested 1,400 characters but an observed 1,399, CRLF, and two failures: 😀 at index 16 and 🚀 at index 17. Approved substitutions reduced the count to zero, and the UTF-8 and Shift_JIS copies reopened with matching text. The one-character discrepancy and unrecorded private-use candidate keep the status partial.
The article’s role is to repair recorded failures and verify a separately named UTF-8 copy. Candidate classification, BOM detection, and saving with independently chosen BOM and line-ending policies have different inputs and exit conditions. The current Mac feature scope is available on the Rune Studio product page.