Fix Garbled Text by Repairing Unconvertible Characters First

This guide covers a narrow form of “garbled text”: a source that opens readably but contains characters that cannot be represented in the required output encoding. It does not recover a file already decoded into nonsense, reconstruct damaged bytes, or guess the original text. If most of the document is unreadable, check BOM and open-time encoding before changing characters.

The controlled Rune Studio example used a readable 1,800-character copy. Five conversion candidates were isolated, only approved substitutes were applied, and a separately named UTF-8 file reopened with 1,800 characters.

A warning is a location, not an editorial answer

An unconvertible character may be a private-use name glyph, an emoji in a quotation, or a rare place-name form. Replacing all five with the same placeholder removes the warning while damaging meaning.

Use a decision sheet:

Location Candidate Context Approver Substitute State
line 4 private glyph A author name author pending hold
line 12 emoji quoted post editor agreed word approved
line 21 rare glyph B place name copy editor common form approved

The surrounding sentence and approver are as important as the candidate itself.

Repair the copy in five deliberate moves

Stop if the source is not readable

Copy source/legacy-manuscript.txt to work/character-repair.txt. Record the source encoding and 1,800-character count. If opening the file produces widespread mojibake, do not continue with substitution. Return to encoding detection.

Capture all five contexts

For every candidate, store its line or character position and about 20 characters on each side. Mark names, quotations, and specialist terminology. A count of five without locations cannot be audited after editing.

Obtain approval before changing text

Ask the author about name forms, the editor about quoted emoji, and the subject owner about technical terms. Rune Studio can surface conversion problems; it does not know the intended semantics.

Save to a new UTF-8 destination

Apply approved substitutions to the copy and save it as out/manuscript-approved-utf8.txt. Keep the received file unchanged. Add the destination path to the candidate sheet.

Reopen and revisit each location

Close the output and reopen it from out/. The controlled result reopened at 1,800 characters after the five approved decisions. Visit each recorded context again and confirm that meaning, not merely character count, survived.

Define “fixed” with six fields

The repair passes when every candidate has a location, context, approver, substitute, destination, and reopened result. Equal character count is supporting information, not a semantic guarantee. A justified spelling note may change the count and still be correct.

Rune Studio’s published feature set describes detection of major Japanese encodings, surfacing unconvertible characters, and pausing autosave while problems remain. The observed result here supports candidate isolation and the separate UTF-8 reopen only.

It does not support recovery of already damaged data, automatic correction of a wrong decoding, OCR repair, binary repair, or automatic choice of replacement meaning.

Choose the right starting point

If the manuscript is readable and only a handful of characters fail the target encoding, begin with the candidate sheet. If the whole document is garbled, begin with BOM and open-time encoding instead.

For the tested 1,800-character case, five approved substitutions led to a separately named UTF-8 copy that reopened at the original count. Readers who need Japanese encoding warnings can review Rune Studio’s current product details before working on a client copy.