Prevent EPUB Save Failures Caused by Unsupported Characters

An abstract editorial 3D still life illustrating Prevent EPUB Save Failures Caused by Unsupported Characters

To prevent unsupported characters from damaging an EPUB source workflow, do not replace every character that merely looks unusual. List the characters that the selected encoding cannot represent, record their positions, approve each replacement, and save the result under a new name.

The controlled sample was intended to contain two private-use characters, one emoji, and two variant-character candidates in 1,000 characters. The run found only one Shift_JIS-unrepresentable character: 😀 at index 18. After the approved substitution, the count fell to zero, and the UTF-8 and Shift_JIS copies reopened with matching repaired text. Because the other four candidates were not classified in the recorded result, the status remains partial.

Separate candidate labels from encoding failures

“Private-use,” “variant,” and “unsupported” are not interchangeable labels. A character becomes an encoding failure only in relation to the encoding selected for the save. The safe general method is to classify each candidate as keep, replace, or out of scope, without bulk-replacing characters simply because they look uncommon.

This article focuses on classification and approval. The zero-warning conversion gate belongs to the nearby missing-characters workflow.

Freeze the 1,000-character sample and its positions

Keep the source untouched and work on a named copy. Record the source encoding, line endings, character count, candidate list, and detected positions before making any change. If the detector returns fewer candidates than expected, preserve that discrepancy instead of altering the sample to produce a preferred result.

The intended sample contained five candidates. The recorded operation identified one actual Shift_JIS failure. The remaining candidates therefore stay unverified.

Replace only the approved character

The run substituted the detected emoji at index 18, then wrote separately named UTF-8 and Shift_JIS copies. Both files were reopened and compared. The purpose was not to “remove all gaiji,” but to preserve a one-to-one link between a detected position and an approved change.

If the replacement is not approved, stop before saving. If the reopened copies differ, discard the copies and return to the point immediately before the substitution.

The measured result: one finding, zero remaining, matching reopen

The title-specific Stage 4 record contains three commands. The requested and observed lengths were both 1,000 characters. It found one unrepresentable character, 😀, at index 18. The post-repair count was zero. The output encodings read back as UTF-8 and Shift_JIS, the source line ending was LF, and repaired_text_matches=true.

A separate common Stage 4 run inspected the same short text in five files: UTF-8, Shift_JIS, and EUC-JP each read as 39 characters with LF; UTF-16LE and UTF-16BE each read as 40 characters with LF. The working tab read back UTF-8. That common run does not classify the five article-specific candidates.

Why the result is still partial

The CLI did not verify Rune Studio’s pre-save warning, red highlighting, visual mis-decoding state, or autosave stop and resume. The record also does not show how the two private-use and two variant candidates were represented or classified.

Complete recovery of already damaged data, OCR, and binary repair are outside scope.

What Rune Studio contributes

Current Rune Studio documentation says the Mac app detects several Japanese encodings, highlights Unicode characters that the selected encoding cannot represent, and pauses autosave while those errors remain. That is the documented Stage 3 product scope.

This Stage 4 run verified positions, counts, output encodings, and reopened text through the CLI. It did not verify the warning or highlighting visually.

Conclusion: approve the detected character, not the label

The useful result is narrow and reproducible: one emoji at index 18 was unrepresentable in Shift_JIS, the approved substitution reduced the count to zero, and two separately named copies reopened consistently. The four other intended candidates remain unclassified, so the entire sample does not pass.

For EPUB unsupported characters, carry forward the position-and-approval method, then rebuild the unresolved part of the sample before applying it to an authoritative manuscript. The current Mac feature scope is available on the Rune Studio product page.