
Historically, gaiji means characters supplied outside a standard character set; it is not a precise synonym for every variant character or every encoding failure. This article narrows the search term to characters that Shift_JIS cannot represent. The line between what passes and what does not is invisible, so use an encoding check rather than relying on appearance alone.
This article covers which characters get caught, why appearance tells you nothing, and at what point in the workflow finding them costs least. How characters actually disappear, and how to reopen a garbled manuscript, are covered separately.
"Cannot be represented" is not the same as "looks unusual"
An encoding is a table mapping characters to numbers. A character not in the table cannot be written in that encoding.
What defies intuition is that near-identical forms are treated differently. I put three names into a test manuscript — one using a variant form of "taka", one using a variant "yoshi" written with an earth radical, and one using an old form of "se" — and checked which could be represented in Shift_JIS.
- the variant "taka": representable
- the old-form "se": representable
- the earth-radical "yoshi": not representable
One common name variant passes; another does not. Both look equally like "an ordinary variant used in personal names". Memorising where the line falls is not realistic.
Looking by eye, or banning variants outright
The naive approach is rereading the manuscript hunting for suspicious characters — listing your cast and flagging anyone with a variant form in their name.
The other approach is deciding up front never to use variants at all, normalising every name to its standard form. Nothing gets caught that way.
The later you find it, the more places you fix
The problem with looking by eye is not missed characters. It is when you find out.
Such characters can display normally while you write because your Mac can render them. Trouble can begin when the manuscript is saved in an encoding with a smaller repertoire. EPUB text handling, store acceptance, and device rendering are separate checks.
Which means the further down the workflow you get, the more places need fixing. At manuscript stage it is one edit. After building the EPUB it is an edit plus a rebuild. After submitting it is an edit, a rebuild, and a replacement process.
Banning variants has its own cost. Character names are part of the work. Replacing a name that must be written with a particular variant is a compromise a writer feels — and made without knowing which characters actually fail, it includes compromises that were never necessary.
rune Studio returns unrepresentable characters with their positions
rune Studio handles the major Japanese encodings and reports characters that the chosen encoding cannot represent.
I verified this by driving the development build from the command line. With a test manuscript containing all three name variants, checking against Shift_JIS returned exactly one unrepresentable character — the earth-radical "yoshi" — with its position and the character itself, so the place to edit is unambiguous. The other two were not reported.
Attempting to save the same manuscript as Shift_JIS produced no file: the save did not happen, and the result stated that one character could not be represented in that encoding. There is no path where the file quietly saves with a character missing.
The check runs at manuscript stage
This check applies to the manuscript, before any EPUB exists. Its result covers that version of the manuscript against the selected encoding. Recheck after changing the text or target encoding, and treat device rendering as a separate question.
For a long work, running it as each chapter is finished is the practical rhythm — the fix stays inside that chapter.
Encoding detection belongs in the same look
rune Studio detects the encoding when it reads a file. A manuscript saved as Shift_JIS was detected as Shift_JIS. Worrying about gaiji without knowing what encoding your manuscript is in is meaningless, so treat the two as one check.
Who this suits, and who is fine without it
It suits anyone writing names with variant characters — historical fiction, books about real people, works dense with family and place names. Gaiji are unavoidable there.
With a small cast written entirely in standard forms, you can ignore this. Stay in UTF-8 throughout and almost nothing is unrepresentable.
Scope note: what was verified is that characters in the manuscript can be checked. Whether a store accepts a particular character, and whether a device renders it, is outside this check.
Summary
- Gaiji cannot be judged by appearance; one common name variant passes while another fails
- Looking by eye pushes discovery later, which multiplies the places you have to fix
- rune Studio reports unrepresentable characters with positions, and refuses to save while any remain
- The check runs at manuscript stage; running it per chapter keeps the current correction scope contained
For a first step, list your characters' names and see whether all of them are written in standard forms. If even one uses a variant, running the check at manuscript stage is worth the minute it takes.
See the current product scope on the Rune Studio product page.


