The Case Against Kanji
On Accepting when Tradition no longer Serves us.. in this Very Specific Domain
The Precedence:
Hangul, the Korean alphabet, was debuted in 1446 by King Sejong the Great. The system was designed to replace Chinese characters, which were deemed too difficult for the common man to learn. It took several hundred years for the wheels of progress to turn, but eventually Korea finally adopted what I argue is the world’s greatest writing system. It’s a phonetic alphabet, organized into syllabic blocks. And it’s gorgeous.
저는 한국어를 못해서 구글 번역기의 도움을 받았어요. 한글 멋지지 않나요?
Japan, like Korea, first developed writing through contact with China, originally writing only in Chinese characters. It kind-of, sort-of worked for a bit. But the core issue was that Chinese characters were not made to fit with Japanese speech. Slowly, Japan developed two syllabaries to make it simpler to demarcate pronunciation and to take quick notes, and over time these became standardized and integrated more and more into the language. However, unlike Korean, Japanese still uses many of their inherited Chinese characters. These characters laugh in the face of a phonetics. It’s (arguably) the least phonetic commonly used writing system in the world.
Unfortunately, the resultant system in Japanese ended up being, frankly, quite bizarre. In the average Japanese sentence, you may have a mix of kanji (logographs), hiragana (the normal syllabary), and sometimes katakana (a syllabary used largely for foreign imported words and for emphasis). Collectively the two syllabaries are called, appropriately, kana. Thankfully, the kana are fully-phonetic.
すこし日本語が話せるので、今度グーグル翻訳を使いませんでした。
To oversimplify, hiragana are simple “swirly” characters, katakana are simple “angular” characters, and kanji… uhh… they can look as simple as kana (look at 日 above), but are often quite complicated and dense (翻訳).
Kanji are Hard:
Kanji aren’t just visually complicated though. If you were hoping that you can figure out how to pronounce a kanji just by looking at its components, I’m sorry to say that’s not the case. Each kanji word has kana “inside it”, called its reading. Without first memorizing the word, you can’t be sure of what the reading is just by looking at the kanji. That remains true even if you’ve seen the same kanji before in a different word. At best, you can guess based on various guidelines, but nothing particularly foolproof.
So Japanese is in the position where a native speaker may know exactly what a word means, and how exactly it is pronounced, but could be totally unable to read or write it1.
It gets worse. Reading (and writing) kanji can be so difficult that being able to do so at a high level becomes a signal of general competence and having passed high levels of the Kanji Kentei test (level 2 or the more difficult level 1) is often placed on resumes of native Japanese speakers.
“Ahh” you say, “That’s kind of like a spelling bee for English, but it’s taken more seriously”. And… kind of, but not really. In high-level spelling bees, you might have to spell words the average person has never heard of, like “pococurante”. The difference being that English spelling fail relatively nicely. If you mispronounce a word due to a strange spelling, you are generally understood. The same is not nearly the case with kanji.
With level 2 of the Kanji Kentei, you are tested on all the kanji you’re supposed to know to by the time you graduate high school. Imagine writing “Native English speaker, fluent in spelling ‘acquiesce,’ ‘conscientious,’ and 2,000 other high-school level words” on your resume2. With kanji though, it is an accomplishment to master them, because kanji are hard.
Why not get rid of them?
From a naive perspective, it seems silly to have some other language’s script (that maps poorly to your own way of speaking) still being the main way that you communicate in text. Are the costs of kanji really worth the benefits? Couldn’t we switch to another, better designed system?
Not everyone thinks so, and those “Not everyone”s come with arguments!
Some arguments rely on appeals to tradition or rely on the ineffable aesthetics of kanji. I won’t be covering those here.
Another argument is that switching over to a different writing system would be hard. And while I admit that costs would be incurred, I think that these temporary costs would be dwarfed by the costs saved by not teaching every future generation thousands of kanji. If anything, I think this argument points more in the direction of “How fast should a transition be?” rather than “Should there be one at all?”.
More interesting arguments to me would be ones in which we shouldn’t switch even if there were no costs to do so. In other words, the claim is that the benefits of kanji are worth the cost to teach them. Proponents of this view argue:
Kanji can provide a “clue” to what an unknown word means.
Reading kana-only, while possible, is hard to parse quickly.
Kanji can differentiate homophones, similar to different English spellings like “pear” and “pair”.
To some extent, each of these arguments is fair! Kanji do absolutely have benefits that would be silly to outright dismiss. That being said, I think they’re much weaker points than kanji advocates tend to assume, and the benefits don’t make up for the costs. But let’s look into these one-by-one.
Argument 1) Kanji can provide a clue to what a word means:
This is true, and to steel-man the argument, I will even point out that it can happen even when the reading doesn’t provide a good clue on its own.
Take a simple word like 図書館 (library). 図 means drawing/diagram, 書 means to write, and 館 is a public building. To the credit of proponents, it’s not something you would get from hearing the word spoken (or seeing the kana) since the first two characters don’t even share the same reading as the common words. Here, you just recognize the kanji. And you can kind of see how those three kanji might mean “library”. Realistically though, if you told me diagram-writing-public-building, I most likely wouldn’t converge on “library”. I might think “cartographer’s office” or “art gallery” or something. In this case, at least, it’s the sort of thing that after I tell you the answer you can sort of see why the kanji work.
Sometimes, there are words that do give some extra meaning through their kanji (ex: 桃色(momoiro) is literally peach-color), but this word would also be determinable through its reading. もも(momo) is “peach” and いろ(iro) is “color”. In a world without kanji, one would still be able to understand the meaning of those quite easily. The kanji is doing very little to help a learner in these cases.
Other times, words are made up of kanji that only exist in one common word. These kanji never a net benefit in helping you figure out the word since it would take less time to learn the word itself than it would to learn the kanji plus the word. It’s like including a word inside it’s definition; unhelpful to those who don’t know the word, and those who do know it don’t need the definition. 挨拶 (aisatsu) is a good example. Both of those kanji only exist in that word.
To be fair, there are words where these objections don’t apply. Famously, the word for “volcano” (火山) is composed of the kanji for “fire” and “mountain”. You really could figure that one out just by the kanji, even in a way that you couldn’t figure it out by the sounds かざん(kazan).
But most words aren’t like that. At best, they provide a kind-of-sort-of clue that might point you vaguely in the direction of the correct word, and maybe, given context, you could figure it out. Is that just my supposition? Am I just subconsciously cherry-picking words in my memory that back up my argument?
Maybe so, so I did my best to test it. The question I want to answer is “What marginal advantage do kanji give in decoding an unknown word?” In order to figure that out, we want to look for the 火山(kazan) type words.
To get the marginal advantage, we want to look at words where we can determine the meaning given both the individual kanji meanings (called glosses) and the context of the sentence, but we don’t want to count words that...
are discernible from context alone.
are discernible from their kana (桃色 (momoiro) type words).
are made of kanji that aren’t used elsewhere (挨拶 (aisatsu) type words).
Eliminating those three types of words leaves us with the 火山(kazan) type words. To do this, we have:
One hundred Japanese sentences with 2-character kanji that are generated by an LLM (so I am as blinded as possible). Each step is done by a different agent blind to the purpose of the entire process, to minimize any chance of it choosing words specifically to support/refute the hypothesis.
It takes the sentences and translates them into English, leaving a blank where the word should be.
I test myself twice:
For the first test, I look at the context of the sentence and try to guess what the (English) word is.
For the second test, I look at the context and the English glosses (grabbed from a Japanese-English dictionary) and give my best guess again.
The answer is revealed, and I compare which ones I get right for both tests. We can see which ones the glosses helped in determining the whole word’s meaning.
Finally, I categorize words as 桃色 (momoiro), 挨拶 (aisatsu), or 火山 (kazan) type words.
Results:
78 words were incorrect both before and after the hints (the kanji weren’t helpful).
2 words had different answers, where I actually switched from the correct answer to an incorrect one after the hint (the kanji were the opposite of helpful).
Of the remaining 20 where I guessed correctly on the 2nd time, 11 of those words were either 桃色 or 挨拶 type words, or some combination of both (see details here3) .
The remaining 9 words were 火山 (kazan) type words, where the kanji did genuine marginal work in helping determine the meaning.
Is this a foolproof test? Absolutely not, but I think it’s decently representative and sure beats “I came up with some kanji off the top of my head to prove the point I’m making”.
So that’s approximately 9% of words where the kanji made a marginal difference. And, to be fair, no kanji advocate would say that one should be able to figure out 100% of all kanji words based on the kanji meanings alone. But, surely, they would think that its much more than ~10%? It would have to be in order for this objection to have teeth. Is the ability to guess ~10% of unknown written words really worth hundreds of hours of study? Remember, you can only do this after you’ve memorized the glosses of kanji you see. If vocab acquisition was what you wanted to accomplish, wouldn’t it just make more sense to memorize more words with those hours?
Ultimately, Japanese is a natural spoken language with kanji retrofitted on top due to historical happenstance. It wasn’t designed, top-down, to have small concepts that chain together to form larger concepts, so we shouldn’t expect individual kanji to be particularly useful in helping us decode unknown words. As it stands, Japanese is no Toki Pona4.
Moving to a fully phonetic system does mean losing something here, but its minor compared to the hundreds of hours that students spend learning kanji.
Argument 2) Having both kana and kanji break up words and make it easier to read more quickly:
First, I have to point out that an argument against kanji isn’t necessarily an argument for a kana-based system. There are other possible non-kanji systems that don’t involve kana at all. Still, taking the argument as written, I can’t deny that kanji help break up long strings of kana. In fact, research seems to back this up.
But you know what else can break up huge strings of kana? Spaces. English has spaces, Korean has spaces, Arabic has spaces. All very different written languages, yet the utility of spaces shines through in each. In other words, by writing in standard Japanese (without spaces), the research above is hugely confounded for the question we’re interested in.
The counter-argument to spaces is that if you replace kanji with kana, the text will become wider (because each kanji has 1-3 kana “inside” it). If you add spaces, it will become wider still. “When is enough, enough?”, asks the kanji advocate.
And how are we to respond? Should we suggest reducing the width of kana characters, since they would no longer have to be as large as the (heavily detailed) kanji? Some sort of half-width characters? That would be madness! Madness, I say!
Here’s standard Japanese in red, kana in blue:
And the English translation in green, because why not:
If you go around the Internet, using half-width katakana in various places, eventually you’re going to end up pissing off a lot of people.
Above, there’s 40 characters in Standard Japanese, and 52(!) in the kana-only version. Both versions have equal height, but as hinted at above, since the kana don’t have to be comically wide, they are easily legible even at half-width5 and they end up taking up less space than Standard Japanese (and yes, I would put spaces in front of particles, fight me).
But, what if there is some reason, unbeknownst to me, that half-width characters are inherently hard to read, and a kana-only system is unworkable as a result? If that’s the case, should we just stick with kanji? I’d say a firm “no” to that. Half-width kana isn’t even my most preferred solution! See the final section for other solutions to this problem.
Argument 3) Kanji can differentiate between Japanese’s many homophones:
Oh boy, here’s the big one. It is 100% true that Japanese has a lot of homophones, more than other languages, and it’s also true that kanji can help disambiguation. But homophones are not nearly as big of a deal as kanji proponents implicitly claim.
Consider that spoken Japanese (short of, sometimes, pitch-accent6) doesn’t differentiate between homophones, and millions of people end up communicating just fine. Proponents of kanji would retort that spoken Japanese is simpler than written Japanese, and talk about instances of Japanese people spelling out kanji in the air when they need to differentiate two homophones in speech. Believe it or not, this is quite rare. Indeed, if it wasn’t, the Japanese would be known for their wild gesticulations while speaking with each other, rather than bowing a comical amount of times during first introductions.
And we have context to rely on, as well. Koshou(こしょう)is often given as a word that has many homophones. Yet still, it’s hard to imagine a non-contrived situation where even something as context-deprived as saying 「こしょうは?」(As for the _____?) isn’t going to be immediately understood in context.
Again, Korean functions just fine without using Chinese characters to disambiguate homophones. Given that, it would seem that the argument must be that Japanese is in the “just-so” space where there are so many homophones that kanji are needed, but not so many that verbal communication is impossible.
“No!” the more astute kanji advocates exclaim, “the issue isn’t the normal, spoken, every-day interactions. It’s the highly technical fields: Science, Law, Medicine! These are where most ambiguous homophones live.”
After grumbling a complaint including the phrase “Russell’s Teapot”, I would simply note that Japanese doctors, lawyers, and scientists seem to communicate just fine to colleagues/clients/patients/judges/juries verbally. And, indeed, medical/scientific terms lean towards foreign imported words, often from English. Perhaps law is the exception, but again… grumble … Russell’s Teapot… burden of proof.
Given that nearly every other language on Earth deals with homophones without resorting to logographs, the burden of proof should be on kanji proponents to demonstrate that they are sufficiently problematic in Japanese to justify the existence of kanji. “There are more homophones in Japanese” doesn’t cut it.
But, again, let’s say I’m completely wrong here. Perhaps, a wife’s written reminder to her husband asking him to turn off a burner (ひをけして) ends with him trying to けして(extinguish) ひ(the day) itself. Maybe a young boy asks for some かみ (paper) and his mother slices off her かみ (hair), while his sister begins to pray to the かみ (spirits/gods). In that case, do I admit defeat and welcome kanji, as a dear friend?
Nope.
Rather than being an existential threat to reform, the worst we get from homophones is some confusing homophonic clashes7. Linguistically, we know that people either structure their sentences to avoid these clashes, or coin new words to avoid the problem entirely. Creating (or importing) new words is not a big ask here. The Japanese love to import foreign words. There’s already an entire syllabary that is dedicated mostly to handling this. They don’t even have to get rid of the native words! If ピンク (pink) can happily live alongside 桃色 (pink), then why can’t ペッパー (pepper) live alongside 胡椒 (pepper)?
Finally, let’s look at the problem from the reverse perspective. Would anyone see a problem like homophone ambiguity in a language without logographs and decide that the solution is to import them? Imagine going up to Sequoyah as he was developing the Cherokee script, and saying “Aha, I see you have a nice, simple syllabic script here for a spoken language that has very few syllables…. Might I suggest adding Chinese logographs to it? You know, mainly so that you can tell apart homophones. And, naturally, you should be using those logographs even in words that don’t have homophones, and the logographs should be read differently based on what word its in. Oh, and make sure they’re very detailed so they take a long time to write and are hard to read. Use this one for reference → 鬱”
Clearly, if Japanese had transitioned to a kanji-free script hundreds of years ago, no serious person would be suggesting that they reincorporate kanji, just like no serious person suggests that Korea does so today. It’s simply a poor tool to solve the problem.
Some Suggested Solutions:
In my dream world, we’re not arguing about whether kanji are worth it (I think my position on that is clear), but about what should replace it and how best to transition to that new system. So, for the sake of completeness, here are my own, personal, most-to-least preferred Japanese writing systems. I’m happy from A to C, and sad from D to G.
A) Linguists design a functional, pretty-looking, compact writing system, with spaces, that still looks distinctly “Japanese”. Bonus points if it’s an alpha-syllabary.
B) Spaced half-width hiragana and katakana (we’d need to add unicode characters for these).
C) Use Korean Hangul characters, slightly adjusted for Japanese phonetics.
D) Heavily reduce the amount of kanji, add spaces.
E) Use the roman alphabet, with adjustments made (like Vietnamese did).
F) The coward’s path: Continue with the current system, forever dooming future generations to kanji drills.
G) Remove the kana, return to the old days where they just write in kanji.
It took Korea about 500 years from the development of Hangul to full implementation. Japan, this is your opportunity to blow them completely out of the water. Two generations, max. You too can experience the pure efficiency of a phonetic writing system. Reject kanji. Embrace progress.
In the digital age, its becoming less common to remember how to write some of the less commonly used kanji and people are more often having to look them up on their phones when they have to hand-write them.
I’m being slightly unfair here. One difficult part of the Kanji Kentei test is getting the correct “stroke order” of the kanji, not a necessity for understanding (or even writing the kanji) in the least. Still, since students are taught stroke order for each kanji in school, it’s worth noting as a huge time-sink when comparing a kanji versus no-kanji system.
巨大 │ きょだい │桃色 type-word
世銀 │ せぎん │挨拶 and 桃色 type-word
蔵元 │ くらもと │挨拶 and 桃色 type-word
忠孝 │ ちゅうこう │挨拶 type-word
花婿 │ はなむこ │桃色 type-word
駅長 │ えきちょう │桃色 type-word
外部 │ がいぶ │桃色 type-word
太鼓 │ たいこ │挨拶 type-word
内需 │ ないじゅ │挨拶 and 桃色 type-word
全容 │ ぜんよう │桃色 type-word
本国 │ ほんごく │桃色 type-word
一斉 │ いっせい │火山 type-word
歌手 │ かしゅ │火山 type-word
対談 │ たいだん │火山 type-word
自社 │ じしゃ │火山 type-word
狂犬 │ きょうけん │火山 type-word
血行 │ けっこう │火山 type-word
上旬 │ じょうじゅん │火山 type-word
今期 │ こんき │火山 type-word
区内 │ くない │火山 type-word
https://tokipona.org/ - Toki Pona is a constructed (“made-up”) minimalist language that tries to form all the concepts you need by chaining together ~140 basic words (ex: big-paper-house meaning library). Normally it isn’t logographic, but it does have an official logograms (sitelen pona) made by the creator of the language. It’s worth mentioning that despite being designed exactly for the purpose of building concepts on top of each other, the language is notoriously ambiguous, to the point where anything more than basic communication is strained.
For weird historical ASCII reasons, there aren’t currently half-width hiragana characters in Unicode (hence the picture of text I manually made above). There are half-width katakana characters, but they have weird bits too (specifically with the dakuten゛and handakuten゜). This could all be easily changed if Japan decided to switch.
You would be forgiven for thinking that pitch accent is central to Japanese the same way that “tones” are central to Chinese languages, but they are in completely different leagues. In terms of importance to the language, pitch accents are more like the stress patterns in English words, which are never denoted outside linguistic circles. Indeed, pitch accents differ between dialects and this causes little confusion among speakers. If pitch accent were so important to Japanese, it could be much more elegantly denoted in a well-designed writing system.
https://www.cambridge.org/core/journals/language-and-cognition/article/whos-afraid-of-homophones-a-multimethodological-approach-to-homophony-avoidance/9F89DC60F67812808F9D5850BBD739D1 - This study looked at the Dutch and how they change the structure of their speech to avoid homophones based on the subject of the sentence (some verb tenses are homophones with other verb tenses with “we” as the subject, but not with “I”, for example).


You say, "Korean functions just fine without using Chinese characters to disambiguate homophones," but this isn't quite true. In some texts, particularly technical texts where your audience is expected to be educated it seems (native speakers feel free to clarify), you will often have hanja in parentheses after some hangul in order to disambiguate it.
This reminds me of "Ze drem vil finali kum tru": http://ashvital.freeservers.com/ze_dream.htm, where there might be a neat logical explanation for everything, but it ends up mangling the language overall. If you're going to go that far, might as well just convert everything to English or Esperanto.