CJK Unified Ideographs Extension B

From HandWiki
Short description: Unicode character block
CJK Unified Ideographs Extension B
RangeU+20000..U+2A6DF
(42,720 code points)
PlaneSIP
ScriptsHan
Assigned42,718 code points
Unused2 reserved code points
Unicode version history
3.142,711 (+42,711)
13.042,718 (+7)
Note: [1][2]
Historic single-glyph (UCS2003) code chart
File:CJK-B Fontforge.png
A screenshot of a CJK extension B on Fontforge

CJK Unified Ideographs Extension B is a Unicode block containing rare and historic CJK ideographs for Chinese, Japanese, Korean, and Vietnamese submitted to the Ideographic Research Group between 1998 and 2000, plus seven gongche characters for kunqu added in Unicode 13.0, and two characters for the Macao Supplementary Character Set added in Unicode 14.0.[3]

Variation selector sequences

The block has dozens of variation sequences defined for standardized variants.[4] It also has thousands of ideographic variation sequences registered in the Unicode Ideographic Variation Database (IVD).[5][6] These sequences specify the desired glyph variant for a given Unicode character.

Development process

SuperCJK

The repertoire of Extension B was developed by the Ideographic Rapporteur Group in successive revisions of a document named "SuperCJK".[7] This lists all characters included at the time in the Unified Repertoire and Ordering, in Extension A and in the draft Extension B, ordered by radical and residual stroke count. Each character was included with up to four glyph bitmaps, and references to character sources including the Kangxi Dictionary, the Hanyu Da Zidian, and various national character set standards.[8]

"UCS2003" glyphs

Extension B was the only CJK Unified Ideographs Extension block with a UCS2003 source identifier. Since Extension B contained too many characters, the original code charts were produced with a single glyph for all regions. The glyphs were designed by Beijing Zhongyi Electronic Ltd. After the introduction of multi-column code charts on Unicode 5.2, the original glyphs were retained under the UCS2003 source identifier; they were then removed in Unicode 14.0, being redundant as well as misleading.[9] The glyphs are packaged in the "SimSun-ExtB" font distributed with the Simplified Chinese versions of Windows, and do not adhere to the glyphs for the Mainland China region.

Known issues

Ken Lunde noted that what happens in Extension B stays in Extension B. Peter Constable pointed out that we have been dealing with many historical errors in this block over the years.

UTC #187 meeting minutes[10]

Size

Extension B contains 42720 (initially 42711) characters, more than double the 20992 included in the Unified Repertoire and Ordering, and far more than any other CJK extension block (the second-largest of which, as of Unicode 17.0, is CJK Unified Ideographs Extension F, with 7473 characters).[11] This also makes it the only block, besides Private Use Areas, to occupy more than half of the 65534 allocatable codepoints in a Unicode plane.

Owing to its size, a number of mistakes have been discovered in the original version of Extension B since it was encoded.[10][12] To facilitate better review of proposed characters, the working set size for subsequent CJK extension blocks has been kept much smaller; for example, the Ideographic Research Group currently requests that its member bodies should usually submit no more than 2500 characters each, so as to keep the size of the working set below 10000 characters if possible.[13]

Unifiable variants and exact duplicates in Extension B

In CJK Unified Ideographs Extension B, hundreds of glyph variants were encoded.[14] In addition to the deliberate encoding of close glyph variants, seven exact duplicates (where the same character has inadvertently been encoded twice) and two semi-duplicates (where the CJK-B character represents a de facto disunification of two glyph forms unified in the corresponding BMP character) were encoded by mistake:[15]

  • U+34A8 㒨 = U+20457 𠑗 : U+20457 is the same as the China-source glyph for U+34A8, but it is significantly different from the Taiwan-source glyph for U+34A8
  • U+3DB7 㶷 = U+2420E 𤈎 : same glyph shapes
  • U+8641 虁 = U+27144 𧅄 : U+27144 is the same as the Korean-source glyph for U+8641, but it is significantly different from the Chinese Mainland-, Taiwan- and Japan-source glyphs for U+8641
  • U+204F2 𠓲 = U+23515 𣔕 : same glyph shapes, but ordered under different radicals
  • U+21018 𡀘 = U+2103C 𡀼 : same glyph shapes
  • U+249BC 𤦼 = U+249E9 𤧩 : same glyph shapes
  • U+24BD2 𤯒 = U+2A415 𪐕 : same glyph shapes, but ordered under different radicals
  • U+26842 𦡂 = U+26866 𦡦 : same glyph shapes
  • U+FA23 﨣 = U+27EAF 𧺯 : same glyph shapes (U+FA23 﨣 is a unified CJK ideograph, despite its name "CJK COMPATIBILITY IDEOGRAPH-FA23.")

Block

History

The following Unicode-related documents record the purpose and process of defining specific characters in the CJK Unified Ideographs Extension B block:

References

  1. ↑ "Unicode character database". The Unicode Standard. https://www.unicode.org/ucd/. Retrieved 2023-07-26. 
  2. ↑ "Enumerated Versions of The Unicode Standard". The Unicode Standard. https://www.unicode.org/versions/enumeratedversions.html. Retrieved 2023-07-26. 
  3. ↑ "18.1: Han (§ Blocks Containing Han Ideographs)", The Unicode Standard: Core Specification, Version 15.0, 2022, pp. 741–744, ISBN 978-1-936213-32-0, https://www.unicode.org/versions/Unicode15.0.0/ch18.pdf#page=6 
  4. ↑ "Unicode Character Database: Standardized Variation Sequences". The Unicode Consortium. https://www.unicode.org/Public/UNIDATA/StandardizedVariants.txt. 
  5. ↑ "Ideographic Variation Database". Unicode Consortium. https://www.unicode.org/ivd/. 
  6. ↑ "UTS #37, Unicode Ideographic Variation Database". Unicode Consortium. https://www.unicode.org/reports/tr37/. 
  7. ↑ Japan National Body (2011-03-01). "Re-confirm on how to use Super CJK document and Japan's concern". https://www.unicode.org/irg/docs/n1754-SuperCJK.pdf. 
  8. ↑ Ideographic Rapporteur Group. "SuperCJK 14.0". https://www.unicode.org/irg/docs/n0802-SuperCJKv14.pdf. 
  9. ↑ "The Unicode Standard, Version 14.0". Unicode Consortium. https://www.unicode.org/versions/Unicode14.0.0/ch24.pdf. 
  10. ↑ 10.0 10.1 "Draft Minutes of UTC Meeting 187: Berkeley, California, United States — April 21-23, 2026". 2026-04-29. https://www.unicode.org/L2/L2026/26093.htm. 
  11. ↑ Lunde, Ken (2025-09-09). "Unicode Version 17.0 CJK Unified Ideographs & CJK Compatibility Ideographs". https://www.unicode.org/irg/docs/n2856-CJKUIv17.pdf. 
  12. ↑ West, Andrew (2007-12-02). "CJK-B Case Study #1: U+272F0". https://babelstone.co.uk/Blog/2007/12/cjk-b-case-study-1-u272f0.html. 
  13. ↑ Lunde, Ken (2026-04-17), IRG Principles & Procedures Version 18, Ideographic Research Group / Unicode Consortium 
  14. ↑ "unifiable glyph variants". http://www.cse.cuhk.edu.hk/~irg/irg/irg25/IRGN1155_Possible_Duplicates.pdf. 
  15. ↑ Cook, Richard. "Defect Report on Duplicate Encoded CJK Forms". https://www.unicode.org/wg2/docs/n2644.pdf.