ICU4C and ICU4J are Very Comparable
International Elements for Unicode (ICU) is an open-source venture of mature C/C++ and Java libraries for Unicode help, software internationalization, Pattaya Properties and software globalization. [2] ICU has been included as an ordinary element with Microsoft Home windows since Home windows 10 version 1703.[3] ICU is broadly portable to many working programs and environments. It provides functions the same outcomes on all platforms and between C, C++, and Java software program. The ICU undertaking is a technical committee of the Unicode Consortium and sponsored, supported, and used by IBM and plenty of different companies.
ICU provides the following services: Unicode text handling, full character properties, and character set conversions; Unicode common expressions; full Unicode sets; character, word, and line boundaries; language-sensitive collation and looking; normalization, higher and lowercase conversion, and script transliterations; complete locale information and useful resource bundle structure via the Frequent Locale Knowledge Repository (CLDR); multiple calendars and time zones; and rule-based formatting and parsing of dates, occasions, numbers, currencies, and messages. ICU provided advanced textual content structure service for Arabic, Hebrew, Indic, and Thai historically, but that was deprecated in model 54, and was utterly eliminated in model 58 in favor of HarfBuzz.[4]
ICU gives more extensive internationalization facilities than the usual libraries for C and C++. Future ICU 75 deliberate for April 2024 would require C++17 (up from C++11) or C11 (up from C99), relying on what languages is used. ICU has traditionally used UTF-16, and nonetheless does just for Java; while for C/C++ UTF-8 is supported,[5][6] including the correct handling of “illegal UTF-8”.[7]
ICU 73.2 has improved important modifications for GB18030-2022 compliance help, i.e. for Chinese (that up to date Chinese language GB18030 Unicode Transformation Format customary is slightly incompatible); has “a modified character conversion table, mapping some GB18030 characters to Unicode characters that had been encoded after GB18030-2005” and has plenty of other modifications reminiscent of enhancing Japanese and Korean short-text line breaking, and in “English, the name “Türkiye” is now used for the country as an alternative of “Turkey” (the alternate spelling can be obtainable in the data).”[8]
ICU seventy four “updates to Unicode 15.1, together with new characters, emoji, security mechanisms, and corresponding APIs and implementations. [..] ICU seventy four and CLDR 44 are major releases, together with a new model of Unicode and major locale data enhancements.”[9] Of the numerous changes some are for individual identify formatting, or for improved language support, e.g. for Low German, and there’s e.g. a brand new spoof checker API, following the (latest version) Unicode 15.1.Zero UTS #39: Unicode Security Mechanism.
Older version details
[edit]
ICU seventy two up to date to Unicode 15 (and 73.2 to newest 15.1). “In lots of formatting patterns, ASCII areas are changed with Unicode areas (e.g., a “thin space”).” ICU (ICU4J) now requires Java 8 but “Many of the ICU seventy two library code ought to still work with Java 7 / Android API level 21, but we not check with Java 7.”[10] ICU 71 added e.g. phrase-primarily based line breaking for Japanese (earlier strategies didn’t work properly for brief Japanese textual content, comparable to in titles and headings) and support for Hindi written in Latin letters (hello_Latn), also referred to as “Hinglish”. Help for AIX, Solaris and z/OS could even be restricted in later versions (i.e. building relies on compiler assist). ICU 64.2 added help for Unicode 12.1, i.e. the single new symbol for current Japanese Reiwa era (but support for it has also been backported to older ICU variations down to ICU 4.8.2). ICU 58 (with Unicode 9.Zero assist) is the final model to help older platforms similar to Windows XP and Home windows Vista. ICU 70 added e.g. support for emoji properties of strings and might now be constructed and used with C++20 compilers (and “ICU operator==() and operator!=() capabilities now return bool as a substitute of UBool, as an adjustment for incompatible changes in C++20”),[11] and as of that model the minimal Home windows model is Windows 7. ICU 67 handles removing of Great Britain from the EU. [12]
Origin and growth
[edit]
After Taligent turned part of IBM in early 1996, Solar Microsystems decided that the new Java language should have better help for internationalization. [13] A big portion of this code still exists within the java.text and java.util packages. Further internationalization options have been added with every later launch of Java. Since Taligent had experience with such technologies and were shut geographically, their Textual content and International group have been requested to contribute the worldwide classes to the Java Growth Equipment as a part of the JDK 1.1 internationalization APIs.
The Java internationalization classes were then ported to C++ and C[14] as a part of a library often known as ICU4C (“ICU for C”). Each frameworks have been enhanced over time to assist new amenities and new features of Unicode and customary Locale Data Repository (CLDR). The ICU project also provides ICU4J (“ICU for Java”), which adds options not present in the standard Java libraries. ICU4C and ICU4J are very comparable, though not similar; for example, ICU4C includes a daily Expression API, while ICU4J doesn’t.
ICU was released as an open-source mission in 1999 below the identify IBM Courses for Unicode. [16] [15] In Could 2016, the ICU undertaking joined the Unicode consortium as technical committee ICU-TC, and the library sources at the moment are distributed beneath the Unicode license. It was later renamed to International Elements For Unicode.
MessageFormat
[edit]
Part of ICU is the MessageFormat class, a formatting system that permits for any number of arguments to control the plural form (plural, selectordinal) or extra basic switch-case-fashion selection (choose) for things like grammatical gender. These statements can be nested. ICU MessageFormat was created by adding the plural. Selection system to an identically-named system in Java SE.
Alternatives
[edit]
An alternate for utilizing ICU with C++, or to utilizing it instantly, is to use Enhance.Locale, which is a C++ wrapper for ICU (while additionally allowing other backends[18]). The claim for utilizing it reasonably than ICU straight is that “is completely unfriendly to C++ builders. It ignores standard C++ idioms (the STL, RTTI, exceptions, etc), as a substitute largely mimicking the Java API.”[19][20] One other declare, that ICU only supports UTF-16 (and thus a reason to avoid using ICU) is not true with ICU now also supporting UTF-eight for C and C++.[5]
Apple Superior Typography
Apple Sort Services for Unicode Imaging
gettext
Graphite (good font expertise)
NetRexx (ICU license)
OpenType
Pango
Uconv
Uniscribe
^ unicode-org. “Launch ICU 78.3 · unicode-org/icu“. Retrieved 18 March 2026.
^ “ICU – Worldwide Components for Unicode”. site.icu-undertaking.org. Archived from the original on 2021-08-27. Retrieved 2011-11-14.
^ Chen, Raymond (27 Might 2021). “How can I convert between IANA time zones and Home windows registry-primarily based time zones?”. The Outdated New Thing. Microsoft.
^ “Structure Engine – ICU Consumer Information”. userguide.icu-undertaking.org.
^ a b “UTF-8”. ICU Documentation. Retrieved 2022-05-24.
^ “UTF-8 – ICU User Guide”. userguide.icu-project.org. Retrieved 2018-04-03.
^ “#13311 (change illegal-UTF-8 handling to Unicode “finest follow”)”. bugs.icu-undertaking.org. Retrieved 2018-04-03.
^ “ICU – International Elements for Unicode – ICU 73”. icu.unicode.org. Retrieved 2023-09-24.
^ “ICU – Worldwide Components for Unicode – ICU 74″. icu.unicode.org. Retrieved 2023-11-29.
^ “ICU – International Parts for Unicode – ICU 72″. icu.unicode.org. Retrieved 2023-01-24.
^ “ICU – Worldwide Parts for Unicode – ICU 70″. icu.unicode.org. Retrieved 2023-01-24.
^ “Download ICU 64 – ICU – Worldwide Components for Unicode”. site.icu-venture.org. Retrieved 2019-10-20.
^ Laura Werner (1999). “Getting Java ready for the world: A short history of IBM and Sun’s internationalization efforts”. Archived from the unique on 2021-11-17. Retrieved 2007-05-23.
^ “ICU Consumer Guide”. userguide.icu-venture.org.
^ “ICU Project Management Committee”. Archived from the original on 2021-08-28. Retrieved 2012-08-17.
^ “ICU joins the Unicode Consortium”. Unicode, Inc. 2016-05-16. Retrieved 2016-08-01.
^ “Formatting Messages”. ICU User Guide.
^ “Boost.Locale: Utilizing Localization Backends”. www.increase.org. Retrieved 2022-05-24.
^ “Increase.Locale: Design Rationale”. www.increase.org. Retrieved 2022-05-24.
^ “ICU vs Enhance Locale in C++”. Stack Overflow. Retrieved 2022-05-24.
Worldwide Components for Unicode transliteration services
ICU Editor with Visible Preview
Unicode Consortium
ISO/IEC 10646 (Universal Character Set)
Variations
Code
points
Universal Character Set
Character charts
Character property
Airplane
Private Use Space
Pairs
Compatibility characters
Homoglyph
Precomposed character
listing
Z-variant
Regional indicator symbol
Emoji skin coloration
Particular
objective
BOM
Combining grapheme joiner
Left-to-right mark – Right-to-left mark
Soft hyphen
Variant type
Word joiner
Zero-width joiner
Zero-width non-joiner
Zero-width area
Lists
Characters
CJK Unified Ideographs
Combining character
Duplicate characters
Numerals
Halfwidth – fullwidth
Alias names – abbreviations
Whitespace characters
Processing
Algorithms
Bidirectional textual content
Collation
ISO/IEC 14651
Equivalence
Variation sequences
Encoding
comparability
BOCU-1
CESU-eight
Punycode
SCSU
UTF-1
UTF-7
UTF-8
UTF-16/UCS-2
UTF-32/UCS-4
UTF-EBCDIC
Use
Domain names (IDN)
Fonts
HTML
entity references
numeric references
Enter
International Ideographs Core
Related
requirements
Widespread Locale Information Repository (CLDR)
GB 18030
ISO/IEC 8859
DIN 91379
ISO 15924
Related
subjects
Anomalies
ConScript Unicode Registry
Ideographic Analysis Group
Individuals concerned with Unicode
Han unification
Scripts and symbols in Unicode
Scripts
Frequent,
inherited
Combining marks
Diacritics
Punctuation marks
Spaces
Numbers
Fashionable
Adlam
Arabic
Armenian
Balinese
Bamum
Batak
Bengali
Beria Erfe
Bopomofo
Braille
Buhid
Burmese
Canadian Aboriginal
Chakma
Cham
Cherokee
CJK Unified Ideographs (Han)
Cyrillic
Deseret
Devanagari
Garay
Geʽez
Georgian
Greek
Gujarati
Gunjala Gondi
Gurmukhi
Gurung Khema
Hangul
Hanifi Rohingya
Hanja
Hanunuoo
Hebrew
Hiragana
Javanese
Kanji
Kannada
Katakana
Kayah Li
Khmer
Kirat Rai
Lao
Latin
Lepcha
Limbu
Lisu (Fraser)
Lontara
Malayalam
Masaram Gondi
Mende Kikakui
Medefaidrin
Miao (Pollard)
Mongolian
Mru
N’Ko
Nag Mundari
New Tai Lue
Nüshu
Nyiakeng Puachue Hmong
Odia
Ol Chiki
Ol Onal
Osage
Osmanya
Pahawh Hmong
Pau Cin Hau
Pracalit (Newa)
Ranjana
Rejang
Samaritan
Saurashtra
Shavian
Sinhala
Sorang Sompeng
Sundanese
Sunuwar
Syriac
Tagbanwa
Tai Le
Tai Tham
Tai Viet
Tai Yo
Tamil
Tangsa
Telugu
Thaana
Thai
Tibetan
Tifinagh
Tirhuta
Tolong Siki
Toto
Vai
Wancho
Warang Citi
Yi
Historical,
historic
Ahom
Anatolian hieroglyphs
Historical North Arabian
Avestan
Bassa Vah
Bhaiksuki
Brāhmī
Carian
Caucasian Albanian
Coptic
Cuneiform
Cypriot
Cypro-Minoan
Dives Akuru
Dogra
Egyptian hieroglyphs
Elbasan
Elymaic
Glagolitic
Gothic
Grantha
Hatran
Imperial Aramaic
Inscriptional Pahlavi
Inscriptional Parthian
Kaithi
Kawi
Kharosthi
Khitan small script
Khojki
Khudawadi
Khwarezmian (Chorasmian)
Linear A
Linear B
Lycian
Lydian
Mahajani
Makasar
Mandaic
Manichaean
Marchen
Meetei Mayek
Meroitic
Modi
Multani
Nabataean
Nandinagari
Ogham
Old Hungarian
Old Italic
Previous Permic
Old Persian cuneiform
Old Sogdian
Outdated Turkic
Old Uyghur
Palmyrene
ʼPhags-pa
Phoenician
Psalter Pahlavi
Runic
Sharada
Siddham
Sidetic
Sogdian
South Arabian
Soyombo
Sylheti Nagri
Tagalog (Baybayin)
Takri
Tangut
Todhri
Tulu Tigalari
Ugaritic
Vithkuqi
Yezidi
Zanabazar Square
Notational
Duployan
SignWriting
Symbols
Cultural, political, religious symbols
Foreign money symbols
Control Footage
Mathematical operators, symbols
Glossary
Phonetic symbols (together with IPA)
Emoji
Class: Unicode
Class: Unicode blocks
Retrieved from “https://en.wikipedia.org/w/index.php?title=International_Elements_for_Unicode&oldid=1336484803”
Unicode
Element-based mostly software program engineering
Digital typography
Pattern matching
Internationalization and localization
Free pc libraries
This web page was last edited on four February 2026, at 01:Fifty one (UTC).
International Parts for Unicode