ICU4C and ICU4J are Very Comparable

ICU4C and ICU4J are Very Comparable

"Samui Airport"International Elements for Unicode (ICU) is an open-source venture of mature C/C++ and Java libraries for Unicode help, software internationalization, Pattaya Properties and software globalization. [2] ICU has been included as an ordinary element with Microsoft Home windows since Home windows 10 version 1703.[3] ICU is broadly portable to many working programs and environments. It provides functions the same outcomes on all platforms and between C, C++, and Java software program. The ICU undertaking is a technical committee of the Unicode Consortium and sponsored, supported, and used by IBM and plenty of different companies.

ICU provides the following services: Unicode text handling, full character properties, and character set conversions; Unicode common expressions; full Unicode sets; character, word, and line boundaries; language-sensitive collation and looking; normalization, higher and lowercase conversion, and script transliterations; complete locale information and useful resource bundle structure via the Frequent Locale Knowledge Repository (CLDR); multiple calendars and time zones; and rule-based formatting and parsing of dates, occasions, numbers, currencies, and messages. ICU provided advanced textual content structure service for Arabic, Hebrew, Indic, and Thai historically, but that was deprecated in model 54, and was utterly eliminated in model 58 in favor of HarfBuzz.[4]

ICU gives more extensive internationalization facilities than the usual libraries for C and C++. Future ICU 75 deliberate for April 2024 would require C++17 (up from C++11) or C11 (up from C99), relying on what languages is used. ICU has traditionally used UTF-16, and nonetheless does just for Java; while for C/C++ UTF-8 is supported,[5][6] including the correct handling of “illegal UTF-8”.[7]

ICU 73.2 has improved important modifications for GB18030-2022 compliance help, i.e. for Chinese (that up to date Chinese language GB18030 Unicode Transformation Format customary is slightly incompatible); has “a modified character conversion table, mapping some GB18030 characters to Unicode characters that had been encoded after GB18030-2005” and has plenty of other modifications reminiscent of enhancing Japanese and Korean short-text line breaking, and in “English, the name “Türkiye” is now used for the country as an alternative of “Turkey” (the alternate spelling can be obtainable in the data).”[8]

ICU seventy four “updates to Unicode 15.1, together with new characters, emoji, security mechanisms, and corresponding APIs and implementations. [..] ICU seventy four and CLDR 44 are major releases, together with a new model of Unicode and major locale data enhancements.”[9] Of the numerous changes some are for individual identify formatting, or for improved language support, e.g. for Low German, and there’s e.g. a brand new spoof checker API, following the (latest version) Unicode 15.1.Zero UTS #39: Unicode Security Mechanism.

Older version details
[edit]
ICU seventy two up to date to Unicode 15 (and 73.2 to newest 15.1). “In lots of formatting patterns, ASCII areas are changed with Unicode areas (e.g., a “thin space”).” ICU (ICU4J) now requires Java 8 but “Many of the ICU seventy two library code ought to still work with Java 7 / Android API level 21, but we not check with Java 7.”[10] ICU 71 added e.g. phrase-primarily based line breaking for Japanese (earlier strategies didn’t work properly for brief Japanese textual content, comparable to in titles and headings) and support for Hindi written in Latin letters (hello_Latn), also referred to as “Hinglish”. Help for AIX, Solaris and z/OS could even be restricted in later versions (i.e. building relies on compiler assist). ICU 64.2 added help for Unicode 12.1, i.e. the single new symbol for current Japanese Reiwa era (but support for it has also been backported to older ICU variations down to ICU 4.8.2). ICU 58 (with Unicode 9.Zero assist) is the final model to help older platforms similar to Windows XP and Home windows Vista. ICU 70 added e.g. support for emoji properties of strings and might now be constructed and used with C++20 compilers (and “ICU operator==() and operator!=() capabilities now return bool as a substitute of UBool, as an adjustment for incompatible changes in C++20”),[11] and as of that model the minimal Home windows model is Windows 7. ICU 67 handles removing of Great Britain from the EU. [12]

Origin and growth
[edit]
After Taligent turned part of IBM in early 1996, Solar Microsystems decided that the new Java language should have better help for internationalization. [13] A big portion of this code still exists within the java.text and java.util packages. Further internationalization options have been added with every later launch of Java. Since Taligent had experience with such technologies and were shut geographically, their Textual content and International group have been requested to contribute the worldwide classes to the Java Growth Equipment as a part of the JDK 1.1 internationalization APIs.

The Java internationalization classes were then ported to C++ and C[14] as a part of a library often known as ICU4C (“ICU for C”). Each frameworks have been enhanced over time to assist new amenities and new features of Unicode and customary Locale Data Repository (CLDR). The ICU project also provides ICU4J (“ICU for Java”), which adds options not present in the standard Java libraries. ICU4C and ICU4J are very comparable, though not similar; for example, ICU4C includes a daily Expression API, while ICU4J doesn’t.

ICU was released as an open-source mission in 1999 below the identify IBM Courses for Unicode. [16] [15] In Could 2016, the ICU undertaking joined the Unicode consortium as technical committee ICU-TC, and the library sources at the moment are distributed beneath the Unicode license. It was later renamed to International Elements For Unicode.

MessageFormat
[edit]
Part of ICU is the MessageFormat class, a formatting system that permits for any number of arguments to control the plural form (plural, selectordinal) or extra basic switch-case-fashion selection (choose) for things like grammatical gender. These statements can be nested. ICU MessageFormat was created by adding the plural. Selection system to an identically-named system in Java SE.

Alternatives
[edit]
An alternate for utilizing ICU with C++, or to utilizing it instantly, is to use Enhance.Locale, which is a C++ wrapper for ICU (while additionally allowing other backends[18]). The claim for utilizing it reasonably than ICU straight is that “is completely unfriendly to C++ builders. It ignores standard C++ idioms (the STL, RTTI, exceptions, etc), as a substitute largely mimicking the Java API.”[19][20] One other declare, that ICU only supports UTF-16 (and thus a reason to avoid using ICU) is not true with ICU now also supporting UTF-eight for C and C++.[5]

Apple Superior Typography

Apple Sort Services for Unicode Imaging

gettext

Graphite (good font expertise)

NetRexx (ICU license)

OpenType

Pango

Uconv

Uniscribe

^ unicode-org. “Launch ICU 78.3 · unicode-org/icu“. Retrieved 18 March 2026.

^ “ICU – Worldwide Components for Unicode”. site.icu-undertaking.org. Archived from the original on 2021-08-27. Retrieved 2011-11-14.

^ Chen, Raymond (27 Might 2021). “How can I convert between IANA time zones and Home windows registry-primarily based time zones?”. The Outdated New Thing. Microsoft.

^ “Structure Engine – ICU Consumer Information”. userguide.icu-undertaking.org.

^ a b “UTF-8”. ICU Documentation. Retrieved 2022-05-24.

^ “UTF-8 – ICU User Guide”. userguide.icu-project.org. Retrieved 2018-04-03.

^ “#13311 (change illegal-UTF-8 handling to Unicode “finest follow”)”. bugs.icu-undertaking.org. Retrieved 2018-04-03.

^ “ICU – International Elements for Unicode – ICU 73”. icu.unicode.org. Retrieved 2023-09-24.

^ “ICU – Worldwide Components for Unicode – ICU 74″. icu.unicode.org. Retrieved 2023-11-29.

^ “ICU – International Parts for Unicode – ICU 72″. icu.unicode.org. Retrieved 2023-01-24.

^ “ICU – Worldwide Parts for Unicode – ICU 70″. icu.unicode.org. Retrieved 2023-01-24.

^ “Download ICU 64 – ICU – Worldwide Components for Unicode”. site.icu-venture.org. Retrieved 2019-10-20.

^ Laura Werner (1999). “Getting Java ready for the world: A short history of IBM and Sun’s internationalization efforts”. Archived from the unique on 2021-11-17. Retrieved 2007-05-23.

^ “ICU Consumer Guide”. userguide.icu-venture.org.

^ “ICU Project Management Committee”. Archived from the original on 2021-08-28. Retrieved 2012-08-17.

^ “ICU joins the Unicode Consortium”. Unicode, Inc. 2016-05-16. Retrieved 2016-08-01.

^ “Formatting Messages”. ICU User Guide.

^ “Boost.Locale: Utilizing Localization Backends”. www.increase.org. Retrieved 2022-05-24.

^ “Increase.Locale: Design Rationale”. www.increase.org. Retrieved 2022-05-24.

^ “ICU vs Enhance Locale in C++”. Stack Overflow. Retrieved 2022-05-24.

Worldwide Components for Unicode transliteration services

ICU Editor with Visible Preview

Unicode Consortium

ISO/IEC 10646 (Universal Character Set)

Variations

Code
points

Universal Character Set

Character charts

Character property

Airplane

Private Use Space

Pairs

Compatibility characters

Homoglyph

Precomposed character

listing

Z-variant

Regional indicator symbol

Emoji skin coloration

Particular
objective

BOM

Combining grapheme joiner

Left-to-right mark – Right-to-left mark

Soft hyphen

Variant type

Word joiner

Zero-width joiner

Zero-width non-joiner

Zero-width area

Lists

Characters

CJK Unified Ideographs

Combining character

Duplicate characters

Numerals

Halfwidth – fullwidth

Alias names – abbreviations

Whitespace characters

Processing

Algorithms

Bidirectional textual content

"Koh Chang Thailand"Collation

ISO/IEC 14651

Equivalence

Variation sequences

Encoding
comparability

BOCU-1

CESU-eight

Punycode

SCSU

UTF-1

UTF-7

UTF-8

UTF-16/UCS-2

UTF-32/UCS-4

UTF-EBCDIC

Use

Domain names (IDN)

Fonts

HTML

entity references

numeric references

Enter

International Ideographs Core

Related
requirements

Widespread Locale Information Repository (CLDR)

GB 18030

ISO/IEC 8859

DIN 91379

ISO 15924

Related
subjects

Anomalies

ConScript Unicode Registry

Ideographic Analysis Group

Individuals concerned with Unicode

Han unification

Scripts and symbols in Unicode

Scripts

Frequent,
inherited

Combining marks

Diacritics

Punctuation marks

Spaces

Numbers

Fashionable

Adlam

Arabic

Armenian

Balinese

Bamum

Batak

Bengali

Beria Erfe

Bopomofo

Braille

Buhid

Burmese

Canadian Aboriginal

Chakma

Cham

Cherokee

CJK Unified Ideographs (Han)

Cyrillic

Deseret

Devanagari

Garay

Geʽez

Georgian

Greek

Gujarati

Gunjala Gondi

Gurmukhi

Gurung Khema

Hangul

Hanifi Rohingya

Hanja

Hanunuoo

Hebrew

Hiragana

Javanese

Kanji

Kannada

Katakana

Kayah Li

Khmer

Kirat Rai

Lao

Latin

Lepcha

Limbu

Lisu (Fraser)

Lontara

Malayalam

Masaram Gondi

Mende Kikakui

Medefaidrin

Miao (Pollard)

Mongolian

Mru

N’Ko

Nag Mundari

New Tai Lue

Nüshu

Nyiakeng Puachue Hmong

Odia

Ol Chiki

Ol Onal

Osage

Osmanya

Pahawh Hmong

Pau Cin Hau

Pracalit (Newa)

Ranjana

Rejang

Samaritan

Saurashtra

Shavian

Sinhala

Sorang Sompeng

Sundanese

Sunuwar

Syriac

Tagbanwa

Tai Le

Tai Tham

Tai Viet

Tai Yo

Tamil

Tangsa

Telugu

Thaana

Thai

Tibetan

Tifinagh

Tirhuta

Tolong Siki

Toto

Vai

Wancho

Warang Citi

Yi

Historical,
historic

Ahom

Anatolian hieroglyphs

Historical North Arabian

Avestan

Bassa Vah

Bhaiksuki

Brāhmī

Carian

Caucasian Albanian

Coptic

Cuneiform

Cypriot

Cypro-Minoan

Dives Akuru

Dogra

Egyptian hieroglyphs

Elbasan

Elymaic

Glagolitic

Gothic

Grantha

Hatran

Imperial Aramaic

Inscriptional Pahlavi

Inscriptional Parthian

Kaithi

Kawi

Kharosthi

Khitan small script

Khojki

Khudawadi

Khwarezmian (Chorasmian)

Linear A

Linear B

Lycian

Lydian

Mahajani

Makasar

Mandaic

Manichaean

Marchen

Meetei Mayek

Meroitic

Modi

Multani

Nabataean

Nandinagari

Ogham

Old Hungarian

Old Italic

Previous Permic

"Apex Market Analysis"Old Persian cuneiform

Old Sogdian

Outdated Turkic

Old Uyghur

Palmyrene

ʼPhags-pa

Phoenician

Psalter Pahlavi

Runic

Sharada

Siddham

Sidetic

Sogdian

South Arabian

Soyombo

Sylheti Nagri

Tagalog (Baybayin)

Takri

Tangut

Todhri

Tulu Tigalari

Ugaritic

Vithkuqi

Yezidi

Zanabazar Square

Notational

Duployan

SignWriting

Symbols

Cultural, political, religious symbols

Foreign money symbols

Control Footage

Mathematical operators, symbols

Glossary

Phonetic symbols (together with IPA)

Emoji

Class: Unicode

Class: Unicode blocks

Retrieved from “https://en.wikipedia.org/w/index.php?title=International_Elements_for_Unicode&oldid=1336484803”

Unicode

Element-based mostly software program engineering

Digital typography

Pattern matching

Internationalization and localization

Free pc libraries

This web page was last edited on four February 2026, at 01:Fifty one (UTC).

International Parts for Unicode