String length() and Unicode
In this article, we will find answers to the following questions:
- coding standards and encodings
- unicode and UTF
- UTF-8, UTF-16, UTF-32
- char and code unit
- code point
- string length with emoji
- character encodings
- difference between UTF-8 and UTF-16
- compact strings in Java 9
- how to count emoji
- surrogate pairs in Java strings
- ASCII and Unicode in Java
- working with string encodings in Java
Coding standards and encodings, Unicode and UTF
To begin with, it is necessary to be able to distinguish between encoding standards and encoding formats.
Encoding standard assigns a unique number (code point) to each character regardless of platform, program or language and defines which symbol corresponds to which number.
What symbols exist at all (letters, digits, emoji, hieroglyphs, etc.) and which unique number (code point) corresponds to each character β is set by the standard.
One such standard is Unicode.
Encodings β these are concrete ways of representing these code points as sequences of bytes for storage in computer memory or transmission over a network.
Encoding family implementing the Unicode standard: UTF-8, UTF-16, UTF-32.
ASCII β standard or encoding?
But we have ASCII. What is it then?
In fact, ASCII is 2-in-1: both. ASCII stands for American Standard Code for Information Interchange.
ASCII as a standard
This is a character table, like Unicode, but much smaller: only 128 characters.
It includes:
- English letters (AβZ, aβz)
- Digits (0β9)
- Punctuation marks and special characters (!, @, #, ~, etc.)
- Control characters (for example, line feed \n, carriage return \r)
Each character has a number from 0 to 127 β this is the ASCII code point.
ASCII as an encoding
These same numbers (0β127) are directly mapped to single-byte values.
- That is, character A (code point 65) β byte 01000001
- No additional conversions β one-to-one.
Therefore, ASCII is both a set of characters and an encoding, because it itself defines both what to encode and how to encode it.
ASCII is 7-bit. The last symbol matches the byte 01111111.
The Most Significant Bit (MSB) is the leftmost bit in a byte. It is not part of the ASCII standard, but is still always sent.
On old systems it was used for parity checking, but now it is always 0.
National versions of ASCII
As we have found out, original ASCII is 7-bit.
There are two types of ASCII modifications:
- 8-bit ASCII extensions using the top bit of the original ASCII
- The idea is to add the local alphabet of the country (Windows-1252, KOI8-R, and others)
- 7-bit interpretations of the original ASCII
- The idea is that some symbols (e.g. #, @, [, , ], {, }) can be replaced by national letters.
8-bit ASCII extensions are called by their names: Windows-1252, KOI8-R, etc.
But to avoid confusion with national 7-bit interpretations of original ASCII used in other countries, it is recommended to designate the original code variant as ‘US-ASCII’.
Let’s give an example from Java. In the StandardCharsets class there is no constant ASCII, but there is US-ASCII:
String text = "!";
byte[] bytes = text.getBytes(StandardCharsets.US_ASCII);
System.out.println(bytes.length); // 1 byte
System.out.println(new String(bytes, StandardCharsets.US_ASCII));
Code unit, code point
As we have found out, the encoding standard defines the mapping symbol <-> number, and the encoding is the specific representation of this number in computer memory.
Code point is the numeric value of a character in Unicode, regardless of how it is represented. For example, for the character “A”, the code point is U+0041 (65 decimal).
Code unit is the minimum storage unit in a particular encoding (for example, 8 bits in UTF-8, 16 bits in UTF-16).
Depending on the encoding and on how large the number encoding the symbol can be, in computer memory it may be represented differently:
- Constant 4 bytes for all symbols.
- Here code point = code unit = 4 bytes.
- Two bytes for most symbols, but sometimes 2-byte pairs in case of complex symbols.
- Here, code point is either 2 bytes or 4 bytes, and code unit is always 2 bytes.
Now that we distinguish code points from code units, let’s move on to encodings.
UTF-8, UTF-16 and UTF-32
UTF-8
- The most popular encoding
- Variable number of bytes for each symbol
- That is, code point = code unit = 1…4 bytes.
- Symbols in the ASCII range (the first 128 Unicode symbols) take 1 byte.
- This means that UTF-8 is a superset of ASCII, and all ASCII symbols are valid in UTF-8.
- More complex symbols take 2, 3 or 4 bytes in UTF-8.
Example of a 1-byte symbol (capital Latin A):
- Unicode code point: U+0041 (decimal 65)
- UTF-8 representation: 01000001 (1 code point / 1 byte / 1 code unit)
Example of a 2-byte symbol (lowercase Cyrillic Ρ):
- Unicode code point: U+044F (decimal 1103)
- UTF-8 representation: 11010001 10001111 (1 code point / 2 bytes / 2 code units)
Example of a 3-byte symbol (summation sign β):
- Unicode code point: U+2211 (decimal 8721)
- UTF-8 representation: 11100010 10001000 10010001 (1 code point / 3 bytes / 3 code units)
Example of a 4-byte symbol (Gothic letter π):
- Unicode code point: U+10348 (decimal 66376)
- UTF-8 representation: 11110000 10010000 10001101 10001000 (1 code point / 4 bytes / 4 code units)
Let us note the structure of bytes in UTF-8:
1-byte symbols: 0xxxxxxx
2-byte symbols: 110xxxxx 10xxxxxx
3-byte symbols: 1110xxxx 10xxxxxx 10xxxxxx
4-byte symbols: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
Advantage: compatibility with ASCII, compactness for most texts (for example, English text will take up as much space as in ASCII).
UTF-16
- Uses 16-bit (2-byte) code units
- Most frequently used symbols (Basic Multilingual Plane) are encoded with one code unit (2 bytes)
- Symbols outside BMP (e.g. emoji, rare symbols) are encoded with a surrogate pair β two code units (4 bytes)
Thus, in UTF-16 one code point can be encoded with either one or two code units. But the code unit in UTF-16 is always 2 bytes.
UTF-16 is not directly compatible with ASCII. Although the lower byte of UTF-16 for all ASCII symbols will match ASCII encoding, UTF-16 requires a high byte.
If in ASCII, the code-point (and therefore code-unit) for the letter ‘A’ is 01000001, then in UTF-16 it is 00000000 01000001.
UTF-32
- Uses 32-bit (4-byte) code units
- Each Unicode character always takes exactly 4 bytes
- Here always 1 code point = 1 code unit = 4 bytes
- Advantage: simplicity and predictability in processing
Disadvantage: takes 2-4x more space than UTF-8 or UTF-16 for most texts.
What’s wrong with char in Java?
Java under the hood uses UTF-16, where code unit is always two bytes, and code point = either one code unit or two code units (surrogate pair).
Char always stores one code unit
The char (Character) type in Java actually stores not code point but code unit.
“By lucky coincidence” for most symbols this is enough to know the symbol, BUT NOT ALWAYS. Let’s dwell on this in more detail.
Java has methods to determine whether the current two bytes are a leading (a.k.a. high) code unit in a surrogate pair:
Character.isHighSurrogate();
Or a following (a.k.a. low) code unit in a surrogate pair:
Character.isLowSurrogate();
Char and emoji
For example, emoji in UTF-16 are represented by a surrogate pair of code units. Since we have already established that char = 1 code unit, this means that an emoji simply won’t fit in char!
char ch = 'π'; // ERROR: Too many characters in character literal
Generation of surrogate pair code units
How to get the high surrogate and low surrogate for an emoji? Extract a char[] array from a String:
String str = "π";
char[] arr = str.toCharArray();
char high = arr[0];
char low = arr[1];
We can confirm this:
Character.isHighSurrogate(high); // true
Character.isLowSurrogate(low); // true
When trying to print high or low separately to the screen, we will not see the expected result.
Assembling the symbol (code unit) from surrogate pair
How to “assemble” the emoji back? Use the String constructor:
System.out.println(new String(arr));
In our manipulations above Java implicitly uses UTF-16, as this encoding is used under the hood in Java.
Note that when converting byte[] to String you can specify the encoding:
new String(byteArr, StandardCharsets.UTF-8);
But when converting char[] to String β this possibility does not exist, because UTF-16 is used:
new String(charArr);
Note about encodings in Java / JVM
For full understanding, it is necessary to distinguish between the default encoding (what is used for IO and can be configured) and the internal encoding (always UTF-16, used when JVM works with strings).
Default encoding:
- Used for input/output operations such as reading from files or writing to streams.
- Depends on the locale settings of the operating system or can be explicitly set using JVM parameters (e.g. ‘-Dfile.encoding=UTF-8’).
- May differ in different environments and configurations.
Internal encoding:
- Used to represent String objects at runtime in Java.
- Always UTF-16, regardless of system or JVM settings.
- Uniform in all Java environments, ensuring consistent String data handling.
Converting character to byte[] and BOM (Byte Order Mark)
String <-> char[] is always UTF-16, but when converting String <-> byte[] the default encoding is UTF-8!
// don't specify encoding
byte[] bytes = "abc123π".getBytes();
// specify UTF-8, everything works
System.out.println(new String(bytes, StandardCharsets.UTF_8));
You can convert a string (character) to bytes in UTF-16. But first it is important to note that StandardCharsets class has three constants corresponding to UTF-16:
- ‘StandardCharsets.UTF_16’
- ‘StandardCharsets.UTF_16BE’
- ‘StandardCharsets.UTF_16LE’
If we convert a single character to a byte array:
byte[] bytes = "b".getBytes(StandardCharsets.UTF_16);
the expected size of byte[] will be 2 bytes, since code unit in UTF-16 is 2 bytes.
But if we check it:
byte[] bytes = "b".getBytes(StandardCharsets.UTF_16);
System.out.println(bytes.length); // output: 4
It turns out to be four. Why?
When using ‘StandardCharsets.UTF_16’, the resulting byte array will contain a byte order mark (BOM) at the beginning, which is typically 2 bytes (either ‘FE FF’ or ‘FF FE’ depending on byte order).
Thus, for one symbol we actually get 4 bytes:
- 2 bytes for specification
- 2 bytes for the symbol
And when using ‘StandardCharsets.UTF_16BE’ (big endian) or ‘StandardCharsets.UTF_16LE’ (little endian), the resulting array size is as expected β 2 bytes:
// char (2 bytes) to 2 bytes in UTF-16BE <- big endian, high comes first
// using String
byte[] bytes = new String(new char[] {ch}).getBytes(StandardCharsets.UTF_16BE);
// char (2 bytes) to 2 bytes in UTF-16LE <- little endian, low comes first
// using String
byte[] bytes = new String(new char[] {ch}).getBytes(StandardCharsets.UTF_16LE);
But that’s not all yet
How to count the number of characters in a string? ‘string.length()’?
Strings in Java are char[] arrays. And ‘string.length()’ is actually the length of the char[] array under the hood.
Therefore the length of such a string will be 2:
System.out.println("π".length()); // output: 2
If you read the official javadoc for ‘string.length()’:
The length is equal to the number of Unicode code units in the string.
Everything falls into place. But how to count the string length after all?
Since each symbol is always encoded by a number (code point), to count the number of symbols in a string, it is enough to find out the number of code points.
So the correct answer to the question is:
String str = "π";
System.out.println(str.codePoints().count());
Java 9+ Compact Strings
Before Java 9, all String objects in Java stored characters in UTF-16 format, using two bytes per character regardless of actual content.
In Java 9 a new Compact Strings optimization appeared, which dynamically chooses between Latin-1 encoding (one byte) and UTF-16 (two bytes) depending on the string content.
- JVM analyzes all characters in the string
- The question is asked: “Can each character in this string be represented using only Latin-1?” (characters with code points 0-255)
- If YES β Latin-1 encoding is used (1 byte per character)
- If NO β UTF-16 encoding is used (2 bytes per character)
If the string contains emoji, then Java 9+ will use UTF-16 (code unit = 2 bytes) instead of Latin-1 (code unit = 1 byte).
Bonus β converting char to a single byte
We already know that:
- Java uses UTF-16 under the hood for strings
- Char stores one code unit
- Code unit for UTF-16 is always 2 bytes
- So char size is 2 bytes
- For some symbols you need two chars, not one
- ASCII symbol code points (English letters, digits and basic special characters) match the code points in Unicode
And this means that basic symbols (ASCII) such as letters (e.g. b) and digits (e.g. 3), even when represented in UTF-16, fit in one byte:
byte b = (byte) 'a';
That is, the UTF-16 representation of the symbol b:
0000 0000 0110 0010
Fits into a byte variable due to truncation of the high byte:
0110 0010
Bonus code unit / code point / char
int -> char
// int (code unit) -> char
char ch = (char) 97;
// int (code point) -> char[]
int codePoint = 0x1F60A; // π
char[] chars = Character.toChars(codePoint);