#393
Medium Algorithms Utf 8 validation
Array Bit Manipulation
46.1% acceptance
Jan 13, 2026
946
2888
Given an integer array data representing the data, return whether it is a valid UTF-8 encoding (i.e. it translates to a sequence of valid UTF-8 encoded characters).
A character in UTF8 can be from 1 to 4 bytes long, subjected to the following rules:
For a 1-byte character, the first bit is a 0, followed by its Unicode code.
For an n-bytes character, the first n bits are all one's, the n + 1 bit is 0, followed by n - 1 bytes with the most significant 2 bits being 10.
This is how the UTF-8 encoding would work:
Number of Bytes | UTF-8 Octet Sequence
| (binary)
--------------------+-----------------------------------------
1 | 0xxxxxxx
2 | 110xxxxx 10xxxxxx
3 | 1110xxxx 10xxxxxx 10xxxxxx
4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
x denotes a bit in the binary form of a byte that may be either 0 or 1.
Note: The input is an array of integers. Only the least significant 8 bits of each integer is used to store the data. This means each integer represents only 1 byte of data.
Solution
Rust
Time O(n²)
Space O(1)
impl Solution {
pub fn valid_utf8(data: Vec<i32>) -> bool {
let mut i = 0;
while i < data.len() {
let byte = data[i] as u8;
let num_bytes = if byte & 0b10000000 == 0 {
1
} else if byte & 0b11100000 == 0b11000000 {
2
} else if byte & 0b11110000 == 0b11100000 {
3
} else if byte & 0b11111000 == 0b11110000 {
4
} else {
return false;
};
if i + num_bytes > data.len() {
return false;
}
for j in 1..num_bytes {
let continuation = data[i + j] as u8;
if continuation & 0b11000000 != 0b10000000 {
return false;
}
}
i += num_bytes;
}
true
}
}