Skip to main content
Back to problems
#393
Medium Algorithms

Utf 8 validation

Array Bit Manipulation
46.1% acceptance
Jan 13, 2026
946
2888
Given an integer array data representing the data, return whether it is a valid UTF-8 encoding (i.e. it translates to a sequence of valid UTF-8 encoded characters). A character in UTF8 can be from 1 to 4 bytes long, subjected to the following rules: For a 1-byte character, the first bit is a 0, followed by its Unicode code. For an n-bytes character, the first n bits are all one's, the n + 1 bit is 0, followed by n - 1 bytes with the most significant 2 bits being 10. This is how the UTF-8 encoding would work: Number of Bytes | UTF-8 Octet Sequence | (binary) --------------------+----------------------------------------- 1 | 0xxxxxxx 2 | 110xxxxx 10xxxxxx 3 | 1110xxxx 10xxxxxx 10xxxxxx 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx x denotes a bit in the binary form of a byte that may be either 0 or 1. Note: The input is an array of integers. Only the least significant 8 bits of each integer is used to store the data. This means each integer represents only 1 byte of data.

Solution

Rust
Time O(n²)
Space O(1)
LeetCode
solution.rs
impl Solution {
  pub fn valid_utf8(data: Vec<i32>) -> bool {
    let mut i = 0;
    while i < data.len() {
      let byte = data[i] as u8;
      let num_bytes = if byte & 0b10000000 == 0 {
        1
      } else if byte & 0b11100000 == 0b11000000 {
        2
      } else if byte & 0b11110000 == 0b11100000 {
        3
      } else if byte & 0b11111000 == 0b11110000 {
        4
      } else {
        return false;
      };
      
      if i + num_bytes > data.len() {
        return false;
      }
      
      for j in 1..num_bytes {
        let continuation = data[i + j] as u8;
        if continuation & 0b11000000 != 0b10000000 {
          return false;
        }
      }
      
      i += num_bytes;
    }
    true
  }
}