AI Language Models Modified to Process Text Byte by Byte
Researchers at Ludwig-Maximilians-Universität München (LMU) have developed a method to convert existing large language models (LLMs) into byte-level models using less than one percent of the typical training effort.

Researchers at Ludwig-Maximilians-Universität München (LMU) have developed a new method that enables the conversion of existing artificial intelligence language models, such as those powering ChatGPT, to process text byte by byte. This "byteification" process requires a fraction of the standard training time and enhances the models' ability to handle text at a granular level.
Most current language models break down text into words or word fragments before processing. This method, known as subword tokenization, can limit the models' understanding of individual letters or characters. The new technique allows models to operate at the byte level, which roughly corresponds to processing single characters.
The research, published in the journal Nature, was conducted in collaboration with the Allen Institute for AI, the University of Cambridge, the University of Washington, and Imperial College London. The "byteification" process retains most of the original model's performance while adding the advantages of byte-level processing. These new models have proven to outperform previous byte-level models of comparable size.
Valentin Hofmann, Junior Professor at LMU and lead of the study, believes this method can address long-standing deficiencies in LLM models. "Conventional LLM models struggle with tasks requiring character-level abilities, such as spelling a word backward. Our 'byteified' models are significantly better at this," Hofmann stated. The models, code, and training data have been made publicly available.