AI Language Models Can Now Process Text Letter by Letter
Researchers have developed a method to convert existing AI language models to process text at the byte level, enhancing their understanding of language structure with minimal retraining.

Researchers at Ludwig-Maximilians-Universität München (LMU) have developed a novel method that enables large language models (LLMs) to process text at the individual character (byte) level. Unlike traditional LLMs that segment text into words or sub-word units, missing direct character input, this new "byteification" process retrofits existing models. The procedure requires less than one percent of the typical retraining cost while largely preserving the original model's performance.
This advancement addresses a significant limitation in current LLMs, which struggle with tasks requiring character-level understanding, such as spelling words backward. The newly converted models demonstrate superior performance in these specific areas. This development is considered a crucial step toward achieving deeper language comprehension in artificial intelligence and is expected to unlock new avenues for research.
The study, a collaboration involving researchers from LMU, the Allen Institute for AI, the University of Cambridge, the University of Washington, and Imperial College London, has been published in the journal Nature. The resulting models, along with the code and training data, are publicly accessible, promoting open science and future development.
Valentin Hofmann, Junior Professor at LMU Munich and the study's last author, stated that byte-level LLMs have the potential to overcome long-standing shortcomings. He highlighted that a precise representation of low-level text structure is critical for many scientific applications, including working with code or biological sequences.