Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations

NeurIPSSpotlight2025

Authors
Brian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase, Yejin Choi, Noah A. Smith
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Modern tokenizers employ deterministic algorithms to map text into a single ``canonical" token sequence, yet the same string can be encoded as many non-canonical tokenizations using the language model…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive