From Markov to Laplace: How Mamba In-Context Learns Markov Chains
ICLROral2026
TL;DR
We uncover an interesting phenomenon where a single-layer Mamba represents the Bayes optimal Laplacian smoothing estimator when trained on Markov chains and we demonstrate it theoretically and empirically.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
mamba