
Hourglass is a hierarchical Transformer that shortens and expands intermediate representations using explicit downsampling and upsampling layers. It studies how hierarchy can reduce computation while retaining long-range sequence modelling capacity.
The model improves efficiency over standard Transformer baselines on ImageNet32 generation and the enwik8 language-modelling benchmark.