← Back to Research

Hierarchical Transformers Are More Efficient Language Models

Findings of NAACL 2022

Hourglass hierarchical Transformer architecture

Hourglass is a hierarchical Transformer that shortens and expands intermediate representations using explicit downsampling and upsampling layers. It studies how hierarchy can reduce computation while retaining long-range sequence modelling capacity.

The model improves efficiency over standard Transformer baselines on ImageNet32 generation and the enwik8 language-modelling benchmark.

Citation

Nawrot, P., Tworkowski, S., Tyrolski, M., Kaiser, Ł., Wu, Y., Szegedy, C., & Michalewski, H. (2022). Hierarchical Transformers Are More Efficient Language Models. Findings of the Association for Computational Linguistics: NAACL 2022, 1559–1571. https://doi.org/10.18653/v1/2022.findings-naacl.117