
This work analyses weaknesses in existing MHC-I presentation benchmarks, including train–test overlap and limited tests of generalisation to unseen peptides and alleles. It introduces a stricter benchmark and HLABERT, a pretrained Transformer model for pan-specific MHC-I presentation prediction.
An extended abstract was presented at the Machine Learning for Drug Discovery Workshop at ICLR 2023. The canonical peer-reviewed publication is the final Methods article.