Why Is Attention Sparse In Particle Transformer?

Timothy Legge; Aaron Wang; Jacob Ortiz; Victor Limouzi; Zihan Zhao; Abhijith Gandrakota; Elham E. Khoda; Jennifer Ngadiuba; Javier Duarte; Richard Cavanaugh

doi:10.48550/arXiv.2512.00210

← Recent

AG-2025.11-1573·hep-ph·cross-listed: hep-exphysics.data-an

Why Is Attention Sparse In Particle Transformer?

Authors

Timothy Legge
Aaron Wang
Jacob Ortiz
Victor Limouzi
Zihan Zhao
Abhijith Gandrakota
Elham E. Khoda
Jennifer Ngadiuba
Javier Duarte
Richard Cavanaugh

Abstract

Transformer-based models have achieved state-of-the-art performance in jet tagging at the CERN Large Hadron Collider (LHC), with the Particle Transformer (ParT) representing a leading example of such models. A striking feature of ParT is its sparse, nearly binary, attention structure, raising questions about the origin of this behavior and whether it encodes physically meaningful correlations. In this work, we investigate the source of ParT's sparse attention by comparing models trained on multiple benchmark datasets and examine the relative contributions of the attention term and the physics-inspired interaction matrix before softmax. We find that binary sparsity arises primarily from the attention mechanism itself, with the interaction matrix playing a secondary role. Moreove, we show that ParT is able to identify key jet substructure elements, such as leptons in semileptonic top decays, even without explicit particle identification inputs. These results provide new insight into the interpretability of transformer-based jet taggers and clarify the conditions under which sparse attention patterns emerge in ParT.

Submitted

28 November 20258 months ago

Version

v1

License

CC-BY-4.0

DOI

10.48550/arXiv.2512.00210

Cite this preprint

BibTeX RIS

Imports into BibLaTeX, Zotero, Mendeley, EndNote.

PDF

Open PDF

Opens in a new tab · v1.

Chat with this PDF

Ask questions, probe assumptions, request a plain-English summary. Answers cite sections from the preprint itself.

Community

Questions and answers about this paper from other readers. No formal peer review — just a place to think out loud.