Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 2 of 2 for “"efficient attention"”.
-
Random Features for Efficient Attention Approximation
… of the softmax kernel appearing in self-attention, the main component in the Transformer backbone. The obtained efficient Transformer architecture is referred to as Performer. Compared to other developments in the area of efficient Transformers, Performer is theory-grounded and the …
-
Attention-based representation learning on graphs
… of graph learning operators. Instead, I propose attention as a universal mechanism capable of augmenting and even superseding current architectural choices. My first contribution targets graph neural networks and consists of replacing classical readout functions with neural network-based adaptive …