Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 42 for “"Self-attention"”.

  1. Self-attention policy architectures for reinforcement learning under partial observability

    … which necessarily constitutes noise. We explore self-attention-based policy architectures as a solution to this problem, demonstrating their robustness under conditions of high partial observability on different rein-forcement learning benchmark tasks, and explore the advantages and disadvantages …

    cape-town Repository record for Self-attention policy architectures for reinforcement learning under partial observability (opens in a new tab)

  2. Molecular graph Self attention and graph convolution for drug discovery

    … undirected graphs and use graph convolutions and self-attention to predict molecular properties. With a series of ablation studies, we demonstrate the added value of several key components in our network. We analyze two standard datasets: BBBP, which includes classication data on whether molecules …

    mit Repository record for Molecular graph Self attention and graph convolution for drug discovery (opens in a new tab)

  3. Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms

    … by incorporating a cycling layer into self-attention mechanisms. Traditional techniques often struggle to capture the cyclical nature of time series data, impacting predictive model accuracy. By integrating a cycling layer, this thesis aims to enhance the ability of models to recognize …

    york Repository record for Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms (opens in a new tab)

  4. Random Features for Efficient Attention Approximation

    … approximation of the softmax kernel appearing in self-attention, the main component in the Transformer backbone. The obtained efficient Transformer architecture is referred to as Performer. Compared to other developments in the area of efficient Transformers, Performer is theory-grounded and the …

    cambridge Repository record for Random Features for Efficient Attention Approximation (opens in a new tab)

  5. An attention-enhanced student–teacher framework for structural and logical anomaly detection in industrial settings

    … logical anomaly detection scheme (c. 2024) with self-attention mechanisms, enabling more effective relational modeling. We demonstrate consistent improvements on the MVTec LOCO Anomaly Detection benchmark. Specifically, AeCSAD employs a global student network with self-attention for reasoning …

    uoit Repository record for An attention-enhanced student–teacher framework for structural and logical anomaly detection in industrial settings (opens in a new tab)

  6. Advanced modelling and analytics for effective change and anomaly detection in hyperspectral images.

    … Additionally, this study introduces a novel 2D self-attention module, leading to the development of two lightweight deep learning networks focused on extracting local spatial-spectral features for more accurate change detection. The first network, namely CBANet, integrates a cross-band feature …

    rgu Repository record for Advanced modelling and analytics for effective change and anomaly detection in hyperspectral images. (opens in a new tab)

  7. Computer-Aided Detection of Clinically Significant Prostate Cancer using Bi-Parametric Magnetic Resonance Imaging

    … for PCa detection. Furthermore, the standard self-attention mechanism used to build the Transformer is memory and computationally inefficient, which hurts its performance and limits its application in actual clinical practice. In this study, we proposed a hybrid segmentation network that …

    poli-torino Repository record for Computer-Aided Detection of Clinically Significant Prostate Cancer using Bi-Parametric Magnetic Resonance Imaging (opens in a new tab)

  8. Benchmarking Graph Transformers Toward Scalability for Large Graphs

    … on graph-structured data. In particular, the self-attention mechanism of GTs mitigates the fundamental limitations of over-squashing, over-smoothing, and limited expressiveness that GNNs face. Furthermore, like transformers used for natural language processing and computer vision, GTs have the …

    mit Repository record for Benchmarking Graph Transformers Toward Scalability for Large Graphs (opens in a new tab)

  9. Neural attentions for natural language understanding and modeling

    In this thesis, we explore the use of neural attention mechanisms for improving natural language representation learning, a fundamental concept for modern natural language processing. With the proposed attention algorithms, our model made significant improvements in both language modeling and …

    mit Repository record for Neural attentions for natural language understanding and modeling (opens in a new tab)

  10. Attention-Based Learning for Combinatorial Optimization

    … of model, the Decision Transformer, which is a Self-Attention Transformer architecture that was recently developed for training on reinforcement learning problems. To analyze the model, we structure the Traveling Salesman problem as a reinforcement learning problem and, by continuously varying …

    mit Repository record for Attention-Based Learning for Combinatorial Optimization (opens in a new tab)

  11. Machine learning based speech quality prediction

    … such as CNNs, LSTM networks, and Transformer/self-attention networks were combined and compared. It was found that a network with CNN, Self-Attention, and a proposed attention-pooling delivers the best single-ended speech quality predictions on the considered dataset. Furthermore, a …

    tu-berlin Repository record for Machine learning based speech quality prediction (opens in a new tab)

  12. Exploration of Techniques for Working with Sparse Data when Applying Natural Language Processing to Assist a Qualitative Data Analysis of a COVID-19 Open Innovation Community

    … Neural Networks (CNN), and particularly the Self-Attention model. The proposed framework, identified for its superior performance, demonstrates a noteworthy ability to process and interpret sparse qualitative data, surpassing both traditional approaches in its effectiveness. Furthermore, the …

    calgary Repository record for Exploration of Techniques for Working with Sparse Data when Applying Natural Language Processing to Assist a Qualitative Data Analysis of a COVID-19 Open Innovation Community (opens in a new tab)

  13. Neural Representation for 3D Building Reconstruction from Point Clouds

    … point embedding module, linearized multi-head self-attention layers, and reversible functions, all designed to reduce computation time and space complexities. Additionally, the architectural design of the PReFormer follows a ∇-shape, which improves (object) size-invariant feature extraction and …

    calgary Repository record for Neural Representation for 3D Building Reconstruction from Point Clouds (opens in a new tab)

  14. VADViT:Vision Transformer-Driven Memory Forensics for Malicious Process Detection and Explainable Threat Attribution

    … them using a Vision Transformer (ViT) with self-attention to enhance detection accuracy. We also introduce BCCC-MalMem-SnapLog-2025, a dataset logging process identifier (PID) for precise VAD extraction without dynamic analysis. Experimental results show 99% accuracy in binary classification …

    york Repository record for VADViT:Vision Transformer-Driven Memory Forensics for Malicious Process Detection and Explainable Threat Attribution (opens in a new tab)

  15. Improving LLM Long Context Understanding via Synthetic Data and Adaptive Compression

    … are constrained by the quadratic scaling of the self-attention mechanism, which restricts most popular LLMs to a context length of several thousand tokens. Many methods have been introduced to extend the context of LLMs, including the Activation Beacon approach. In this work, we propose two key …

    mit Repository record for Improving LLM Long Context Understanding via Synthetic Data and Adaptive Compression (opens in a new tab)

  16. Gated Transformer-Based Architecture for Automatic Modulation Classification

    … encoder architecture incorporating a multi-head self-attention mechanism. We train our architecture extensively across a diverse range of signal-to-noise ratios (SNRs) from the RadioML 2018.01A dataset. We introduce a novel transformer-based architecture with a gated mechanism, designed as a …

    vt Repository record for Gated Transformer-Based Architecture for Automatic Modulation Classification (opens in a new tab)

  17. Investigation of Future Voluntary Movement Prediction for Pathological Tremor-Alleviating Exoskeletons

    … on a convolutional neural network (CNN) and self-attention mechanism to accurately extract and predict patients' voluntary movement intentions from tremor-affected motion data. This model enables real-time motion planning for the exoskeleton, achieving both tremor suppression and zero-latency …

    vt Repository record for Investigation of Future Voluntary Movement Prediction for Pathological Tremor-Alleviating Exoskeletons (opens in a new tab)

  18. Human Mesh Recovery Using Radio Signals

    … strong and weak supervision, 2) a multi-headed self-attention mechanism that attends differently to temporal information in the radio signal, and 3) an adversarially trained temporal discriminator that imposes a prior on the dynamics of human motion. Our results show that RF-Avatar accurately …

    mit Repository record for Human Mesh Recovery Using Radio Signals (opens in a new tab)

  19. VRAG: Region Attention Graph for Content-Based Video Retrieval

    … In this paper, we introduce Video Region Attention Graph Networks (VRAG) that improves the state-of-the-art of video-level methods. We represent videos at a finer granularity viaregion-level features and encode video spatio-temporal dy-namics through region-level relations. Our VRAG …

    nus Repository record for VRAG: Region Attention Graph for Content-Based Video Retrieval (opens in a new tab)

  20. Hardware software co-design of machine learning accelerators using univariate functions

    … architectures. Next, we introduce Personal Self-Attention (PSA), a novel method for learning non- linear univariate functions. Using PSA with linear transformations, we demonstrate a 2×compression in hidden size of Multi-Layer Perceptrons (MLP), while matching accuracy. Applying this to an …

    umn Repository record for Hardware software co-design of machine learning accelerators using univariate functions (opens in a new tab)

Page 1 of 3