Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 21 for “"Annotated dataset"”.
-
High throughput livestock monitoring using computer vision
… monitoring of livestock. We first consider a pig dataset acquired in controlled lab environments and apply state-of-the-art action recognition methods to it. The 'Pig Novelty Preference Behavioral Dataset' is used to train and validate the performance of different models. We obtain an accuracy of …
-
Information extraction with neural networks
… time-consuming to develop and fine-tune for new datasets. In this thesis, we propose the first de-identification system based on artificial neural networks (ANNs), which achieves state-of-the-art results without any human-engineered features. The ANN architecture is extended to incorporate …
-
A Comprehensive Analysis Of Straight Line Estimation With A Novel Noise Dataset
… detectors. This thesis presents a novel noise dataset called Artificially Generated Objects with Noise (AGON), with the goal of advancing the state-of-the art in straight line segment detection and noise corruption research. Generating 27,195 synthetic images under 11 noise model distributions …
-
Grounded SCAN Human: A Benchmark for Zero-Shot Generalizations
In this work, we collect a new human annotated dataset called Grounded SCAN Human (gSCAN Human) as an extension of the original Grounded SCAN (gSCAN) dataset. The original gSCAN dataset was created to test various compositional generalizations by holding out certain examples during train time. …
-
Joint Biomedical Event Extraction and Entity Linking via Iterative Collaborative Training
… these two tasks together as there is no existing dataset that contains annotations for both tasks. To solve these challenges, we propose joint biomedical entity linking and event extraction by regarding the event structures and entity references in knowledge bases as latent variables and updating …
-
Visual and auditory scene parsing
… benchmark based on a large scale, densely annotated dataset ADE20K. This benchmark, together with the state-of-the-art models we open source, offers a powerful tool for the research community to solve semantic and instance segmentation tasks. Then I investigate the challenge of parsing a …
-
Self-training for cyberbully detection: Achieving high accuracy with a balanced multi-class dataset
… involves the meticulous curation of a balanced dataset specifically designed for training the ML/ DL models. To overcome the challenge of limited labeled data, we employ a semi-supervised self-training algorithm, which effectively expands the size of the labeled dataset. By leveraging real-world …
-
Unsupervised summarization of public talk radio
… I create a novel spoken opinion summarization dataset consisting of compressed versions of "representative," opinion-containing utterances extracted from a hand-curated and crowd-source-annotated dataset of 275 snippets. I use this evaluation dataset to show that my model quantitatively …
-
Drivable Area Segmentation on LiDAR Range View Images for Autonomous Driving
… and inference speed. Because no LiDAR-based datasets for road segmentation are available, both training and evaluation are based both on the BDD100K dataset and on a proprietary dataset developed specifically for this thesis. To construct it, a hybrid annotation pipeline is introduced, …
-
Mapping Informality: An Approach to Classifying Sidewalk Informal Practices and Elements Through Street View Imagery
… this thesis also developed a taxonomy and annotated dataset of informality which was used to reveal spatial inequities in sidewalk use. By converting curbside complexity into structured, updateable categories, the framework enables planners to recognize the adaptive value of informal …
-
Metagenomic approaches for examining the diversity of large DNA viruses in the biosphere
… binning tool) in recovering viral genomes using annotated dataset. We used a metagenome simulator (CAMISIM) to generate simulated short reads with known composition to assess these processes. Moreover, I emphasized the importance of binning contigs for viral genomes to fully recover the genomes …
-
COCO-Bridge: Common Objects in Context Dataset and Benchmark for Structural Detail Detection of Bridges
… intelligence (AI) platforms. COCO-Bridge is an annotated dataset which can be trained using a convolutional neural network (CNN) to identify specific structural details. Many annotated datasets have been developed to detect regions of interest in images for a wide variety of applications and …
-
Ανάλυση συναισθήματος προφορικού λόγου από αθλητικές εκπομπές
… ενός επαρκώς σχολιασμένου συνόλου δεδομένων (annotated dataset) και την εκπαίδευση βασικών μοντέλων μηχανικής μάθησης (Machine Learning) για την αυτόματη αναγνώριση και ταξινόμηση συναισθημάτων. Αρχικά, έγινε η διαδικασία συλλογής και σύνθεσης ενός σώματος κειμένων (corpus), αποτελούμενου από …
-
Knowledge base integration in biomedical natural language processing applications
… processing in the biomedical field, the lack of annotated data due to regulations and expensive labor remains an issue. In this work, we study the potential of knowledge bases for biomedical language processing to compensate for the shortage of annotated data. Accordingly, we experiment with the …
-
Vision-based human action recognition using machine learning techniques
… Network model already trained on large-scale annotated dataset is transferable to action recognition task with limited training dataset. The comparative analysis also confirms its superior performance over handcrafted feature-based methods in terms of accuracy on same datasets. The second …
-
Automated Image-Based Inspection of Masonry Arch Bridges
… This has involved the creation of a pixel-wise annotated dataset of masonry arch bridge surfaces for both different defect classes and mortar joints. This dataset is believed to be unparalleled in both scope and scale, compared to other works in the literature and therefore serves as an …
-
Automatic identification of representative content on Twitter
… 22 different election topics on a manually annotated dataset. Tweet2Vec outperformed state-of-the- art algorithms on widely used semantic relatedness and sentiment classification evaluation tasks. To demonstrate the value of the framework, we analyzed tweets leading up to a primary debate …
-
Sentence Simplification for Text Processing
… linking and bounding functions. I present the annotated resources used to train and evaluate this sign tagger (Chapter 2) and the machine learning method used to implement it (Chapter 3). The second syntactic analysis method exploits the sign tagger and identifies the spans of compound clauses …
-
Biographical information extraction: A language-agnostic methodology for datasets and models
… using a specific architecture and a specific annotated dataset. These specific datasets typically aim to represent common patterns that the model is to learn, albeit at the cost of manual annotation, which can be costly and time-consuming. In addition, due to the nature of the training …
-
Extracting Coronary Lesion Information from Angiogram Reports for Patient Screening Applications
… patient research. We collected and curated a dataset of 72 diagnostic coronary angiogram reports from health systems which contributed data to the Abiomed cVAD registry. Of these, 39 reports from 6 sites were used as a training set and 13 reports from the same 6 sites as a development set for …
Page 1 of 2