Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 5813 for “"Dataset"”.

  1. Medical abstract inference dataset

    In this thesis, I built a dataset for predicting clinical outcomes from medical abstracts and their title. Medical Abstract Inference consists of 1,794 data points. Titles were filtered to include the abstract's reported medical intervention and clinical outcome. Data points were annotated with the …

    mit Repository record for Medical abstract inference dataset (opens in a new tab)

  2. Dataset Deduplication with Datamodels

    Large curated datasets have been essential to the development of deep learning models across many disciplines. Consequently, the properties of these datasets have a large impact on the behavior of these models. As machine learning pipelines increasingly leverage more unlabelled datasets—which tend …

    mit Repository record for Dataset Deduplication with Datamodels (opens in a new tab)

  3. Hierarchical Bayesian Dataset Selection

    … depends on access to large, high-quality datasets, which are often challenging to identify. To address this, we introduce <b>H</b>ierarchical <b>B</b>ayesian <b>D</b>ataset <b>S</b>election (<b>HBDS</b>), the first dataset selection algorithm that utilizes hierarchical Bayesian modeling, …

    vt Repository record for Hierarchical Bayesian Dataset Selection (opens in a new tab)

  4. Finding local experts from Yelp dataset

    … of its data for experimentation and we use this dataset to test our hypotheses. Finding local experts can be used in many ways, such as generating weighted, more accurate reviews for businesses, and to create recommendations for new users. If you visit Paris for the first time, you would now be …

    uiuc Repository record for Finding local experts from Yelp dataset (opens in a new tab)

  5. Privacy-Preserving Natural Language Dataset Generation

    … we propose the generation of private synthetic datasets to replace the original datasets in training and testing the model. These synthetic datasets will have the same semantic and statistical distribution as the original dataset, but will be differentially private, thus preventing individuals …

    mit Repository record for Privacy-Preserving Natural Language Dataset Generation (opens in a new tab)

  6. Analysis and Comparison of a Detailed Land Cover Dataset versus the National Land Cover Dataset (NLCD) in Blacksburg, Virginia

    … accuracy assessments on the National Land Cover Dataset (NLCD), little research has utilized a detailed digitized land cover dataset, like that available for the Town of Blacksburg, for this comparison. This study aims to evaluate the information available from a detailed land cover dataset and …

    vt Repository record for Analysis and Comparison of a Detailed Land Cover Dataset versus the National Land Cover Dataset (NLCD) in Blacksburg, Virginia (opens in a new tab)

  7. STREETS: a benchmark dataset for suburban traffic forecasting

    … and benchmark STREETS, a novel traffic flow dataset from publicly available web cameras in the suburbs of Chicago, IL. STREETS addresses multiple limitations of existing vehicular traffic datasets. Many current datasets lack a coherent traffic network graph to describe the relationship …

    uiuc Repository record for STREETS: a benchmark dataset for suburban traffic forecasting (opens in a new tab)

  8. Semantic amodal video segmentation using a synthetic dataset

    … modal instance-level object segmentation dataset. This new dataset provides ample opportunities to train models for instance-level segmentation, both modal and amodal. Moreover, in this work, we also present results for instance-level segmentation using ResNet-based DeepLab, a …

    uiuc Repository record for Semantic amodal video segmentation using a synthetic dataset (opens in a new tab)

  9. An empirical study identifying bias in Yelp dataset

    … behaviors in users' rating habits in the Yelp dataset. The Surprise recommender system is utilized to produce expected ratings for the test set, training the model with 75% of the original dataset to learn the rating trends. Then, the ordinary least squares (OLS) linear regression is applied to …

    mit Repository record for An empirical study identifying bias in Yelp dataset (opens in a new tab)

  10. Fueling Conflict: A Global Dataset of Energy Protests

    … results in the creation of the first global dataset on energy protests. This novel source of evidence, in turn, will open new avenues for research on the conflict-energy nexus, particularly on the impact of market shocks on civilian unrest and instability in low- and middle-income countries – …

    mit Repository record for Fueling Conflict: A Global Dataset of Energy Protests (opens in a new tab)

  11. Attention: not just another dataset for patch-correctness checking

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms

    uiuc Repository record for Attention: not just another dataset for patch-correctness checking (opens in a new tab)

  12. Dataset Interfaces: Diagnosing Model Failures Using Controllable Counterfactual Generation

    … In this work, we introduce the notion of a dataset interface: a framework that, given an input dataset and a user-specified shift, returns instances from that input distribution that exhibit the desired shift. We study a number of natural implementations for such an interface, and find that …

    mit Repository record for Dataset Interfaces: Diagnosing Model Failures Using Controllable Counterfactual Generation (opens in a new tab)

  13. Spoken ObjectNet: Creating a Bias-Controlled Spoken Caption Dataset

    Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained on that data. We introduce Spoken ObjectNet, which is …

    mit Repository record for Spoken ObjectNet: Creating a Bias-Controlled Spoken Caption Dataset (opens in a new tab)

  14. A Comparison of Imperviousness Derived from a Detailed Land Cover Dataset (DLCD) versus the National Land Cover Dataset (NLCD) at Two Time Periods

    … accuracy concerns of the National Land Cover Dataset (NLCD), this case study compares impervious surface from the NLCD to a Detailed Land Cover Dataset (DLCD) for the Town of Blacksburg, Virginia over two time periods (2005/2006 and 2011) at spatial aggregation scales (fine to coarse) and …

    vt Repository record for A Comparison of Imperviousness Derived from a Detailed Land Cover Dataset (DLCD) versus the National Land Cover Dataset (NLCD) at Two Time Periods (opens in a new tab)

  15. Understanding the human psychology towards relationships using the Reddit dataset

    There is an increasing interest in exploiting social media data to address problems in various domains, including social science and psychology. Reddit is a well-known social media platform that promotes user interactions on various topics and accumulates massive amounts of data daily. Reddit users …

    missouri Repository record for Understanding the human psychology towards relationships using the Reddit dataset (opens in a new tab)

  16. Object Detection on Unmanned Arial Vehicles Dataset Using Adaptive HydraNet

    … small objects on Unmanned Aerial Vehicle (UAV) datasets remains a significant challenge due to the limitations of the current backbone architecture of these methods. This limitation arises from the architecture's multilabel classification step, which lacks precision in detecting small objects …

    calgary Repository record for Object Detection on Unmanned Arial Vehicles Dataset Using Adaptive HydraNet (opens in a new tab)

  17. On the Role of the Source Dataset in Transfer Learning

    … suggests that removing data from the source dataset can actually help too. In this work, we take a closer look at the role of the source dataset's composition in transfer learning and present a framework for probing its impact on downstream performance. Our framework gives rise to new …

    mit Repository record for On the Role of the Source Dataset in Transfer Learning (opens in a new tab)

  18. Addressing two issues in machine learning : interpretability and dataset shift

    … by a large sacrifice in accuracy on real world datasets. I then briefly discuss possible extensions that allow one to directly optimize rank statistics over rule lists, and handle ordinal data. In the second, I address a shortcoming of a popular approach to handling covariate shift, in which the …

    mit Repository record for Addressing two issues in machine learning : interpretability and dataset shift (opens in a new tab)

Page 1 of 291