Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 55 for “"data cleaning"”.
-
Data Cleaning Framework: An Extensible Approach to Data Cleaning
… rise in errors in this information. The stored data can be critically important, necessitating new ways of correcting anomalous records. Current cleaning techniques are very domain-specific and hard to extend, hindering their use in some areas. This work proposes an extensible framework for data …
-
Data cleaning techniques for software engineering data sets
Data quality is an important issue which has been addressed and recognised in research communities such as data warehousing, data mining and information systems. It has been agreed that poor data quality will impact the quality of results of analyses and that it will therefore impact on decisions …
-
A conceptual model for transparent, reusable, and collaborative data cleaning
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms
-
PClean : Bayesian data cleaning at scale with domain-specific probabilistic programming
Data cleaning is naturally framed as probabilistic inference in a generative model, combining a prior distribution over ground-truth databases with a likelihood that models the noisy channel by which the data are filtered, corrupted, and joined to yield incomplete, dirty, and denormalized datasets. …
-
Human Versus Computer Algorithmic Measurements of Caloric Response: Implications for Test Analysis
… to bithermal caloric test (BCT) analysis by data cleaning and determine how variation differs between examiners. Methods: Analysis of 435 consecutive BCTs performed by 6 examiners using identical protocols on adults with dizziness. Outcomes of total eye speed (TES) and unilateral weakness …
-
Data curation with ontology functional dependences
Poor data quality has become a pervasive issue due to the increasing complexity and size of modern datasets. Functional dependencies have been used in existing cleaning solutions to model syntactic equivalence. They are not able to model semantic equivelence, however. We advance the state of data …
-
Can an LLM find its way around a Spreadsheet?
… contexts, and one of the most vexing challenges data analysts face is performing data cleaning prior to analysis and evaluation. The ad-hoc and arbitrary nature of data cleaning problems, such as typos, inconsistent formatting, missing values, and a lack of standardization, often creates the need …
-
New techniques for improving biological data quality through information integration
As databases become more pervasive through the biological sciences, various data quality concerns are emerging. Biological databases tend to develop data quality issues regarding data legacy, data uniformity and data duplication. Due to the nature of this data, each of these problems is non-trivial …
-
Uma metodologia para tratamento de dados de curvas de carga baseada em técnicas de inteligência artificial
Data quality is critical in the short-term load forecasting. Frequently, load data show aberrant values (outliers), discontinuities, and gaps (missing data) caused by the abnormal operation of the electrical system or failures and problems in the measurement system. The presence of corrupted data …
-
Organizing historical agricultural data and identifying data integrity zones to assess agricultural data quality
… transitions into decision agriculture, data driven decision- making has become the focus of the industry and data quality will be increasingly important. Traditionally, yield data cleaning techniques have removed individual data points based on criteria primarily focused on the yield …
-
EFFICIENT DATA CURATION AND UTILIZATION FOR DEEP LEARNING
… construction efficiency of large-scale vision datasets, aiming to reduce computational and annotation costs. We propose InfoBatch, an unbiased dynamic data pruning framework that losslessly accelerates training and saves 20–40% of computation across diverse vision tasks. To address dataset …
-
Examination timetabling at the University of Cape Town: a tabu search approach to automation
… on the UCT November 2014 examination timetabling data with tabu search proving to be more effective, capable of producing feasible solutions from randomised initial solutions. To make this research more accessible, a user-friendly app was developed which showcased the optimisation techniques in a …
-
The Perception of Partner’s Pornography Use as a Betrayal: The Role of Trust, Investment, Commitment, and Forgiveness
… (N = 49). The final sample size after the data cleaning procedure was N = 135, which was further separated into two subsamples, the Betrayed Subsample (n = 47) and the Not Betrayed Subsample (n = 86). The results of the study revealed that investment and commitment had significant negative …
-
Work-family conflict and organisational commitment amongst fathers in the South African National Defence Force
… in the South African National Defence Force. Data was collected using a paper-based survey. After data cleaning, there were 132 usable questionnaires from uniformed members of the SA National Defence Force (9 SAI) based in Cape Town.The correlation analysis revealed no significant …
-
Overdue invoice forecasting and data mining
… The main procedures of the research work are data cleaning and processing, statistical analysis, building machine learning models and evaluating model performance. The analytical and modeling of the study are based on the real-world invoice data from a Fortune 500 company. The thesis also …
-
Contour Extraction of Drosophila Embryos Using Active Contours in Scale Space
… applied on the images to refine embryo contours. Data cleaning methods are applied to smooth the jaggy contours caused by blurred embryo boundaries. The scale space theory is applied to improve the performance of the result. The active contour adjusts better to the object for finer scales. The …
-
Workload-aware compressed linear algebra for data-centric machine learning pipelines
Compression is an effective technique for fitting data in available memory, reducing I/O across the storage-memory-cache hierarchy, decreasing energy consumption, and increasing instruction parallelism. Modern machine learning (ML) systems exploit the approximate nature of ML and mostly use lossy …
-
Machine learning and deep learning techniques for natural language processing with application to audio recordings
… companies need to rely on research focusing on data analysis methods that can assist them to analyse their unstructured data which holds information that could help them to better assign their collection agents to high repayment probable accounts. These types of accounts are characterised by the …
-
Psychedelic Use and Psychological Flexibility: The Role of Decentering, Mystical Experiences, Ego-Dissolution, and Insight
… recruited from social media. The sample after data cleaning was N = 427, however, following pairwise deletions the final sample size across analyses ranged from 114 to 149. The results of this study revealed that both self-perceived meaningful intention and feelings of comfort/safety were …
-
Public policy and barriers influencing SMEs' market expansion
… web-based questionnaire was used for collecting data. To ensure quality results, the data collected from 178 managers of formal manufacturing SMEs was reduced to 79 through a rigorous data cleaning process. The multiple linear regression test results suggest that South African SMEs are still …
Page 1 of 3