Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 72 for “"Hadoop"”.
-
DATA MINING: TRACKING SUSPICIOUS LOGGING ACTIVITY USING HADOOP
… could have on the system. This project uses Hadoop framework to search and store log files for logging activities and then performs a ‘Map Reduce’ programming code to finally compute and analyze the results.</p>
-
An Apache Hadoop Framework for Large-Scale Peptide Identification
… thesis describes the implementation of Apache Hadoop framework for large-scale peptide identification using C-Ranker. The Apache Hadoop data processing software is immersed in a complex environment composed of massive machine clusters, large data sets, and several processing jobs. The framework …
-
A Framework for Hadoop Based Digital Libraries of Tweets
… were designed to operate on top of the Hadoop and Spark libraries of tools. The first set of data structures is an abstract representation of a tweet at a basic level, as well as several concrete implementations which represent varying levels of detail to correspond with common sources …
-
Sharing the love : a generic socket API for Hadoop Mapreduce
<p>Hadoop is a popular software framework written in Java that performs data-intensive distributed computations on a cluster. It includes Hadoop MapReduce and the Hadoop Distributed File System (HDFS). HDFS has known scalability limitations due to its single NameNode which holds the entire file …
-
The application of the Hadoop software framework in Bioinformatics programs
… sequence de-duplication - by 1) adopting the Hadoop MapReduce algorithms and distributed file system and 2) implementing the fully automated Hadoop programs into a user friendly graphical user interface (GUI). In addition, the researcher was also interested in investigating the advantages and …
-
Data analytics with mapreduce in apache spark and hadoop systems
MapReduce comes from a traditional problem solving method: separating a big problem and solving each small parts. With the target of computing larger dataset in more efficient and cheaper way, this is implement into a programming mode to deal with massive quantity of data. The users get a map …
-
Improving Energy-Efficiency through Smart Data Placement in Hadoop Clusters
<p>Hadoop, a pioneering open source framework, has revolutionized the big data world because of its ability to process vast amounts of unstructured and semi-structured data. This ability makes Hadoop the ‘go-to’ technology for many industries that generate big data, thus it also aids in being cost …
-
P2PHDFS: AN IMPLEMENTATION OF STATISTIC MULTIPLEXED COMPUTING ARCHITECTURE IN HADOOP FILE SYSTEM
The Peer to Peer Hadoop Distributed File System (P2PHDFS) is designed to store and process extremely large-scale data sets reliably. This is a first attempt implementation of the Statistic Multiplexed Computing Architecture concept proposed by Dr. Shi for the existing Hadoop File System (HDFS) to …
-
Reusing software tests for configuration testing: A case study of the Hadoop project
… yield an effective rate of 41.7% and 48.8%, for Hadoop Common and HDFS, respectively. Further, this thesis conducts in-depth analysis on tests that cannot be effectively reused, categorizes code patterns of these tests that have false positives or negatives, and provides examples on rewriting …
-
The Tessera D&R computational environment: Designed experiments for R-Hadoop performance and Bitcoin analysis
… computational environment is RHIPE, the R and Hadoop Integrated Programming Environment. R is a widely used interactive language for data analysis. Hadoop is a distributed, parallel computational environment consisting of a distributed file system (HDFS) and distributed compute engine …
-
Application-Aware Network Design Using Software Defined Networking for Application Performance Optimization for Big Data and Video Streaming
… For applications, we consider two classes: Hadoop MapReduce and video streaming. The Hadoop MapReduce (M/R) framework has become the de facto standard for Big Data analytics. However, the lack of network-awareness of the default MapReduce resource manager in a traditional IP network can …
-
Cross-layer scheduling in cloud computing systems
… our framework in a batch-processing framework (Hadoop) and a stream-processing framework (Storm [2]). Our experimental results show that we are able to improve throughput of jobs in Storm and Hadoop by up to 32% and 29% respectively.
-
Phurti: application and network-aware flow scheduling for MapReduce
… the goal of decreasing the completion time for Hadoop MapReduce jobs. Phurti communicates both with the Hadoop framework to retrieve job-level network traffic information and the OpenFlow-based switches to learn about network topology. Phurti implements a novel heuristic called Smallest Maximum …
-
Platform leadership in open source software
… those challenges by analyzing the Android and Hadoop ecosystems through an augmented version of Porter's Five Forces framework proposed by Intel's Andrew Grove. The analysis finds that platform contenders in open source behave differently depending on whether they focus on competing against …
-
Performance optimization and energy efficiency of big-data computing workflows
… by extensive simulation-based results in Hadoop/YARN in comparison with existing workflow mapping models and algorithms. Considering that large-scale workflows for big data analytics have become a main consumer of energy in data centers, this dissertation also delves into the problem of …
-
Implementation of digital forensics readiness in big data wireless medical networks
… supported by an automated, user-friendly Linux-Hadoop Forensics Extractor (LHFX) tool tailored for big data Linux-Hadoop environments. Two prototypes were developed: a partial implementation of the framework’s big data DFR (BdDFR) environment and a fully functional LHFX tool. The proposed …
-
Performance models for parallel applications under failures
… of the dissertation evaluates the performance of Hadoop MapReduce applications, with different execution parameters and under different failure scenarios. The dissertation introduces performance models for Hadoop MapReduce applications considering node and process failures. Having a performance …
-
Reducing Cluster Power Consumption by Dynamically Suspending Idle Nodes
… center clusters. Then, I installed a version of Hadoop modified with a novel power management system on the cluster. The power management system uses different algorithms to determine when to turn off idle nodes in the cluster. Using the experimental cluster running a modified Hadoop …
-
What about big data?
… the two major current proposed solutions: Apache Hadoop and Apache Spark. As we will see, both are open source big data processing frameworks, but each one uses the batch processing solution in a different manner.
-
Exploiting Heterogeneity in Distributed Software Frameworks
… Software Frameworks (DSFs), such as MapReduce, Hadoop, Dryad, and Pregel, for supporting data-intensive scientific and enterprise applications on emerging heterogeneous compute, storage and network infrastructure. Large DSF deployments in the cloud continue to grow both in size and number, given …
Page 1 of 4