Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 17 of 17 for “"map-reduce"”.

  1. Non-parametric Bayesian models for structured output prediction

    … structured output problems. We first study a map-reduce implementation of a stochastic inference method designed for the infinite hidden Markov model, applied to a computational linguistics task, part-of-speech tagging. We show that mainstream map-reduce frameworks do not easily support highly …

    cambridge Repository record for Non-parametric Bayesian models for structured output prediction (opens in a new tab)

  2. An Approach for Fast Score Computation in Bayesian Network Structure Learning Over Large-Scale Distributed Data

    … We show that DiSC can significantly outperform map-reduce style score computations executed by the distributed computation framework Apache Spark on a variety of synthetic and real datasets with a low accuracy trade-off.

    umkc Repository record for An Approach for Fast Score Computation in Bayesian Network Structure Learning Over Large-Scale Distributed Data (opens in a new tab)

  3. Data cleansing and integration operators for a parallel data analytics platform

    The data quality of real-world datasets need to be constantly monitored and maintained to allow organizations and individuals to reliably use their data. Especially, data integration projects suffer from poor initial data quality and as a consequence consume more effort and money. Commercial …

    potsdam-diss Repository record for Data cleansing and integration operators for a parallel data analytics platform (opens in a new tab)

  4. DATA MINING: TRACKING SUSPICIOUS LOGGING ACTIVITY USING HADOOP

    … for logging activities and then performs a ‘Map Reduce’ programming code to finally compute and analyze the results.</p>

    csusb Repository record for DATA MINING: TRACKING SUSPICIOUS LOGGING ACTIVITY USING HADOOP (opens in a new tab)

  5. Low latency data retrieval solutions for big data

    … Caching approach, we analyze real world map reduce traces from Facebook and Yahoo in terms of file size and access pattern distribution. We combine it with the existing analysis from literature to develop a new caching algorithm that builds on top of the HDFS caching API recently …

    uiuc Repository record for Low latency data retrieval solutions for big data (opens in a new tab)

  6. A new Internet Naming System

    … It is well known that DNS problems lead to reduced availability of Internet-based services in many different ways. In this thesis, I present four main results. All of them contribute to improvements and deeper understanding of DNS’ dependability issues. First, I discuss, how well established …

    qucosa-diss

  7. Enriching the Web of Data with topics and links

    This thesis presents novel ideas and research findings for the Web of Data – a global data space spanning many so-called Linked Open Data sources. Linked Open Data adheres to a set of simple principles to allow easy access and reuse for data published on the Web. Linked Open Data is by now an …

    potsdam-diss Repository record for Enriching the Web of Data with topics and links (opens in a new tab)

  8. Distributed graph decomposition algorithms on Apache Spark

    … of our algorithm with state-of-the-art Map-Reduce and parallel algorithms using openly available real world network data. Our proposed algorithms have shown substantial performance improvement.

    iupui Repository record for Distributed graph decomposition algorithms on Apache Spark (opens in a new tab)

  9. Privacy preservation in social media environments using big data

    "With the pervasive use of mobile devices, social media, home assistants, and smart devices, the idea of individual privacy is fading. More than ever, the public is giving up personal information in order to take advantage of what is now considered every day conveniences and ignoring the …

    must-thes Repository record for Privacy preservation in social media environments using big data (opens in a new tab)

  10. Performance optimization and energy efficiency of big-data computing workflows

    … of moldable parallel computing jobs, such as MapReduce, with intricate inter-job dependencies. The granularity of task partitioning in each moldable job of such big data workflows has a significant impact on workflow completion time, energy consumption, and financial cost if executed in …

    njit Repository record for Performance optimization and energy efficiency of big-data computing workflows (opens in a new tab)

  11. A scientific workflow framework for scientific data querying and processing

    … and a set of workflow constructs, including Map, Reduce, Tree, Loop, Conditional, and Curry, which are fully compositional one with another. Our workflow composition framework is unique in that workflows are the only operands for composition; in this way, our approach elegantly solves the …

    wayne-thes Repository record for A scientific workflow framework for scientific data querying and processing (opens in a new tab)

  12. On Ranked Approximate Matching Of Large Attributed Graphs

    … as it can be seamlessly transformed to adhere to map-reduce framework.</p> <p>We have conducted thorough experiments on several synthetic and real data sets, and have demonstrated the effectiveness and efficiency of the proposed method.</p>

    wayne-thes Repository record for On Ranked Approximate Matching Of Large Attributed Graphs (opens in a new tab)

  13. High dimensional revenue management

    … for distributed computation referred to as 'Map-Reduce'. An implementation of our solver in a shared memory environment where we can benchmark against solvers such as CPLEX shows that the algorithm outperforms those solvers on the types of LPs that an ad network would have to solve in …

    mit Repository record for High dimensional revenue management (opens in a new tab)

  14. Improving the Prediction Accuracy of Text Data and Attribute Data Mining with Data Preprocessing

    <p>Data Mining is the extraction of valuable information from the patterns of data and turning it into useful knowledge. Data preprocessing is an important step in the data mining process. The quality of the data affects the result and accuracy of the data mining results. Hence, Data preprocessing …

    kennesaw Repository record for Improving the Prediction Accuracy of Text Data and Attribute Data Mining with Data Preprocessing (opens in a new tab)

  15. Image classification and feature selection

    Made available in DSpace on 2012-06-27T21:22:52Z (GMT). No. of bitstreams: 9 Chen_Gang.pdf: 2139361 bytes, checksum: 3e14f03bb001785a473bc6193dc54f8b (MD5) license.txt: 4058 bytes, checksum: 50cddeb1b191bb1536258b22551f35da (MD5) ch4.tex: 87063 bytes, checksum: 8fac0f031e2c4b465952bb040f6131fa …

    uiuc Repository record for Image classification and feature selection (opens in a new tab)

  16. On the Feasibility of MapReduce to Compute Phase Space Properties of Graphical Dynamical Systems: An Empirical Study

    … challenges. To address this, we devise various MapReduce programming paradigms that can be used to characterize system state transitions, compute phase spaces, functional equivalence classes, dynamic equivalence classes and cycle equivalence classes of dynamical systems. We also evaluate these …

    vt Repository record for On the Feasibility of MapReduce to Compute Phase Space Properties of Graphical Dynamical Systems: An Empirical Study (opens in a new tab)