{"id":{"repo_id":"bielefeld","oai_identifier":"oai:pub.uni-bielefeld.de:2990802"},"canonical_url":"https://search.dev.ndltd.org/etd/bielefeld/oai:pub.uni-bielefeld.de:2990802","repository":{"repo_id":"bielefeld","name":"Universität Bielefeld","base_url":"https://pub.uni-bielefeld.de/oai"},"display":{"title":"Large-scale storage, analysis and integration of metabolomics data","abstract":"Continuous improvements in technologies and experimental methods in the field of metabolomics lead to rapidly increased amounts of data originating from different sources. These large amounts of data introduce a challenge on large scale metabolomic analytics as current computational methods have reached their limits on the volume of data being processed. Novel computational approaches with short processing times and robust processing methods are required in order to optimize the handling, integration and biological interpretation of these large amounts of data. Thus, in this thesis, I developed an automated and flexible cloud-based bioinformatics platform called MetHoS which utilizes a set of software tools to facilitate these needs. MetHoS is based on big data frameworks to enable storage, processing, analysis and integration of great amounts of mass spectrometry-based metabolomics data originating from different metabolomics studies with reduced processing and analysis times.<br/><br/>In order to tackle with these challenges the functionality of the platform is based on three concepts: (i) parallel processing, (ii) distributed storage and (iii) distributed analysis of metabolomics data.<br/><br/>For parallel processing I made use of the KNIME Analytics Platform software in order to develop a set of different workflows that process metabolomic experiments. Using these pre-defined KNIME workflows, MetHoS is able to quantify and identify metabolite features. Apache Spark is responsible for the distribution of these processing jobs to the nodes of the compute cluster where they are parallelized. Every node in the cluster downloads an experiment from the object storage, where it was previously uploaded, and processes it with a KNIME workflow. The results from the processing of every experiment are written into an Apache Cassandra database automatically.<br/><br/>Apache Cassandra is responsible for the distributed storage of the results of a pre-defined workflow across the cluster. The results originated from an experiment are saved in one Cassandra node and copied two more times to neighbouring nodes to ensure reliability and fault tolerance.<br/><br/>For the distributed analysis, Apache Spark Machine Learning Library is responsible for the implementations of several statistical tests in a distributed manner. In addition, for every test a set of choices is provided that defines the depth of analysis and handles the missing values.<br/><br/>In order to present the capabilities of MetHoS, thousands of experiments from different studies were downloaded from MetaboLights database and were used to perform a large-scale processing, storage and statistical analysis in a matter of hours. MetHoS, is capable of handling terabytes of metabolomics data and provides users an efficient and user-friendly handling of their own experimental metabolomic data.","abstract_html":"Continuous improvements in technologies and experimental methods in the field of metabolomics lead to rapidly increased amounts of data originating from different sources. These large amounts of data introduce a challenge on large scale metabolomic analytics as current computational methods have reached their limits on the volume of data being processed. Novel computational approaches with short processing times and robust processing methods are required in order to optimize the handling, integration and biological interpretation of these large amounts of data. Thus, in this thesis, I developed an automated and flexible cloud-based bioinformatics platform called MetHoS which utilizes a set of software tools to facilitate these needs. MetHoS is based on big data frameworks to enable storage, processing, analysis and integration of great amounts of mass spectrometry-based metabolomics data originating from different metabolomics studies with reduced processing and analysis times.&lt;br/&gt;&lt;br/&gt;In order to tackle with these challenges the functionality of the platform is based on three concepts: (i) parallel processing, (ii) distributed storage and (iii) distributed analysis of metabolomics data.&lt;br/&gt;&lt;br/&gt;For parallel processing I made use of the KNIME Analytics Platform software in order to develop a set of different workflows that process metabolomic experiments. Using these pre-defined KNIME workflows, MetHoS is able to quantify and identify metabolite features. Apache Spark is responsible for the distribution of these processing jobs to the nodes of the compute cluster where they are parallelized. Every node in the cluster downloads an experiment from the object storage, where it was previously uploaded, and processes it with a KNIME workflow. The results from the processing of every experiment are written into an Apache Cassandra database automatically.&lt;br/&gt;&lt;br/&gt;Apache Cassandra is responsible for the distributed storage of the results of a pre-defined workflow across the cluster. The results originated from an experiment are saved in one Cassandra node and copied two more times to neighbouring nodes to ensure reliability and fault tolerance.&lt;br/&gt;&lt;br/&gt;For the distributed analysis, Apache Spark Machine Learning Library is responsible for the implementations of several statistical tests in a distributed manner. In addition, for every test a set of choices is provided that defines the depth of analysis and handles the missing values.&lt;br/&gt;&lt;br/&gt;In order to present the capabilities of MetHoS, thousands of experiments from different studies were downloaded from MetaboLights database and were used to perform a large-scale processing, storage and statistical analysis in a matter of hours. MetHoS, is capable of handling terabytes of metabolomics data and provides users an efficient and user-friendly handling of their own experimental metabolomic data.","abstract_has_math":false,"creators":["Tzanakis, Konstantinos"],"institution":"Universität Bielefeld","degree_name":null,"degree_level":"thesis.doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-06-13","date_published":"2024-06-13","updated_at":"2026-07-27T18:50:07Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://pub.uni-bielefeld.de/record/2990802","outbound_label":"Repository record","outbound_source":"source_url"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Tzanakis, Konstantinos"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:publisher","label":"Institution","values":["Universitätsbibliothek Bielefeld"]},{"key":"dc:type","label":"Dc Type","values":["doctoralThesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["thesis.doctoral"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Universität Bielefeld"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Continuous improvements in technologies and experimental methods in the field of metabolomics lead to rapidly increased amounts of data originating from different sources. These large amounts of data introduce a challenge on large scale metabolomic analytics as current computational methods have reached their limits on the volume of data being processed. Novel computational approaches with short processing times and robust processing methods are required in order to optimize the handling, integration and biological interpretation of these large amounts of data. Thus, in this thesis, I developed an automated and flexible cloud-based bioinformatics platform called MetHoS which utilizes a set of software tools to facilitate these needs. MetHoS is based on big data frameworks to enable storage, processing, analysis and integration of great amounts of mass spectrometry-based metabolomics data originating from different metabolomics studies with reduced processing and analysis times.<br/><br/>In order to tackle with these challenges the functionality of the platform is based on three concepts: (i) parallel processing, (ii) distributed storage and (iii) distributed analysis of metabolomics data.<br/><br/>For parallel processing I made use of the KNIME Analytics Platform software in order to develop a set of different workflows that process metabolomic experiments. Using these pre-defined KNIME workflows, MetHoS is able to quantify and identify metabolite features. Apache Spark is responsible for the distribution of these processing jobs to the nodes of the compute cluster where they are parallelized. Every node in the cluster downloads an experiment from the object storage, where it was previously uploaded, and processes it with a KNIME workflow. The results from the processing of every experiment are written into an Apache Cassandra database automatically.<br/><br/>Apache Cassandra is responsible for the distributed storage of the results of a pre-defined workflow across the cluster. The results originated from an experiment are saved in one Cassandra node and copied two more times to neighbouring nodes to ensure reliability and fault tolerance.<br/><br/>For the distributed analysis, Apache Spark Machine Learning Library is responsible for the implementations of several statistical tests in a distributed manner. In addition, for every test a set of choices is provided that defines the depth of analysis and handles the missing values.<br/><br/>In order to present the capabilities of MetHoS, thousands of experiments from different studies were downloaded from MetaboLights database and were used to perform a large-scale processing, storage and statistical analysis in a matter of hours. MetHoS, is capable of handling terabytes of metabolomics data and provides users an efficient and user-friendly handling of their own experimental metabolomic data."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large-scale storage, analysis and integration of metabolomics data"]}]}],"canonical_facts":{"dc:creator":["Tzanakis, Konstantinos"],"dc:description.abstract":["Continuous improvements in technologies and experimental methods in the field of metabolomics lead to rapidly increased amounts of data originating from different sources. These large amounts of data introduce a challenge on large scale metabolomic analytics as current computational methods have reached their limits on the volume of data being processed. Novel computational approaches with short processing times and robust processing methods are required in order to optimize the handling, integration and biological interpretation of these large amounts of data. Thus, in this thesis, I developed an automated and flexible cloud-based bioinformatics platform called MetHoS which utilizes a set of software tools to facilitate these needs. MetHoS is based on big data frameworks to enable storage, processing, analysis and integration of great amounts of mass spectrometry-based metabolomics data originating from different metabolomics studies with reduced processing and analysis times.<br/><br/>In order to tackle with these challenges the functionality of the platform is based on three concepts: (i) parallel processing, (ii) distributed storage and (iii) distributed analysis of metabolomics data.<br/><br/>For parallel processing I made use of the KNIME Analytics Platform software in order to develop a set of different workflows that process metabolomic experiments. Using these pre-defined KNIME workflows, MetHoS is able to quantify and identify metabolite features. Apache Spark is responsible for the distribution of these processing jobs to the nodes of the compute cluster where they are parallelized. Every node in the cluster downloads an experiment from the object storage, where it was previously uploaded, and processes it with a KNIME workflow. The results from the processing of every experiment are written into an Apache Cassandra database automatically.<br/><br/>Apache Cassandra is responsible for the distributed storage of the results of a pre-defined workflow across the cluster. The results originated from an experiment are saved in one Cassandra node and copied two more times to neighbouring nodes to ensure reliability and fault tolerance.<br/><br/>For the distributed analysis, Apache Spark Machine Learning Library is responsible for the implementations of several statistical tests in a distributed manner. In addition, for every test a set of choices is provided that defines the depth of analysis and handles the missing values.<br/><br/>In order to present the capabilities of MetHoS, thousands of experiments from different studies were downloaded from MetaboLights database and were used to perform a large-scale processing, storage and statistical analysis in a matter of hours. MetHoS, is capable of handling terabytes of metabolomics data and provides users an efficient and user-friendly handling of their own experimental metabolomic data."],"dc:format.medium":["application/pdf"],"dc:publisher":["Universitätsbibliothek Bielefeld"],"dc:title":["Large-scale storage, analysis and integration of metabolomics data"],"dc:type":["doctoralThesis"],"thesis:degree_level":["thesis.doctoral"],"thesis:institution_name":["Universität Bielefeld"]},"updated_at":"2026-07-27T18:50:07Z"}