{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113134"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113134","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Next-generation safety-critical systems using COTS based homogeneous multi-core processors and heterogeneous MPSoCS","abstract":"The embedded computing revolution is pushing the transition from a single-core processor to a multicore processor where multiple single-core processors are packed into a single chip. With recent hardware advancements and the evolution of new challenging applications in autonomy such as self-driving cars and unmanned aerial vehicles (UAVs), hardware needs to perform immense computation in real-time and at low power. To address such requirements, hardware manufacturers are integrating multiple processing elements such as CPU clusters, GPU, programmable logic (PL), and AI accelerators into a single multiprocessor system-on-chip (MPSoC). One of the common attributes of these high-performance multicore processors and MPSoCs is having a shared memory subsystem. Having shared memory for multicore processors suits well for systems where the average-case performance of the system is the performance metric. However, this is not suitable for safety-critical systems where the developer is interested in the worst-case execution time (WCET) of the task on each processing element in the system. The WCET of a task running on one of the processing elements in these high-performance multicore/MPSoCs changes as we activate more processing elements because of the contention on the shared memory subsystem. This thesis proposes hardware/software solutions for commercial-off-the-shelf (COTS) homogeneous multicore architectures and heterogeneous MPSoCs such that the WCET of the tasks running in these environments can be predictable. Moreover, it presents techniques for inter-core communication, reliability, and streaming of compiler-generated segments on CPU only and CPU + accelerators while ensuring predictability. We categorize COTS-based homogeneous multicore architectures into cache-based and scratchpad-based multicore platforms. For cache-based multicore architectures, the primary sources of contention among the cores include shared last-level cache (LLC), the DRAM memory controller, and the DRAM banks. Strict partitioning of shared resources on such platforms at the operating system (OS) has been proposed in the literature for WCET estimation. However, such partitioning prohibits inter-core communication. This thesis proposes an analyzable inter-core communication mechanism for such platforms by relaxing the strict partitioning assumption. For scratchpad-based multicore architectures, this thesis proposes an SPM-centric OS that takes advantage of the scratchpad memory (SPM), DMA, and the I/O subsystems to implement three-phase execution (load, execute and unload) model such that inter-core interference among the cores can be mitigated. The OS has also been extended to handle the inter/intra-core communication and the bit-flip errors that the ECC modules cannot correct. For the newer generation of heterogeneous MPSoCs that employ processing systems (PS) and programmable logic (PL), the thesis proposes a hardware/software co-design approach that isolates the different criticality domains. Different criticality domains take advantage of hypervisor-level cache-coloring to partition the shared cache. The hypervisor also ensures that high and medium criticality cores always execute the tasks from the SPM in the PL using their dedicated interfaces. The dual-ported SPM, together with the DMA and I/O core, implements the three-phase execution model for high and medium criticality domains. The low-criticality domain executes from DRAM shared by DMA. The work has also been extended to provide the APIs that allow the streaming of segments of large tasks that do not fit into half the SPM for the CPUs. Moreover, an OS framework to support segment streaming across different processing units such as CPU and accelerators with DMA in the PL has also been proposed.","abstract_html":"The embedded computing revolution is pushing the transition from a single-core processor to a multicore processor where multiple single-core processors are packed into a single chip. With recent hardware advancements and the evolution of new challenging applications in autonomy such as self-driving cars and unmanned aerial vehicles (UAVs), hardware needs to perform immense computation in real-time and at low power. To address such requirements, hardware manufacturers are integrating multiple processing elements such as CPU clusters, GPU, programmable logic (PL), and AI accelerators into a single multiprocessor system-on-chip (MPSoC). One of the common attributes of these high-performance multicore processors and MPSoCs is having a shared memory subsystem. Having shared memory for multicore processors suits well for systems where the average-case performance of the system is the performance metric. However, this is not suitable for safety-critical systems where the developer is interested in the worst-case execution time (WCET) of the task on each processing element in the system. The WCET of a task running on one of the processing elements in these high-performance multicore/MPSoCs changes as we activate more processing elements because of the contention on the shared memory subsystem. This thesis proposes hardware/software solutions for commercial-off-the-shelf (COTS) homogeneous multicore architectures and heterogeneous MPSoCs such that the WCET of the tasks running in these environments can be predictable. Moreover, it presents techniques for inter-core communication, reliability, and streaming of compiler-generated segments on CPU only and CPU + accelerators while ensuring predictability. We categorize COTS-based homogeneous multicore architectures into cache-based and scratchpad-based multicore platforms. For cache-based multicore architectures, the primary sources of contention among the cores include shared last-level cache (LLC), the DRAM memory controller, and the DRAM banks. Strict partitioning of shared resources on such platforms at the operating system (OS) has been proposed in the literature for WCET estimation. However, such partitioning prohibits inter-core communication. This thesis proposes an analyzable inter-core communication mechanism for such platforms by relaxing the strict partitioning assumption. For scratchpad-based multicore architectures, this thesis proposes an SPM-centric OS that takes advantage of the scratchpad memory (SPM), DMA, and the I/O subsystems to implement three-phase execution (load, execute and unload) model such that inter-core interference among the cores can be mitigated. The OS has also been extended to handle the inter/intra-core communication and the bit-flip errors that the ECC modules cannot correct. For the newer generation of heterogeneous MPSoCs that employ processing systems (PS) and programmable logic (PL), the thesis proposes a hardware/software co-design approach that isolates the different criticality domains. Different criticality domains take advantage of hypervisor-level cache-coloring to partition the shared cache. The hypervisor also ensures that high and medium criticality cores always execute the tasks from the SPM in the PL using their dedicated interfaces. The dual-ported SPM, together with the DMA and I/O core, implements the three-phase execution model for high and medium criticality domains. The low-criticality domain executes from DRAM shared by DMA. The work has also been extended to provide the APIs that allow the streaming of segments of large tasks that do not fit into half the SPM for the CPUs. Moreover, an OS framework to support segment streaming across different processing units such as CPU and accelerators with DMA in the PL has also been proposed.","abstract_has_math":false,"creators":["Tabish, Rohan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Caccamo, Marco","Sha, Lui Raymond","Nahrstedt, Klara","Pellizzoni, Rodolfo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T22:34:50Z","date_published":"2022-01-12T22:34:50Z","updated_at":"2026-07-22T22:24:53Z","subjects":["Multicore, MPSoC, Fault Tolerance, Real-Time Computing, Communication, Predictability, Embedded Computing, Embedded Systems, Cyber-Physical Systems"],"languages":["en"],"rights":["Copyright 2021 Rohan Tabish"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113134","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Caccamo, Marco","Sha, Lui Raymond","Nahrstedt, Klara","Pellizzoni, Rodolfo"]},{"key":"dc:creator","label":"Author","values":["Tabish, Rohan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T22:34:50Z","2024-01-12T22:35:30Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multicore, MPSoC, Fault Tolerance, Real-Time Computing, Communication, Predictability, Embedded Computing, Embedded Systems, Cyber-Physical Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Rohan Tabish"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113134"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The embedded computing revolution is pushing the transition from a single-core processor to a multicore processor where multiple single-core processors are packed into a single chip. With recent hardware advancements and the evolution of new challenging applications in autonomy such as self-driving cars and unmanned aerial vehicles (UAVs), hardware needs to perform immense computation in real-time and at low power. To address such requirements, hardware manufacturers are integrating multiple processing elements such as CPU clusters, GPU, programmable logic (PL), and AI accelerators into a single multiprocessor system-on-chip (MPSoC). One of the common attributes of these high-performance multicore processors and MPSoCs is having a shared memory subsystem. Having shared memory for multicore processors suits well for systems where the average-case performance of the system is the performance metric. However, this is not suitable for safety-critical systems where the developer is interested in the worst-case execution time (WCET) of the task on each processing element in the system. The WCET of a task running on one of the processing elements in these high-performance multicore/MPSoCs changes as we activate more processing elements because of the contention on the shared memory subsystem. This thesis proposes hardware/software solutions for commercial-off-the-shelf (COTS) homogeneous multicore architectures and heterogeneous MPSoCs such that the WCET of the tasks running in these environments can be predictable. Moreover, it presents techniques for inter-core communication, reliability, and streaming of compiler-generated segments on CPU only and CPU + accelerators while ensuring predictability. We categorize COTS-based homogeneous multicore architectures into cache-based and scratchpad-based multicore platforms. For cache-based multicore architectures, the primary sources of contention among the cores include shared last-level cache (LLC), the DRAM memory controller, and the DRAM banks. Strict partitioning of shared resources on such platforms at the operating system (OS) has been proposed in the literature for WCET estimation. However, such partitioning prohibits inter-core communication. This thesis proposes an analyzable inter-core communication mechanism for such platforms by relaxing the strict partitioning assumption. For scratchpad-based multicore architectures, this thesis proposes an SPM-centric OS that takes advantage of the scratchpad memory (SPM), DMA, and the I/O subsystems to implement three-phase execution (load, execute and unload) model such that inter-core interference among the cores can be mitigated. The OS has also been extended to handle the inter/intra-core communication and the bit-flip errors that the ECC modules cannot correct. For the newer generation of heterogeneous MPSoCs that employ processing systems (PS) and programmable logic (PL), the thesis proposes a hardware/software co-design approach that isolates the different criticality domains. Different criticality domains take advantage of hypervisor-level cache-coloring to partition the shared cache. The hypervisor also ensures that high and medium criticality cores always execute the tasks from the SPM in the PL using their dedicated interfaces. The dual-ported SPM, together with the DMA and I/O core, implements the three-phase execution model for high and medium criticality domains. The low-criticality domain executes from DRAM shared by DMA. The work has also been extended to provide the APIs that allow the streaming of segments of large tasks that do not fit into half the SPM for the CPUs. Moreover, an OS framework to support segment streaming across different processing units such as CPU and accelerators with DMA in the PL has also been proposed.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Rohan Tabish, accepted the attached license on 2021-07-12 at 12:27.","The student, Rohan Tabish, submitted this Dissertation for approval on 2021-07-12 at 12:48.","This Dissertation was approved for publication on 2021-07-12 at 16:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16740 on 2022-01-12 at 12:52:45","Made available in DSpace on 2022-01-12T22:34:50Z (GMT). No. of bitstreams: 3 TABISH-DISSERTATION-2021.pdf: 4849631 bytes, checksum: 43c530db266665e19b3e57d62f0e172e (MD5) Final_Thesis.zip: 8766720 bytes, checksum: 60841962903d251f4e0863af1cd2d569 (MD5) LICENSE.txt: 4209 bytes, checksum: 419c64a4abc07c0547b425f6bed184f4 (MD5) Previous issue date: 2021-07-12","Embargo set by: Seth Robbins for item 121060 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Next-generation safety-critical systems using COTS based homogeneous multi-core processors and heterogeneous MPSoCS"]}]}],"canonical_facts":{"dc:contributor":["Caccamo, Marco","Sha, Lui Raymond","Nahrstedt, Klara","Pellizzoni, Rodolfo"],"dc:creator":["Tabish, Rohan"],"dc:date":["2022-01-12T22:34:50Z","2024-01-12T22:35:30Z","2021-07-12","2021-08"],"dc:description":["The embedded computing revolution is pushing the transition from a single-core processor to a multicore processor where multiple single-core processors are packed into a single chip. With recent hardware advancements and the evolution of new challenging applications in autonomy such as self-driving cars and unmanned aerial vehicles (UAVs), hardware needs to perform immense computation in real-time and at low power. To address such requirements, hardware manufacturers are integrating multiple processing elements such as CPU clusters, GPU, programmable logic (PL), and AI accelerators into a single multiprocessor system-on-chip (MPSoC). One of the common attributes of these high-performance multicore processors and MPSoCs is having a shared memory subsystem. Having shared memory for multicore processors suits well for systems where the average-case performance of the system is the performance metric. However, this is not suitable for safety-critical systems where the developer is interested in the worst-case execution time (WCET) of the task on each processing element in the system. The WCET of a task running on one of the processing elements in these high-performance multicore/MPSoCs changes as we activate more processing elements because of the contention on the shared memory subsystem. This thesis proposes hardware/software solutions for commercial-off-the-shelf (COTS) homogeneous multicore architectures and heterogeneous MPSoCs such that the WCET of the tasks running in these environments can be predictable. Moreover, it presents techniques for inter-core communication, reliability, and streaming of compiler-generated segments on CPU only and CPU + accelerators while ensuring predictability. We categorize COTS-based homogeneous multicore architectures into cache-based and scratchpad-based multicore platforms. For cache-based multicore architectures, the primary sources of contention among the cores include shared last-level cache (LLC), the DRAM memory controller, and the DRAM banks. Strict partitioning of shared resources on such platforms at the operating system (OS) has been proposed in the literature for WCET estimation. However, such partitioning prohibits inter-core communication. This thesis proposes an analyzable inter-core communication mechanism for such platforms by relaxing the strict partitioning assumption. For scratchpad-based multicore architectures, this thesis proposes an SPM-centric OS that takes advantage of the scratchpad memory (SPM), DMA, and the I/O subsystems to implement three-phase execution (load, execute and unload) model such that inter-core interference among the cores can be mitigated. The OS has also been extended to handle the inter/intra-core communication and the bit-flip errors that the ECC modules cannot correct. For the newer generation of heterogeneous MPSoCs that employ processing systems (PS) and programmable logic (PL), the thesis proposes a hardware/software co-design approach that isolates the different criticality domains. Different criticality domains take advantage of hypervisor-level cache-coloring to partition the shared cache. The hypervisor also ensures that high and medium criticality cores always execute the tasks from the SPM in the PL using their dedicated interfaces. The dual-ported SPM, together with the DMA and I/O core, implements the three-phase execution model for high and medium criticality domains. The low-criticality domain executes from DRAM shared by DMA. The work has also been extended to provide the APIs that allow the streaming of segments of large tasks that do not fit into half the SPM for the CPUs. Moreover, an OS framework to support segment streaming across different processing units such as CPU and accelerators with DMA in the PL has also been proposed.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Rohan Tabish, accepted the attached license on 2021-07-12 at 12:27.","The student, Rohan Tabish, submitted this Dissertation for approval on 2021-07-12 at 12:48.","This Dissertation was approved for publication on 2021-07-12 at 16:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16740 on 2022-01-12 at 12:52:45","Made available in DSpace on 2022-01-12T22:34:50Z (GMT). No. of bitstreams: 3 TABISH-DISSERTATION-2021.pdf: 4849631 bytes, checksum: 43c530db266665e19b3e57d62f0e172e (MD5) Final_Thesis.zip: 8766720 bytes, checksum: 60841962903d251f4e0863af1cd2d569 (MD5) LICENSE.txt: 4209 bytes, checksum: 419c64a4abc07c0547b425f6bed184f4 (MD5) Previous issue date: 2021-07-12","Embargo set by: Seth Robbins for item 121060 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113134"],"dc:language":["en"],"dc:rights":["Copyright 2021 Rohan Tabish"],"dc:subject":["Multicore, MPSoC, Fault Tolerance, Real-Time Computing, Communication, Predictability, Embedded Computing, Embedded Systems, Cyber-Physical Systems"],"dc:title":["Next-generation safety-critical systems using COTS based homogeneous multi-core processors and heterogeneous MPSoCS"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}