Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 26 for “"Video Understanding"”.

  1. Graphical Models for Video Understanding

    … apply to the variety of tasks ranging from video clustering and stabilization, to video retrieval and building of the similarity measures between the distributions.

    uiuc Repository record for Graphical Models for Video Understanding (opens in a new tab)

  2. Language-Guided Video Understanding with Foundation Models

    Video understanding systems have achieved strong performance on controlled benchmarks, yet their deployment in real-world scenarios remains limited by assumptions about supervision, training data availability, and offline access to complete video sequences. These constraints are particularly …

    trento Repository record for Language-Guided Video Understanding with Foundation Models (opens in a new tab)

  3. Egocentric video understanding across modalities and domains

    L'abstract è presente nell'allegato / the abstract is in the attachment

    poli-torino Repository record for Egocentric video understanding across modalities and domains (opens in a new tab)

  4. Pixel-level video understanding with efficient deep models

    The ability to understand videos at the level of pixels plays a key role in a wide range of computer vision applications. For example, a robot or autonomous vehicle relies on classifying each pixel in the video stream into semantic categories to holistically understand the surrounding environment, …

    bu Repository record for Pixel-level video understanding with efficient deep models (opens in a new tab)

  5. Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications

    … represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex …

    auckland-ms Repository record for Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications (opens in a new tab)

  6. Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding

    … capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably …

    vt Repository record for Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding (opens in a new tab)

  7. Evolving Video Analysis: From Object Perception to Holistic Understanding

    The domain of video content analysis has experienced rapid advancements due to the proliferation of digital video content and the evolving capabilities of computer vision technologies. Despite these advancements, significant challenges remain in both video object perception and holistic video

    uts Repository record for Evolving Video Analysis: From Object Perception to Holistic Understanding (opens in a new tab)

  8. Object understanding in generalized video segmentation

    … object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity …

    uiuc Repository record for Object understanding in generalized video segmentation (opens in a new tab)

  9. Spatio-temporal human action detection and instance segmentation in videos

    With an exponential growth in the number of video capturing devices and digital video content, automatic video understanding is now at the forefront of computer vision research. This thesis presents a series of models for automatic human action detection in videos and also addresses the space-time …

    oxford-brookes Repository record for Spatio-temporal human action detection and instance segmentation in videos (opens in a new tab)

  10. Moving object detection and tracking for event-based video analysis

    … in the computer vision community towards video understanding, in particular towards visual event recognition...This dissertation surveys different taxonomies of motion understanding problems, identifies the major components in an automated visual event recognition system, and presents the …

    must-thes Repository record for Moving object detection and tracking for event-based video analysis (opens in a new tab)

  11. VirtualHome : learning to infer programs from synthetic videos of activities in the home

    … are implemented in the Unity3D game engine, and videos are recorded of an agent executing the collected programs in a simulated household environment. The VirtualHome simulator allows the creation of a large activity video dataset with rich groundtruth, enabling training and testing of video

    mit Repository record for VirtualHome : learning to infer programs from synthetic videos of activities in the home (opens in a new tab)

  12. Hierarchical Visual Content Modelling and Query based on Trees

    In recent years, such vast archives of video information have become available that human annotation of content is no longer feasible; automation of video content analysis is therefore highly desirable. The recognition of semantic content in images is a problem that relies on prior knowledge and …

    essex Repository record for Hierarchical Visual Content Modelling and Query based on Trees (opens in a new tab)

  13. Efficient Algorithms, Hardware Architectures and Circuits for Deep Learning Accelerators

    … weights for energy reduction. Last, we present VideoTime3, an algorithm and accelerator co-design for efficient real-time video understanding with temporal redundancy reduction and temporal modeling. Our proposed techniques enrich accelerator designers’ toolkits, pushing the boundaries of energy …

    mit Repository record for Efficient Algorithms, Hardware Architectures and Circuits for Deep Learning Accelerators (opens in a new tab)

  14. VirtualHome : simulating household activities via programs

    … from natural language descriptions or from videos. We then implement the most common atomic (inter)actions in the Unity3D game engine, and use our programs to "drive" an artificial agent to execute tasks in a simulated household environment. Our VirtualHome simulator allows us to create a …

    mit Repository record for VirtualHome : simulating household activities via programs (opens in a new tab)

  15. VICTORIOUS : video indexing with combined tracking and object recognition for improved object understanding in scenes

    Automatic understanding of video content is a problem which grows in importance every day. Video understanding algorithms require accuracy, robustness, speed, and scalability. Accuracy generates user confidence in usage. Robustness enables greater autonomy and reduced human intervention. …

    mit Repository record for VICTORIOUS : video indexing with combined tracking and object recognition for improved object understanding in scenes (opens in a new tab)

  16. GlitchAgent: Detecting Video Game Glitches from Gameplay Videos

    The increasing complexity of modern video games has made Quality Assurance (QA) a critical yet challenging bottleneck in the video game development and maintenance lifecycle, which relies heavily on expensive, labor-intensive, and inefficient manual testing. Automated glitch detection from gameplay …

    vt Repository record for GlitchAgent: Detecting Video Game Glitches from Gameplay Videos (opens in a new tab)

  17. An efficient neural representation for videos

    With the increasing popularity of videos, it has become crucial to find efficient and compact ways to represent them for easier storage, transmission, and downstream video tasks. Our dissertation proposes an innovative neural representation for videos called NeRV, which stores each video implicitly …

    maryland Repository record for An efficient neural representation for videos (opens in a new tab)

  18. Advancing 3D Segmentation: Deep Learning Techniques for Video and Medical Imaging

    … techniques have long been a cornerstone for understanding and interpreting complex visual data. While 2D segmentation has been extensively explored and has seen significant advancements, extending these successes to 3D segmentation presents a unique set of challenges. The complexity inherent …

    cambridge Repository record for Advancing 3D Segmentation: Deep Learning Techniques for Video and Medical Imaging (opens in a new tab)

  19. Human action recognition in the real world: handling domain shift in open-set, source-free and multi-source scenarios

    Human behavior understanding as an application of artificial intelligence and deep learning has quickly acquired popularity over the past few years, due to the crucial role it plays in trending fields such as human-robot interaction, autonomous driving, drone footage, sports and video surveillance. …

    trento Repository record for Human action recognition in the real world: handling domain shift in open-set, source-free and multi-source scenarios (opens in a new tab)

  20. Towards holistic scene understanding from monocular video

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01

    uiuc Repository record for Towards holistic scene understanding from monocular video (opens in a new tab)

Page 1 of 2