Abstract
dc:description.abstractWith the growing prevalence of portable cameras—such as smartphones, action cameras, and smart glasses—recording first-person videos of daily activities has become increasingly common. However, these recordings often suffer from shaky footage caused by the wearer's continuous movements, making them physically uncomfortable to watch, and include repetitive or irrelevant segments that make them tedious to watch. To address these challenges, hyperlapse methods fast-forward egocentric videos while stabilizing camera motion, and semantic hyperlapse methods additionally preserve the most important segments. Although audio is an important part of watching videos, it is often overlooked in hyperlapse creation, leaving the choice of soundtrack to the user. In this dissertation, we introduce a multimodal hyperlapse algorithm that jointly optimizes semantic content retention, visual stability, and playback alignment with a user-chosen song's loudness. Specifically, the hyperlapse slows down during quiet parts of the song to highlight important frames and speeds up during louder segments to de-emphasize less critical content. We also propose strategies to select songs that best complement the hyperlapse. Our experiments show that this approach outperforms existing methods in semantic retention and loudness–speed correlation, while maintaining comparable camera stability and temporal continuity. Keywords: video summarization; semantic fast-forward; egocentric videos; hyperlapse; loudness
Degree
thesis:*- Grantor dc:publisher
- Universidade Federal de Viçosa
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nepomuceno, Raphael Carmo Silva
- Advisor dc:contributor.advisor
-
- Silva, Michel Melo da
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- Acesso Aberto
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:locus.ufv.br:123456789/34783