Abstract
dc:description.abstractUnderstanding people's actions and interactions typically depends on seeing them. Automating the process of human action recognition and event captioning from visual data has been the topic of much research in the computer vision community. But what if it is too dark, or if the person is occluded or behind a wall? This thesis develops a model that can detect human actions through walls and occlusions, and in poor lighting conditions. The model takes radio frequency (RF) signals as input, generates 3D human skeletons as an intermediate representation, and recognizes actions and interactions of multiple people over time, or even generate language descriptions for the event. By translating the input to an intermediate skeleton-based representation, our model can learn from both vision-based and RF-based datasets. This thesis also introduces a new model for captioning daily life by analyzing RF signal in the home with the home's floormap. It can further observe and caption people's life through walls and occlusions and in dark settings. We show that our model achieves comparable accuracy to vision-based action recognition systems in visible scenarios, yet continues to work accurately when people are not visible, hence addressing scenarios that are beyond the limit of today's vision-based action recognition.
Degree
thesis:*- Name thesis:degree_name
- Master
- Department dc:contributor.department
- Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Fan, Lijie(Biologist)Massachusetts Institute of Technology.
- Advisor dc:contributor.advisor
-
- Dina Katabi.
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided.
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1721.1/127341
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/127341