{"id":{"repo_id":"aachen","oai_identifier":"oai:publications.rwth-aachen.de:52362"},"canonical_url":"https://search.dev.ndltd.org/etd/aachen/oai:publications.rwth-aachen.de:52362","repository":{"repo_id":"aachen","name":"RWTH Aachen University","base_url":"https://publications.rwth-aachen.de/oai2d"},"display":{"title":"Nicht-intrusive Erkennung isolierter Gesten und Gebärden","abstract":"Research in automatic gesture recognition is supposed to yield gesture-based human-machine interfaces and sign language translation systems. Using data gloves for capturing the hands or video cameras with markers attached to the user allows reliable separation of vocabulary sizes between 100 and 200 gestures. From a user friendliness perspective these kinds of methods are regarded as unsuitable since intrusive. Vision-based approaches without markers are non-intrusive, however up to now they are limited to vocabulary sizes of 40 gestures at a maximum recognition rate of 98.1 %. The topic of this thesis is the non-intrusive vision-based recognition of isolated gestures from a monocular camera view. The intention is to distinguish a large vocabulary, which consists of 152 gestures from German sign language, for a given person. Thereby, the main interest lies on the development and analysis of suitable image processing routines for determining the information that is relevant for gesture recognition. The localisation of the hands represents the crucial difficulty of this objective. The hands can change their shapes, positions and orientations in the recorded image sequence in a non-continuous way; they offer no dependable cues for unique identification and can be confused easily due to similarity. Additionally, the hands can overlap each other and the face, which increases the difficulty of localising them within a single image. Furthermore, variations are an intrinsic problem of gestures that must be modelled and considered for a reliable recognition. The solution to the problem is given by the developed tracking system, which takes a complete image sequence into account. Basic elements of the analysis are the image regions that comply with a predefined skin colour model. First, the face of the user is tracked by a combination of a Mean-Shift tracker with an Active Shape Model. Then multiple hypotheses for the spatial and temporal assignment of the hands to the image data are generated with the face as a consistent reference object. The inference with regard to the possible different combinations is based on knowledge about gestures, a body model, and motion predictions generated by a Kalman filter. In case of overlap a method derived from the Expectation-Maximisation algorithm is applied for a position estimate. Finally, some basic features for describing the hands are computed and passed to the classification. The classifier makes use of the statistical properties of Hidden Markov Models, in order to handle possible variations in duration and in the observed features. In this thesis for the first time a vocabulary of the given size is recognised at a rate of 97.6 % with a vision based gesture recognition approach that does not require any simplifying means for image processing like markers. A novelty is the evaluation of the tracking system, for which a substantial database with reference segmentations was created. The processing parameters were adapted and optimised according to quantitative results of experiments with the defined database. Additionally, the suitability of the tracking system for different persons and environments was assessed qualitatively.","abstract_html":"Research in automatic gesture recognition is supposed to yield gesture-based human-machine interfaces and sign language translation systems. Using data gloves for capturing the hands or video cameras with markers attached to the user allows reliable separation of vocabulary sizes between 100 and 200 gestures. From a user friendliness perspective these kinds of methods are regarded as unsuitable since intrusive. Vision-based approaches without markers are non-intrusive, however up to now they are limited to vocabulary sizes of 40 gestures at a maximum recognition rate of 98.1 %. The topic of this thesis is the non-intrusive vision-based recognition of isolated gestures from a monocular camera view. The intention is to distinguish a large vocabulary, which consists of 152 gestures from German sign language, for a given person. Thereby, the main interest lies on the development and analysis of suitable image processing routines for determining the information that is relevant for gesture recognition. The localisation of the hands represents the crucial difficulty of this objective. The hands can change their shapes, positions and orientations in the recorded image sequence in a non-continuous way; they offer no dependable cues for unique identification and can be confused easily due to similarity. Additionally, the hands can overlap each other and the face, which increases the difficulty of localising them within a single image. Furthermore, variations are an intrinsic problem of gestures that must be modelled and considered for a reliable recognition. The solution to the problem is given by the developed tracking system, which takes a complete image sequence into account. Basic elements of the analysis are the image regions that comply with a predefined skin colour model. First, the face of the user is tracked by a combination of a Mean-Shift tracker with an Active Shape Model. Then multiple hypotheses for the spatial and temporal assignment of the hands to the image data are generated with the face as a consistent reference object. The inference with regard to the possible different combinations is based on knowledge about gestures, a body model, and motion predictions generated by a Kalman filter. In case of overlap a method derived from the Expectation-Maximisation algorithm is applied for a position estimate. Finally, some basic features for describing the hands are computed and passed to the classification. The classifier makes use of the statistical properties of Hidden Markov Models, in order to handle possible variations in duration and in the observed features. In this thesis for the first time a vocabulary of the given size is recognised at a rate of 97.6 % with a vision based gesture recognition approach that does not require any simplifying means for image processing like markers. A novelty is the evaluation of the tracking system, for which a substantial database with reference segmentations was created. The processing parameters were adapted and optimised according to quantitative results of experiments with the defined database. Additionally, the suitability of the tracking system for different persons and environments was assessed qualitatively.","abstract_has_math":false,"creators":["Akyol, Suat"],"institution":"Publikationsserver der RWTH Aachen University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Kraiss, Karl-Friedrich"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2003,"date_issued":"2003","date_published":"2003","updated_at":"2026-07-30T19:40:50Z","subjects":["info:eu-repo/classification/ddc/620","Videoverarbeitung","Geste","Gebärde","Bilderkennung","Objektverfolgung","Merkmalsextraktion","Ingenieurwissenschaften","Gestenerkennung","Bildverarbeitung","Tracking","Mustererkennung","Mensch-Maschine Interaktion"],"languages":["ger"],"rights":["info:eu-repo/semantics/openAccess"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-114591%22"],"render_values":[{"text":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-114591%22","href":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-114591%22","code":true}]}]},"links":{"outbound_url":"https://publications.rwth-aachen.de/record/52362","outbound_label":"Repository record","outbound_source":"dc:identifier"},"source_record":{"url":"https://publications.rwth-aachen.de/oai2d?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Apublications.rwth-aachen.de%3A52362","prefix":"oai_dc"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kraiss, Karl-Friedrich"]},{"key":"dc:creator","label":"Author","values":["Akyol, Suat"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:coverage","label":"Dc Coverage","values":["DE"]},{"key":"dc:date","label":"Dc Date","values":["2003"]},{"key":"dc:publisher","label":"Institution","values":["Publikationsserver der RWTH Aachen University"]},{"key":"dc:relation","label":"Dc Relation","values":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-6239"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["info:eu-repo/classification/ddc/620","Videoverarbeitung","Geste","Gebärde","Bilderkennung","Objektverfolgung","Merkmalsextraktion","Ingenieurwissenschaften","Gestenerkennung","Bildverarbeitung","Tracking","Mustererkennung","Mensch-Maschine Interaktion"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["ger"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/record/52362","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-114591%22"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Research in automatic gesture recognition is supposed to yield gesture-based human-machine interfaces and sign language translation systems. Using data gloves for capturing the hands or video cameras with markers attached to the user allows reliable separation of vocabulary sizes between 100 and 200 gestures. From a user friendliness perspective these kinds of methods are regarded as unsuitable since intrusive. Vision-based approaches without markers are non-intrusive, however up to now they are limited to vocabulary sizes of 40 gestures at a maximum recognition rate of 98.1 %. The topic of this thesis is the non-intrusive vision-based recognition of isolated gestures from a monocular camera view. The intention is to distinguish a large vocabulary, which consists of 152 gestures from German sign language, for a given person. Thereby, the main interest lies on the development and analysis of suitable image processing routines for determining the information that is relevant for gesture recognition. The localisation of the hands represents the crucial difficulty of this objective. The hands can change their shapes, positions and orientations in the recorded image sequence in a non-continuous way; they offer no dependable cues for unique identification and can be confused easily due to similarity. Additionally, the hands can overlap each other and the face, which increases the difficulty of localising them within a single image. Furthermore, variations are an intrinsic problem of gestures that must be modelled and considered for a reliable recognition. The solution to the problem is given by the developed tracking system, which takes a complete image sequence into account. Basic elements of the analysis are the image regions that comply with a predefined skin colour model. First, the face of the user is tracked by a combination of a Mean-Shift tracker with an Active Shape Model. Then multiple hypotheses for the spatial and temporal assignment of the hands to the image data are generated with the face as a consistent reference object. The inference with regard to the possible different combinations is based on knowledge about gestures, a body model, and motion predictions generated by a Kalman filter. In case of overlap a method derived from the Expectation-Maximisation algorithm is applied for a position estimate. Finally, some basic features for describing the hands are computed and passed to the classification. The classifier makes use of the statistical properties of Hidden Markov Models, in order to handle possible variations in duration and in the observed features. In this thesis for the first time a vocabulary of the given size is recognised at a rate of 97.6 % with a vision based gesture recognition approach that does not require any simplifying means for image processing like markers. A novelty is the evaluation of the tracking system, for which a substantial database with reference segmentations was created. The processing parameters were adapted and optimised according to quantitative results of experiments with the defined database. Additionally, the suitability of the tracking system for different persons and environments was assessed qualitatively."]},{"key":"dc:source","label":"Dc Source","values":["Aachen : Publikationsserver der RWTH Aachen University XXII, 164 S. : Ill., graph. Darst. (2003). = Aachen, Techn. Hochsch., Diss., 2003"]},{"key":"dc:title","label":"Title","values":["Nicht-intrusive Erkennung isolierter Gesten und Gebärden"]}]}],"canonical_facts":{"dc:contributor":["Kraiss, Karl-Friedrich"],"dc:coverage":["DE"],"dc:creator":["Akyol, Suat"],"dc:date":["2003"],"dc:description":["Research in automatic gesture recognition is supposed to yield gesture-based human-machine interfaces and sign language translation systems. Using data gloves for capturing the hands or video cameras with markers attached to the user allows reliable separation of vocabulary sizes between 100 and 200 gestures. From a user friendliness perspective these kinds of methods are regarded as unsuitable since intrusive. Vision-based approaches without markers are non-intrusive, however up to now they are limited to vocabulary sizes of 40 gestures at a maximum recognition rate of 98.1 %. The topic of this thesis is the non-intrusive vision-based recognition of isolated gestures from a monocular camera view. The intention is to distinguish a large vocabulary, which consists of 152 gestures from German sign language, for a given person. Thereby, the main interest lies on the development and analysis of suitable image processing routines for determining the information that is relevant for gesture recognition. The localisation of the hands represents the crucial difficulty of this objective. The hands can change their shapes, positions and orientations in the recorded image sequence in a non-continuous way; they offer no dependable cues for unique identification and can be confused easily due to similarity. Additionally, the hands can overlap each other and the face, which increases the difficulty of localising them within a single image. Furthermore, variations are an intrinsic problem of gestures that must be modelled and considered for a reliable recognition. The solution to the problem is given by the developed tracking system, which takes a complete image sequence into account. Basic elements of the analysis are the image regions that comply with a predefined skin colour model. First, the face of the user is tracked by a combination of a Mean-Shift tracker with an Active Shape Model. Then multiple hypotheses for the spatial and temporal assignment of the hands to the image data are generated with the face as a consistent reference object. The inference with regard to the possible different combinations is based on knowledge about gestures, a body model, and motion predictions generated by a Kalman filter. In case of overlap a method derived from the Expectation-Maximisation algorithm is applied for a position estimate. Finally, some basic features for describing the hands are computed and passed to the classification. The classifier makes use of the statistical properties of Hidden Markov Models, in order to handle possible variations in duration and in the observed features. In this thesis for the first time a vocabulary of the given size is recognised at a rate of 97.6 % with a vision based gesture recognition approach that does not require any simplifying means for image processing like markers. A novelty is the evaluation of the tracking system, for which a substantial database with reference segmentations was created. The processing parameters were adapted and optimised according to quantitative results of experiments with the defined database. Additionally, the suitability of the tracking system for different persons and environments was assessed qualitatively."],"dc:identifier":["https://publications.rwth-aachen.de/record/52362","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-114591%22"],"dc:language":["ger"],"dc:publisher":["Publikationsserver der RWTH Aachen University"],"dc:relation":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-6239"],"dc:rights":["info:eu-repo/semantics/openAccess"],"dc:source":["Aachen : Publikationsserver der RWTH Aachen University XXII, 164 S. : Ill., graph. Darst. (2003). = Aachen, Techn. Hochsch., Diss., 2003"],"dc:subject":["info:eu-repo/classification/ddc/620","Videoverarbeitung","Geste","Gebärde","Bilderkennung","Objektverfolgung","Merkmalsextraktion","Ingenieurwissenschaften","Gestenerkennung","Bildverarbeitung","Tracking","Mustererkennung","Mensch-Maschine Interaktion"],"dc:title":["Nicht-intrusive Erkennung isolierter Gesten und Gebärden"],"dc:type":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]},"updated_at":"2026-07-30T19:40:50Z"}