{"id":{"repo_id":"uoit","oai_identifier":"oai:ontariotechu.scholaris.ca:10155/1132"},"canonical_url":"https://search.dev.ndltd.org/etd/uoit/oai:ontariotechu.scholaris.ca:10155/1132","repository":{"repo_id":"uoit","name":"Ontario Institute of Technology","base_url":"https://ontariotechu.scholaris.ca/server/oai/request"},"display":{"title":"Multi-character prediction using attention","abstract":"We propose a computational attention approach to localize and classify characters in a sequence in a given image. Our approach combines spatial soft-attention with attention regularization and learns “where-to-look” to carry out the sequence classification task. The image is first passed through a Convolutional Neural Network (CNN) that serves as feature extractor. Then at each Recurrent Neural Network (RNN) time step, the attention mechanism attends to the relevant features sequentially to make predictions. The attention mechanism also includes a start and stop state, which instructs the mechanism to start looking and guides it when to stop (e.g., when the sequence has been exhausted). We demonstrate our approach on two sequence detection tasks—multi-digit classification and CAPTCHA unlocking—using the publicly available Street View House Numbers (SVHN) dataset and a custom CAPTCHA dataset. The experiments confirm our hypothesis that the network learns to attend to relevant features by minimizing the loss between the ground truth attention masks and the predicted attention masks.","abstract_html":"We propose a computational attention approach to localize and classify characters in a sequence in a given image. Our approach combines spatial soft-attention with attention regularization and learns “where-to-look” to carry out the sequence classification task. The image is first passed through a Convolutional Neural Network (CNN) that serves as feature extractor. Then at each Recurrent Neural Network (RNN) time step, the attention mechanism attends to the relevant features sequentially to make predictions. The attention mechanism also includes a start and stop state, which instructs the mechanism to start looking and guides it when to stop (e.g., when the sequence has been exhausted). We demonstrate our approach on two sequence detection tasks—multi-digit classification and CAPTCHA unlocking—using the publicly available Street View House Numbers (SVHN) dataset and a custom CAPTCHA dataset. The experiments confirm our hypothesis that the network learns to attend to relevant features by minimizing the loss between the ground truth attention masks and the predicted attention masks.","abstract_has_math":false,"creators":["Baeenh, Mohmmed"],"institution":"University of Ontario Institute of Technology","degree_name":"Master of Science (MSc)","degree_level":null,"degree_discipline":"Applied Bioscience","degree_department":null,"school":null,"contributors":[],"advisors":["Qureshi, Faisal"],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-01-01","date_published":"2020-01-01","updated_at":"2026-07-24T05:35:16Z","subjects":["Computational Attention","Convolutional Neural Networks","Recurrent Neural Networks","Multi-Digit Classification","CAPTCHA"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10155/1132","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Qureshi, Faisal"]},{"key":"dc:creator","label":"Author","values":["Baeenh, Mohmmed"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-02-27T14:30:26Z","2022-03-29T17:25:50Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-02-27T14:30:26Z","2022-03-29T17:25:50Z"]},{"key":"dc:date.issued","label":"Date","values":["2020-01-01"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Applied Bioscience"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Ontario Institute of Technology"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computational Attention","Convolutional Neural Networks","Recurrent Neural Networks","Multi-Digit Classification","CAPTCHA"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10155/1132"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["We propose a computational attention approach to localize and classify characters in a sequence in a given image. Our approach combines spatial soft-attention with attention regularization and learns “where-to-look” to carry out the sequence classification task. The image is first passed through a Convolutional Neural Network (CNN) that serves as feature extractor. Then at each Recurrent Neural Network (RNN) time step, the attention mechanism attends to the relevant features sequentially to make predictions. The attention mechanism also includes a start and stop state, which instructs the mechanism to start looking and guides it when to stop (e.g., when the sequence has been exhausted). We demonstrate our approach on two sequence detection tasks—multi-digit classification and CAPTCHA unlocking—using the publicly available Street View House Numbers (SVHN) dataset and a custom CAPTCHA dataset. The experiments confirm our hypothesis that the network learns to attend to relevant features by minimizing the loss between the ground truth attention masks and the predicted attention masks."]},{"key":"dc:title","label":"Title","values":["Multi-character prediction using attention"]}]}],"canonical_facts":{"dc:contributor.advisor":["Qureshi, Faisal"],"dc:creator":["Baeenh, Mohmmed"],"dc:date.accessioned":["2020-02-27T14:30:26Z","2022-03-29T17:25:50Z"],"dc:date.available":["2020-02-27T14:30:26Z","2022-03-29T17:25:50Z"],"dc:date.issued":["2020-01-01"],"dc:description.abstract":["We propose a computational attention approach to localize and classify characters in a sequence in a given image. Our approach combines spatial soft-attention with attention regularization and learns “where-to-look” to carry out the sequence classification task. The image is first passed through a Convolutional Neural Network (CNN) that serves as feature extractor. Then at each Recurrent Neural Network (RNN) time step, the attention mechanism attends to the relevant features sequentially to make predictions. The attention mechanism also includes a start and stop state, which instructs the mechanism to start looking and guides it when to stop (e.g., when the sequence has been exhausted). We demonstrate our approach on two sequence detection tasks—multi-digit classification and CAPTCHA unlocking—using the publicly available Street View House Numbers (SVHN) dataset and a custom CAPTCHA dataset. The experiments confirm our hypothesis that the network learns to attend to relevant features by minimizing the loss between the ground truth attention masks and the predicted attention masks."],"dc:identifier.uri":["https://hdl.handle.net/10155/1132"],"dc:language.iso":["en"],"dc:subject":["Computational Attention","Convolutional Neural Networks","Recurrent Neural Networks","Multi-Digit Classification","CAPTCHA"],"dc:title":["Multi-character prediction using attention"],"dc:type":["Thesis"],"thesis:degree_discipline":["Applied Bioscience"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["University of Ontario Institute of Technology"]},"updated_at":"2026-07-24T05:35:16Z"}