{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/99126"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/99126","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards an end-to-end music transcription system using neural networks","abstract":"Transcription is the task of writing down instructions on how to play a particular piece of music, including individual notes, note durations, embellishments and so on. While most major works in the traditional repertoire have readily available transcriptions for various instrument arrangements, this is not as common in genres where improvisation is more prevalent, such as Jazz, or where the piece has a very particular purpose, as in motion picture and video game soundtracks. It has notable parallels with the task of Automatic Speech Recognition (ASR) and indeed from this connection arises some natural Machine Learning-based approaches. However, these methods usually involve carefully designed preprocessing steps, or transcription into less flexible representations, such as piano rolls, which are harder to read for humans. This work investigates the feasibility of designing an end-to-end music transcription system that takes in raw audio recordings and produces Lilypond notation, which can directly generate easily-recognizable sheet music. In keeping with modern ASR methods, this task is modeled as a sequence-to-sequence problem using Convolutional and Recurrent Neural Networks. The system is shown to perform well for both monophonic (single melody on a single instrument) and polyphonic music (parallel melodies on possibly different instruments) for randomly generated pieces played by the piano and various other common orchestra instruments.","abstract_html":"Transcription is the task of writing down instructions on how to play a particular piece of music, including individual notes, note durations, embellishments and so on. While most major works in the traditional repertoire have readily available transcriptions for various instrument arrangements, this is not as common in genres where improvisation is more prevalent, such as Jazz, or where the piece has a very particular purpose, as in motion picture and video game soundtracks. It has notable parallels with the task of Automatic Speech Recognition (ASR) and indeed from this connection arises some natural Machine Learning-based approaches. However, these methods usually involve carefully designed preprocessing steps, or transcription into less flexible representations, such as piano rolls, which are harder to read for humans. This work investigates the feasibility of designing an end-to-end music transcription system that takes in raw audio recordings and produces Lilypond notation, which can directly generate easily-recognizable sheet music. In keeping with modern ASR methods, this task is modeled as a sequence-to-sequence problem using Convolutional and Recurrent Neural Networks. The system is shown to perform well for both monophonic (single melody on a single instrument) and polyphonic music (parallel melodies on possibly different instruments) for randomly generated pieces played by the piano and various other common orchestra instruments.","abstract_has_math":false,"creators":["Correa Carvalho, Ralf Gunter"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Smaragdis, Paris"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-03-02T19:59:47Z","date_published":"2018-03-02T19:59:47Z","updated_at":"2026-07-22T22:24:37Z","subjects":["Automatic music transcription","Machine learning","Deep learning","Signal processing"],"languages":["en"],"rights":["Copyright 2017 Ralf Gunter Correa Carvalho"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/99126","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Smaragdis, Paris"]},{"key":"dc:creator","label":"Author","values":["Correa Carvalho, Ralf Gunter"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-03-02T19:59:47Z","2020-03-03T10:15:32Z","2017-07-19","2017-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automatic music transcription","Machine learning","Deep learning","Signal processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Ralf Gunter Correa Carvalho"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/99126"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Transcription is the task of writing down instructions on how to play a particular piece of music, including individual notes, note durations, embellishments and so on. While most major works in the traditional repertoire have readily available transcriptions for various instrument arrangements, this is not as common in genres where improvisation is more prevalent, such as Jazz, or where the piece has a very particular purpose, as in motion picture and video game soundtracks. It has notable parallels with the task of Automatic Speech Recognition (ASR) and indeed from this connection arises some natural Machine Learning-based approaches. However, these methods usually involve carefully designed preprocessing steps, or transcription into less flexible representations, such as piano rolls, which are harder to read for humans. This work investigates the feasibility of designing an end-to-end music transcription system that takes in raw audio recordings and produces Lilypond notation, which can directly generate easily-recognizable sheet music. In keeping with modern ASR methods, this task is modeled as a sequence-to-sequence problem using Convolutional and Recurrent Neural Networks. The system is shown to perform well for both monophonic (single melody on a single instrument) and polyphonic music (parallel melodies on possibly different instruments) for randomly generated pieces played by the piano and various other common orchestra instruments.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-08-01","The student, Ralf Gunter Correa Carvalho, accepted the attached license on 2017-07-19 at 11:04.","The student, Ralf Gunter Correa Carvalho, submitted this Thesis for approval on 2017-07-19 at 11:08.","This Thesis was approved for publication on 2017-07-19 at 11:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11528 on 2018-03-02 at 13:02:41","Made available in DSpace on 2018-03-02T19:59:47Z (GMT). No. of bitstreams: 2 CORREACARVALHO-THESIS-2017.pdf: 3851001 bytes, checksum: e3fda38e910b10ac5509e311c2a77180 (MD5) LICENSE.txt: 4224 bytes, checksum: 0aca188d2e750205f7e19435aa3991bc (MD5) Previous issue date: 2017-07-19","Embargo set by: Seth Robbins for item 105080 Lift date: 2020-03-02T19:59:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105080 Lift date: 2020-03-02T20:02:46Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105080 on 2020-03-03T10:15:32Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards an end-to-end music transcription system using neural networks"]}]}],"canonical_facts":{"dc:contributor":["Smaragdis, Paris"],"dc:creator":["Correa Carvalho, Ralf Gunter"],"dc:date":["2018-03-02T19:59:47Z","2020-03-03T10:15:32Z","2017-07-19","2017-08"],"dc:description":["Transcription is the task of writing down instructions on how to play a particular piece of music, including individual notes, note durations, embellishments and so on. While most major works in the traditional repertoire have readily available transcriptions for various instrument arrangements, this is not as common in genres where improvisation is more prevalent, such as Jazz, or where the piece has a very particular purpose, as in motion picture and video game soundtracks. It has notable parallels with the task of Automatic Speech Recognition (ASR) and indeed from this connection arises some natural Machine Learning-based approaches. However, these methods usually involve carefully designed preprocessing steps, or transcription into less flexible representations, such as piano rolls, which are harder to read for humans. This work investigates the feasibility of designing an end-to-end music transcription system that takes in raw audio recordings and produces Lilypond notation, which can directly generate easily-recognizable sheet music. In keeping with modern ASR methods, this task is modeled as a sequence-to-sequence problem using Convolutional and Recurrent Neural Networks. The system is shown to perform well for both monophonic (single melody on a single instrument) and polyphonic music (parallel melodies on possibly different instruments) for randomly generated pieces played by the piano and various other common orchestra instruments.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-08-01","The student, Ralf Gunter Correa Carvalho, accepted the attached license on 2017-07-19 at 11:04.","The student, Ralf Gunter Correa Carvalho, submitted this Thesis for approval on 2017-07-19 at 11:08.","This Thesis was approved for publication on 2017-07-19 at 11:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11528 on 2018-03-02 at 13:02:41","Made available in DSpace on 2018-03-02T19:59:47Z (GMT). No. of bitstreams: 2 CORREACARVALHO-THESIS-2017.pdf: 3851001 bytes, checksum: e3fda38e910b10ac5509e311c2a77180 (MD5) LICENSE.txt: 4224 bytes, checksum: 0aca188d2e750205f7e19435aa3991bc (MD5) Previous issue date: 2017-07-19","Embargo set by: Seth Robbins for item 105080 Lift date: 2020-03-02T19:59:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105080 Lift date: 2020-03-02T20:02:46Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105080 on 2020-03-03T10:15:32Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/99126"],"dc:language":["en"],"dc:rights":["Copyright 2017 Ralf Gunter Correa Carvalho"],"dc:subject":["Automatic music transcription","Machine learning","Deep learning","Signal processing"],"dc:title":["Towards an end-to-end music transcription system using neural networks"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:37Z"}