{"id":{"repo_id":"dundee","oai_identifier":"oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"},"canonical_url":"https://search.dev.ndltd.org/etd/dundee/oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f","repository":{"repo_id":"dundee","name":"University of Dundee","base_url":"https://discovery.dundee.ac.uk/ws/oai"},"display":{"title":"Computational structure analysis and prediction of Ser/Thr modified by O-GlcNAc in human proteins","abstract":"This thesis studies the post-translational modifications of proteins by O-GlcNAcylation with a computational biology approach. The O-GlcNAc transferase (OGT), the enzyme that catalyses the protein O-GlcNAcylation, targets specific Serines and Threonines (S/T) of intracellular proteins. However, while other post-translationally modified residues, including phosphorylated ones, occur within sites distinguished by their amino acid sequences, less than 25% of known O-GlcNAc sites match to a sequence pattern. The small signal on the sequence patterns of multiple sites leads to the question whether the sites’ structure defines the pattern recognised by OGT.<br/><br/>The thesis then focuses on the structural features of the modified sites that could help distinguish potential sites from non-modifiable ones. 1622 O-GlcNAc sites were collected from the scientific literature. Next, 143 sites were mapped to protein 3D structure in the PDB. Modified S/T were 1.7 times more likely than unmodified S/T in the same protein to be annotated in the REMARK465 field of the PDB file, which defines missing regions in the protein structure, suggesting that these sites may be in structurally disordered regions. Clustering the structure of O-GlcNAc sites leads to ten distinct groups indicating the sites’ structural diversity. The study was extended by the analysis of features predicted from the sequence of O-GlcNAcylated proteins with Jpred4 and 3 disorder predictors, DisEMBL, IUpred and JRonn. Overall, disorder scores and proportion of S/T in coils confirmed that O-GlcNAc sites tend to be disordered.<br/><br/>A new classifier for O-GlcNAc-site (POGSPSF) was developed and trained with sequence, predicted secondary structure and disordered from 1 283 non-redundant sites. The POGSPSF Random Forest model achieved 71% area under the ROC curve in a blind test. Predictions were applied to around 2.5 million S/T in the human proteome. Nuclear and cytoplasmic protein were over-represented among the top ranking proteins. Top scoring sites were also more likely to be phosphorylated. Also, novel and potential proteins were identified within the predictions.","abstract_html":"This thesis studies the post-translational modifications of proteins by O-GlcNAcylation with a computational biology approach. The O-GlcNAc transferase (OGT), the enzyme that catalyses the protein O-GlcNAcylation, targets specific Serines and Threonines (S/T) of intracellular proteins. However, while other post-translationally modified residues, including phosphorylated ones, occur within sites distinguished by their amino acid sequences, less than 25% of known O-GlcNAc sites match to a sequence pattern. The small signal on the sequence patterns of multiple sites leads to the question whether the sites’ structure defines the pattern recognised by OGT.&lt;br/&gt;&lt;br/&gt;The thesis then focuses on the structural features of the modified sites that could help distinguish potential sites from non-modifiable ones. 1622 O-GlcNAc sites were collected from the scientific literature. Next, 143 sites were mapped to protein 3D structure in the PDB. Modified S/T were 1.7 times more likely than unmodified S/T in the same protein to be annotated in the REMARK465 field of the PDB file, which defines missing regions in the protein structure, suggesting that these sites may be in structurally disordered regions. Clustering the structure of O-GlcNAc sites leads to ten distinct groups indicating the sites’ structural diversity. The study was extended by the analysis of features predicted from the sequence of O-GlcNAcylated proteins with Jpred4 and 3 disorder predictors, DisEMBL, IUpred and JRonn. Overall, disorder scores and proportion of S/T in coils confirmed that O-GlcNAc sites tend to be disordered.&lt;br/&gt;&lt;br/&gt;A new classifier for O-GlcNAc-site (POGSPSF) was developed and trained with sequence, predicted secondary structure and disordered from 1 283 non-redundant sites. The POGSPSF Random Forest model achieved 71% area under the ROC curve in a blind test. Predictions were applied to around 2.5 million S/T in the human proteome. Nuclear and cytoplasmic protein were over-represented among the top ranking proteins. Top scoring sites were also more likely to be phosphorylated. Also, novel and potential proteins were identified within the predictions.","abstract_has_math":false,"creators":["Britto Borges, Thiago"],"institution":"University of Dundee","degree_name":"Doctor of Philosophy","degree_level":"Doctoral Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Barton, Geoffrey"],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016","date_published":"2016","updated_at":"2026-07-24T02:08:26Z","subjects":["Computational biology","Structural biology","Bioinformatics","Data Science","Machine learning","O-GlcNAc"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"],"render_values":[{"text":"oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f","href":null,"code":true}]}]},"links":{"outbound_url":"https://discovery.dundee.ac.uk/en/studentTheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Barton, Geoffrey"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Coordenação de Aperfeiçoamento de Pessoal de Nível Superior"]},{"key":"dc:creator","label":"Author","values":["Britto Borges, Thiago"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016"]},{"key":"dc:date.issued","label":"Date","values":["2016"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Computational Biology"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Dundee"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://discovery.dundee.ac.uk/en/studentTheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral Thesis"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computational biology","Structural biology","Bioinformatics","Data Science","Machine learning","O-GlcNAc"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2019-06-30"]},{"key":"dc:rights.embargoreason","label":"Dc Rights Embargoreason","values":["/dk/atira/pure/core/document/studentthesisembargoreason/commercialexploitation"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f","https://discovery.dundee.ac.uk/en/studentTheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://discovery.dundee.ac.uk/files/12078302/tbrittoborges_thesis_revised.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis studies the post-translational modifications of proteins by O-GlcNAcylation with a computational biology approach. The O-GlcNAc transferase (OGT), the enzyme that catalyses the protein O-GlcNAcylation, targets specific Serines and Threonines (S/T) of intracellular proteins. However, while other post-translationally modified residues, including phosphorylated ones, occur within sites distinguished by their amino acid sequences, less than 25% of known O-GlcNAc sites match to a sequence pattern. The small signal on the sequence patterns of multiple sites leads to the question whether the sites’ structure defines the pattern recognised by OGT.<br/><br/>The thesis then focuses on the structural features of the modified sites that could help distinguish potential sites from non-modifiable ones. 1622 O-GlcNAc sites were collected from the scientific literature. Next, 143 sites were mapped to protein 3D structure in the PDB. Modified S/T were 1.7 times more likely than unmodified S/T in the same protein to be annotated in the REMARK465 field of the PDB file, which defines missing regions in the protein structure, suggesting that these sites may be in structurally disordered regions. Clustering the structure of O-GlcNAc sites leads to ten distinct groups indicating the sites’ structural diversity. The study was extended by the analysis of features predicted from the sequence of O-GlcNAcylated proteins with Jpred4 and 3 disorder predictors, DisEMBL, IUpred and JRonn. Overall, disorder scores and proportion of S/T in coils confirmed that O-GlcNAc sites tend to be disordered.<br/><br/>A new classifier for O-GlcNAc-site (POGSPSF) was developed and trained with sequence, predicted secondary structure and disordered from 1 283 non-redundant sites. The POGSPSF Random Forest model achieved 71% area under the ROC curve in a blind test. Predictions were applied to around 2.5 million S/T in the human proteome. Nuclear and cytoplasmic protein were over-represented among the top ranking proteins. Top scoring sites were also more likely to be phosphorylated. Also, novel and potential proteins were identified within the predictions."]},{"key":"dc:title","label":"Title","values":["Computational structure analysis and prediction of Ser/Thr modified by O-GlcNAc in human proteins"]}]}],"canonical_facts":{"dc:contributor.advisor":["Barton, Geoffrey"],"dc:contributor.sponsor":["Coordenação de Aperfeiçoamento de Pessoal de Nível Superior"],"dc:creator":["Britto Borges, Thiago"],"dc:date":["2016"],"dc:date.issued":["2016"],"dc:description.abstract":["This thesis studies the post-translational modifications of proteins by O-GlcNAcylation with a computational biology approach. The O-GlcNAc transferase (OGT), the enzyme that catalyses the protein O-GlcNAcylation, targets specific Serines and Threonines (S/T) of intracellular proteins. However, while other post-translationally modified residues, including phosphorylated ones, occur within sites distinguished by their amino acid sequences, less than 25% of known O-GlcNAc sites match to a sequence pattern. The small signal on the sequence patterns of multiple sites leads to the question whether the sites’ structure defines the pattern recognised by OGT.<br/><br/>The thesis then focuses on the structural features of the modified sites that could help distinguish potential sites from non-modifiable ones. 1622 O-GlcNAc sites were collected from the scientific literature. Next, 143 sites were mapped to protein 3D structure in the PDB. Modified S/T were 1.7 times more likely than unmodified S/T in the same protein to be annotated in the REMARK465 field of the PDB file, which defines missing regions in the protein structure, suggesting that these sites may be in structurally disordered regions. Clustering the structure of O-GlcNAc sites leads to ten distinct groups indicating the sites’ structural diversity. The study was extended by the analysis of features predicted from the sequence of O-GlcNAcylated proteins with Jpred4 and 3 disorder predictors, DisEMBL, IUpred and JRonn. Overall, disorder scores and proportion of S/T in coils confirmed that O-GlcNAc sites tend to be disordered.<br/><br/>A new classifier for O-GlcNAc-site (POGSPSF) was developed and trained with sequence, predicted secondary structure and disordered from 1 283 non-redundant sites. The POGSPSF Random Forest model achieved 71% area under the ROC curve in a blind test. Predictions were applied to around 2.5 million S/T in the human proteome. Nuclear and cytoplasmic protein were over-represented among the top ranking proteins. Top scoring sites were also more likely to be phosphorylated. Also, novel and potential proteins were identified within the predictions."],"dc:identifier":["oai:discovery.dundee.ac.uk:studenttheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f","https://discovery.dundee.ac.uk/en/studentTheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"],"dc:identifier.uri":["https://discovery.dundee.ac.uk/files/12078302/tbrittoborges_thesis_revised.pdf"],"dc:language":["eng"],"dc:publisher.department":["Computational Biology"],"dc:publisher.institution":["University of Dundee"],"dc:relation.isreferencedby":["https://discovery.dundee.ac.uk/en/studentTheses/15f1c78d-8f24-4e53-a3e4-ac97707b903f"],"dc:rights.embargodate":["2019-06-30"],"dc:rights.embargoreason":["/dk/atira/pure/core/document/studentthesisembargoreason/commercialexploitation"],"dc:subject":["Computational biology","Structural biology","Bioinformatics","Data Science","Machine learning","O-GlcNAc"],"dc:title":["Computational structure analysis and prediction of Ser/Thr modified by O-GlcNAc in human proteins"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral Thesis"],"dc:type.qualificationname":["Doctor of Philosophy"]},"updated_at":"2026-07-24T02:08:26Z"}