{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116293"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116293","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Echelon; meaningful feature extraction and clustering on SQL queries","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Weston, Matthew Charles"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Alawini, Abdu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:56Z","subjects":["SQL","Clustering"],"languages":["en","eng"],"rights":["Copyright 2022 Matthew Weston"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116293","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Alawini, Abdu"]},{"key":"dc:creator","label":"Author","values":["Weston, Matthew Charles"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-22"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["SQL","Clustering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Matthew Weston"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116293"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Matthew Weston, accepted the attached license on 2022-07-22 at 15:09.","The student, Matthew Weston, submitted this Thesis for approval on 2022-07-22 at 15:13.","This Thesis was approved for publication on 2022-07-22 at 15:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18429 on 2022-11-15 at 18:22:10","A core part of Computer Science education, and Database Systems education in particular, is the use of machine problems to both develop and assess students’ abilities. In order to conserve resources, these assignments are often automatically graded via auto-grading systems that verify that they produce the correct outputs. Unfortunately, these systems lack essential insights into the approaches students use to solve the assignments being graded, allowing subtle flaws in student intuition to go unseen. Furthermore, manual analysis of students’ code submissions at scale ranges from costly to impossible, depending on course size and assignment frequency, making these drawbacks difficult to avoid. In this thesis paper, we rigorously define a series of metrics for evaluating a system that captures nuances in students’ approach, and then make use of these metrics to develop a system that is capable of serving as a significant force multiplier for Computer Science faculty. This system, Echelon, functions by extracting features that instructors deem significant from students’ SQL queries and using them to generate clusters that capture the key approaches taken, and then projecting these clusters to an interactive dashboard that can be used to help teaching staff quickly identify the major trends in students’ approaches to a problem. We conclude with a full analysis of Echelon on real data."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Echelon; meaningful feature extraction and clustering on SQL queries"]}]}],"canonical_facts":{"dc:contributor":["Alawini, Abdu"],"dc:creator":["Weston, Matthew Charles"],"dc:date":["2022-08","2022-07-22"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Matthew Weston, accepted the attached license on 2022-07-22 at 15:09.","The student, Matthew Weston, submitted this Thesis for approval on 2022-07-22 at 15:13.","This Thesis was approved for publication on 2022-07-22 at 15:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18429 on 2022-11-15 at 18:22:10","A core part of Computer Science education, and Database Systems education in particular, is the use of machine problems to both develop and assess students’ abilities. In order to conserve resources, these assignments are often automatically graded via auto-grading systems that verify that they produce the correct outputs. Unfortunately, these systems lack essential insights into the approaches students use to solve the assignments being graded, allowing subtle flaws in student intuition to go unseen. Furthermore, manual analysis of students’ code submissions at scale ranges from costly to impossible, depending on course size and assignment frequency, making these drawbacks difficult to avoid. In this thesis paper, we rigorously define a series of metrics for evaluating a system that captures nuances in students’ approach, and then make use of these metrics to develop a system that is capable of serving as a significant force multiplier for Computer Science faculty. This system, Echelon, functions by extracting features that instructors deem significant from students’ SQL queries and using them to generate clusters that capture the key approaches taken, and then projecting these clusters to an interactive dashboard that can be used to help teaching staff quickly identify the major trends in students’ approaches to a problem. We conclude with a full analysis of Echelon on real data."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116293"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Matthew Weston"],"dc:subject":["SQL","Clustering"],"dc:title":["Echelon; meaningful feature extraction and clustering on SQL queries"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}