{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/377462"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/377462","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"The Role of Precedent in Computational Models of Law","abstract":"In common law countries, due to the doctrine of precedent, lawyers need to consider a body of case law which grows with each court decision. As a result, the complexity and number of documents a lawyer should be familiar with for each new decision increases rapidly over time. Robust models of precedential reasoning could help in mitigating this problem. Therefore, the goal of this thesis is to model precedent, understand how it works and how to use it to explain system predictions. Towards this, I will use modern techniques from natural language processing (NLP) and machine learning (ML). In my first investigation, I demonstrate how information-theoretic methods can be used to settle a long-standing jurisprudential debate, namely whether it is facts or legal arguments that have a stronger influence on case outcomes. In terms of mutual information, obtained by training a classification neural models built on top of a pre-trained language model, I find that arguments (0.38 nats) have a stronger bearing on outcomes than facts (0.18 nats). I next design a neural model of legal outcome prediction that significantly outperforms prior work. More importantly, it is also able to model negative outcomes: it can learn which law has been claimed as violated but was found not violated, as opposed to only predicting which law has been violated. Negative outcome is important because it forms legally binding precedent, and models that ignore it will inevitably model law incorrectly. My last content chapter is concerned with explainability. Rather than explaining which linguistic faculties a neural model might have learned or how individual words affect the model's behaviour, my operationalisation of this task is more informative from a legal perspective. I use influence functions, which allow me to measure the change in loss to efficiently estimate the effect of removing a single precedent from the model. The results demonstrate that existing outcome-prediction models lack the ability to employ most of the basic types of precedential reasoning. Since there are many outstanding questions about ethics in legal NLP, I elaborate on these in the final chapter of this thesis. In summary, my thesis can be seen as an exercise in applying insights from one discipline to another: modern NLP techniques can advance our understanding of the law, whereas ideas from the legal domain can be applied to build stronger computational models of law.","abstract_html":"In common law countries, due to the doctrine of precedent, lawyers need to consider a body of case law which grows with each court decision. As a result, the complexity and number of documents a lawyer should be familiar with for each new decision increases rapidly over time. Robust models of precedential reasoning could help in mitigating this problem. Therefore, the goal of this thesis is to model precedent, understand how it works and how to use it to explain system predictions. Towards this, I will use modern techniques from natural language processing (NLP) and machine learning (ML). In my first investigation, I demonstrate how information-theoretic methods can be used to settle a long-standing jurisprudential debate, namely whether it is facts or legal arguments that have a stronger influence on case outcomes. In terms of mutual information, obtained by training a classification neural models built on top of a pre-trained language model, I find that arguments (0.38 nats) have a stronger bearing on outcomes than facts (0.18 nats). I next design a neural model of legal outcome prediction that significantly outperforms prior work. More importantly, it is also able to model negative outcomes: it can learn which law has been claimed as violated but was found not violated, as opposed to only predicting which law has been violated. Negative outcome is important because it forms legally binding precedent, and models that ignore it will inevitably model law incorrectly. My last content chapter is concerned with explainability. Rather than explaining which linguistic faculties a neural model might have learned or how individual words affect the model&#x27;s behaviour, my operationalisation of this task is more informative from a legal perspective. I use influence functions, which allow me to measure the change in loss to efficiently estimate the effect of removing a single precedent from the model. The results demonstrate that existing outcome-prediction models lack the ability to employ most of the basic types of precedential reasoning. Since there are many outstanding questions about ethics in legal NLP, I elaborate on these in the final chapter of this thesis. In summary, my thesis can be seen as an exercise in applying insights from one discipline to another: modern NLP techniques can advance our understanding of the law, whereas ideas from the legal domain can be applied to build stronger computational models of law.","abstract_has_math":false,"creators":["Valvoda, Josef"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Teufel, Simone"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-06-01","date_published":"2023-06-01","updated_at":"2026-07-22T22:24:18Z","subjects":["deep learning","legal reasoning","natural language processing"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/241eb611-a660-41ec-804b-da40d5c43b6e/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.114276","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Teufel, Simone"]},{"key":"dc:creator","label":"Author","values":["Valvoda, Josef"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2023-06-01"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/377462"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["deep learning","legal reasoning","natural language processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/241eb611-a660-41ec-804b-da40d5c43b6e/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.114276"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4924d22-3b4b-497f-abe2-3c951b798164/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In common law countries, due to the doctrine of precedent, lawyers need to consider a body of case law which grows with each court decision. As a result, the complexity and number of documents a lawyer should be familiar with for each new decision increases rapidly over time. Robust models of precedential reasoning could help in mitigating this problem. Therefore, the goal of this thesis is to model precedent, understand how it works and how to use it to explain system predictions. Towards this, I will use modern techniques from natural language processing (NLP) and machine learning (ML). In my first investigation, I demonstrate how information-theoretic methods can be used to settle a long-standing jurisprudential debate, namely whether it is facts or legal arguments that have a stronger influence on case outcomes. In terms of mutual information, obtained by training a classification neural models built on top of a pre-trained language model, I find that arguments (0.38 nats) have a stronger bearing on outcomes than facts (0.18 nats). I next design a neural model of legal outcome prediction that significantly outperforms prior work. More importantly, it is also able to model negative outcomes: it can learn which law has been claimed as violated but was found not violated, as opposed to only predicting which law has been violated. Negative outcome is important because it forms legally binding precedent, and models that ignore it will inevitably model law incorrectly. My last content chapter is concerned with explainability. Rather than explaining which linguistic faculties a neural model might have learned or how individual words affect the model's behaviour, my operationalisation of this task is more informative from a legal perspective. I use influence functions, which allow me to measure the change in loss to efficiently estimate the effect of removing a single precedent from the model. The results demonstrate that existing outcome-prediction models lack the ability to employ most of the basic types of precedential reasoning. Since there are many outstanding questions about ethics in legal NLP, I elaborate on these in the final chapter of this thesis. In summary, my thesis can be seen as an exercise in applying insights from one discipline to another: modern NLP techniques can advance our understanding of the law, whereas ideas from the legal domain can be applied to build stronger computational models of law."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["87eda9de84448d1f82354d60eee3eb5f","d771728c781f9708354565fedda74c75"]},{"key":"dc:title","label":"Title","values":["The Role of Precedent in Computational Models of Law"]}]}],"canonical_facts":{"dc:contributor.advisor":["Teufel, Simone"],"dc:creator":["Valvoda, Josef"],"dc:date.issued":["2023-06-01"],"dc:description.abstract":["In common law countries, due to the doctrine of precedent, lawyers need to consider a body of case law which grows with each court decision. As a result, the complexity and number of documents a lawyer should be familiar with for each new decision increases rapidly over time. Robust models of precedential reasoning could help in mitigating this problem. Therefore, the goal of this thesis is to model precedent, understand how it works and how to use it to explain system predictions. Towards this, I will use modern techniques from natural language processing (NLP) and machine learning (ML). In my first investigation, I demonstrate how information-theoretic methods can be used to settle a long-standing jurisprudential debate, namely whether it is facts or legal arguments that have a stronger influence on case outcomes. In terms of mutual information, obtained by training a classification neural models built on top of a pre-trained language model, I find that arguments (0.38 nats) have a stronger bearing on outcomes than facts (0.18 nats). I next design a neural model of legal outcome prediction that significantly outperforms prior work. More importantly, it is also able to model negative outcomes: it can learn which law has been claimed as violated but was found not violated, as opposed to only predicting which law has been violated. Negative outcome is important because it forms legally binding precedent, and models that ignore it will inevitably model law incorrectly. My last content chapter is concerned with explainability. Rather than explaining which linguistic faculties a neural model might have learned or how individual words affect the model's behaviour, my operationalisation of this task is more informative from a legal perspective. I use influence functions, which allow me to measure the change in loss to efficiently estimate the effect of removing a single precedent from the model. The results demonstrate that existing outcome-prediction models lack the ability to employ most of the basic types of precedential reasoning. Since there are many outstanding questions about ethics in legal NLP, I elaborate on these in the final chapter of this thesis. In summary, my thesis can be seen as an exercise in applying insights from one discipline to another: modern NLP techniques can advance our understanding of the law, whereas ideas from the legal domain can be applied to build stronger computational models of law."],"dc:format.checksum.md5":["87eda9de84448d1f82354d60eee3eb5f","d771728c781f9708354565fedda74c75"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.114276"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4924d22-3b4b-497f-abe2-3c951b798164/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/377462"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/241eb611-a660-41ec-804b-da40d5c43b6e/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["deep learning","legal reasoning","natural language processing"],"dc:title":["The Role of Precedent in Computational Models of Law"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:18Z"}