{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113273"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113273","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The effects of character predictability on eye-movement control in reading Mandarin Chinese texts","abstract":"All the major reading theories, which are primarily based on studies of alphabetic scripts, hypothesize that “word” is the basic unit of processing, but there is no consensus on whether this assumption is also true in Chinese reading. As word boundaries are not visually denoted in Chinese text, “word” is hard to define. Chinese readers do not agree with each other, and often do not agree with themselves if asked to segment words in the same text twice (Lin et al., 2011; Liu et al., 2013). Moreover, without inter-word spaces, it is also unclear how Chinese readers segment words in character strings during online reading. In contrast, characters are equally spaced and can be straightforwardly defined. There is also evidence showing that character properties can influence eye-movements. However, the studies on how character properties influence eye movements have been focused on character frequency and visual complexity. The effects of character predictability have been largely ignored. To fill this gap in the literature, this dissertation investigated how cumulative character- by-character predictability values influenced eye-movement control during Chinese reading. In addition, it proposed and explored evidence for a new hypothesis, the character-property-based word segmentation hypothesis: that character properties can provide word boundary information that enables readers to segment words during online reading. Two experiments were designed for these purposes. First, a large-scale character-by- character running cloze task (cf. Luke & Christianson, 2016) was implemented to collect the character predictability data for the characters (4640 characters in total) in 40 short paragraphs. 40 responses were collected for each paragraph (172 participants total). Second, another 102 participants were recruited for an eye-tracking experiment, in which participants read the same passages used in the cloze surveys. When analyzing the eye-tracking data, two information complexity metrics -- character surprisal and entropy reduction -- were computed from the results of the cloze surveys (cf. Lowder et al., 2018) and were used as measures for character predictability. In several analyses, the results showed significant effects for both, suggesting that character predictability plays an important role in eye-movement control. In most cases, informationally more complex characters were viewed longer and skipped less. But in the analysis for characters in content words, higher character surprisal values were associated with higher skipping probabilities, which was in the opposite direction to the effect of word surprisal, indicating that the impacts of character surprisal and word surprisal were different. A post hoc analysis also implied that character surprisal and entropy reduction might function not only as information complexity metrics, but also as word boundary indicators which guided the readers to land their saccades at word-middle positions. Finally, the character-property-based word segmentation hypothesis was supported by the findings that a linear discriminant analysis model, trained by only character variables, was able to predict word boundaries fairly well (M = 84.66%, for word boundaries before characters; M = 70.62%, for word boundaries after); and that when character properties were included in the analyses, explicitly coding whether there was a word boundary before a character did not explain more variance in the reading time and skipping probability, whereas in the baseline model it had significant effects. These results suggest that the word boundary information provided by a character’s properties is sufficient for judging whether a word begins at that character. While word properties had robust effects in all analyses, the findings in the current study suggest that “word” is not the only important processing unit in reading Chinese. Character properties have independent influences on eye-movement control and may be used for online word segmentation.","abstract_html":"All the major reading theories, which are primarily based on studies of alphabetic scripts, hypothesize that “word” is the basic unit of processing, but there is no consensus on whether this assumption is also true in Chinese reading. As word boundaries are not visually denoted in Chinese text, “word” is hard to define. Chinese readers do not agree with each other, and often do not agree with themselves if asked to segment words in the same text twice (Lin et al., 2011; Liu et al., 2013). Moreover, without inter-word spaces, it is also unclear how Chinese readers segment words in character strings during online reading. In contrast, characters are equally spaced and can be straightforwardly defined. There is also evidence showing that character properties can influence eye-movements. However, the studies on how character properties influence eye movements have been focused on character frequency and visual complexity. The effects of character predictability have been largely ignored. To fill this gap in the literature, this dissertation investigated how cumulative character- by-character predictability values influenced eye-movement control during Chinese reading. In addition, it proposed and explored evidence for a new hypothesis, the character-property-based word segmentation hypothesis: that character properties can provide word boundary information that enables readers to segment words during online reading. Two experiments were designed for these purposes. First, a large-scale character-by- character running cloze task (cf. Luke &amp; Christianson, 2016) was implemented to collect the character predictability data for the characters (4640 characters in total) in 40 short paragraphs. 40 responses were collected for each paragraph (172 participants total). Second, another 102 participants were recruited for an eye-tracking experiment, in which participants read the same passages used in the cloze surveys. When analyzing the eye-tracking data, two information complexity metrics -- character surprisal and entropy reduction -- were computed from the results of the cloze surveys (cf. Lowder et al., 2018) and were used as measures for character predictability. In several analyses, the results showed significant effects for both, suggesting that character predictability plays an important role in eye-movement control. In most cases, informationally more complex characters were viewed longer and skipped less. But in the analysis for characters in content words, higher character surprisal values were associated with higher skipping probabilities, which was in the opposite direction to the effect of word surprisal, indicating that the impacts of character surprisal and word surprisal were different. A post hoc analysis also implied that character surprisal and entropy reduction might function not only as information complexity metrics, but also as word boundary indicators which guided the readers to land their saccades at word-middle positions. Finally, the character-property-based word segmentation hypothesis was supported by the findings that a linear discriminant analysis model, trained by only character variables, was able to predict word boundaries fairly well (M = 84.66%, for word boundaries before characters; M = 70.62%, for word boundaries after); and that when character properties were included in the analyses, explicitly coding whether there was a word boundary before a character did not explain more variance in the reading time and skipping probability, whereas in the baseline model it had significant effects. These results suggest that the word boundary information provided by a character’s properties is sufficient for judging whether a word begins at that character. While word properties had robust effects in all analyses, the findings in the current study suggest that “word” is not the only important processing unit in reading Chinese. Character properties have independent influences on eye-movement control and may be used for online word segmentation.","abstract_has_math":false,"creators":["Li, You"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"E Asian Languages & Cultures","degree_department":null,"school":null,"contributors":["Packard, Jerome","Christianson, Kiel","Shih, Chilin","Sadler, Misumi"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T22:54:09Z","date_published":"2022-01-12T22:54:09Z","updated_at":"2026-07-22T22:24:53Z","subjects":["Mandarin","reading","character","word","predictability","entropy reduction","surprisal","eye-movement"],"languages":["en"],"rights":["Copyright 2021 You Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113273","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Packard, Jerome","Christianson, Kiel","Shih, Chilin","Sadler, Misumi"]},{"key":"dc:creator","label":"Author","values":["Li, You"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T22:54:09Z","2024-01-12T22:56:20Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["E Asian Languages & Cultures"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Mandarin","reading","character","word","predictability","entropy reduction","surprisal","eye-movement"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 You Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113273"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["All the major reading theories, which are primarily based on studies of alphabetic scripts, hypothesize that “word” is the basic unit of processing, but there is no consensus on whether this assumption is also true in Chinese reading. As word boundaries are not visually denoted in Chinese text, “word” is hard to define. Chinese readers do not agree with each other, and often do not agree with themselves if asked to segment words in the same text twice (Lin et al., 2011; Liu et al., 2013). Moreover, without inter-word spaces, it is also unclear how Chinese readers segment words in character strings during online reading. In contrast, characters are equally spaced and can be straightforwardly defined. There is also evidence showing that character properties can influence eye-movements. However, the studies on how character properties influence eye movements have been focused on character frequency and visual complexity. The effects of character predictability have been largely ignored. To fill this gap in the literature, this dissertation investigated how cumulative character- by-character predictability values influenced eye-movement control during Chinese reading. In addition, it proposed and explored evidence for a new hypothesis, the character-property-based word segmentation hypothesis: that character properties can provide word boundary information that enables readers to segment words during online reading. Two experiments were designed for these purposes. First, a large-scale character-by- character running cloze task (cf. Luke & Christianson, 2016) was implemented to collect the character predictability data for the characters (4640 characters in total) in 40 short paragraphs. 40 responses were collected for each paragraph (172 participants total). Second, another 102 participants were recruited for an eye-tracking experiment, in which participants read the same passages used in the cloze surveys. When analyzing the eye-tracking data, two information complexity metrics -- character surprisal and entropy reduction -- were computed from the results of the cloze surveys (cf. Lowder et al., 2018) and were used as measures for character predictability. In several analyses, the results showed significant effects for both, suggesting that character predictability plays an important role in eye-movement control. In most cases, informationally more complex characters were viewed longer and skipped less. But in the analysis for characters in content words, higher character surprisal values were associated with higher skipping probabilities, which was in the opposite direction to the effect of word surprisal, indicating that the impacts of character surprisal and word surprisal were different. A post hoc analysis also implied that character surprisal and entropy reduction might function not only as information complexity metrics, but also as word boundary indicators which guided the readers to land their saccades at word-middle positions. Finally, the character-property-based word segmentation hypothesis was supported by the findings that a linear discriminant analysis model, trained by only character variables, was able to predict word boundaries fairly well (M = 84.66%, for word boundaries before characters; M = 70.62%, for word boundaries after); and that when character properties were included in the analyses, explicitly coding whether there was a word boundary before a character did not explain more variance in the reading time and skipping probability, whereas in the baseline model it had significant effects. These results suggest that the word boundary information provided by a character’s properties is sufficient for judging whether a word begins at that character. While word properties had robust effects in all analyses, the findings in the current study suggest that “word” is not the only important processing unit in reading Chinese. Character properties have independent influences on eye-movement control and may be used for online word segmentation.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2023-08-01","The student, You Li, accepted the attached license on 2021-07-07 at 17:09.","The student, You Li, submitted this Dissertation for approval on 2021-07-07 at 17:15.","This Dissertation was approved for publication on 2021-07-12 at 15:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16761 on 2022-01-12 at 13:03:55","Made available in DSpace on 2022-01-12T22:54:09Z (GMT). No. of bitstreams: 3 LI-DISSERTATION-2021.pdf: 4868651 bytes, checksum: e24d6c032ca6df2d947664e6a0d4d7a5 (MD5) LICENSE.txt: 4203 bytes, checksum: d157268a2d445612da61073848211672 (MD5) PROQUEST_LICENSE.txt: 4549 bytes, checksum: 879f754f9aaebe033ee1a6761470bc58 (MD5) Previous issue date: 2021-07-12","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:54:14Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:55:09Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:56:20Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The effects of character predictability on eye-movement control in reading Mandarin Chinese texts"]}]}],"canonical_facts":{"dc:contributor":["Packard, Jerome","Christianson, Kiel","Shih, Chilin","Sadler, Misumi"],"dc:creator":["Li, You"],"dc:date":["2022-01-12T22:54:09Z","2024-01-12T22:56:20Z","2021-07-12","2021-08"],"dc:description":["All the major reading theories, which are primarily based on studies of alphabetic scripts, hypothesize that “word” is the basic unit of processing, but there is no consensus on whether this assumption is also true in Chinese reading. As word boundaries are not visually denoted in Chinese text, “word” is hard to define. Chinese readers do not agree with each other, and often do not agree with themselves if asked to segment words in the same text twice (Lin et al., 2011; Liu et al., 2013). Moreover, without inter-word spaces, it is also unclear how Chinese readers segment words in character strings during online reading. In contrast, characters are equally spaced and can be straightforwardly defined. There is also evidence showing that character properties can influence eye-movements. However, the studies on how character properties influence eye movements have been focused on character frequency and visual complexity. The effects of character predictability have been largely ignored. To fill this gap in the literature, this dissertation investigated how cumulative character- by-character predictability values influenced eye-movement control during Chinese reading. In addition, it proposed and explored evidence for a new hypothesis, the character-property-based word segmentation hypothesis: that character properties can provide word boundary information that enables readers to segment words during online reading. Two experiments were designed for these purposes. First, a large-scale character-by- character running cloze task (cf. Luke & Christianson, 2016) was implemented to collect the character predictability data for the characters (4640 characters in total) in 40 short paragraphs. 40 responses were collected for each paragraph (172 participants total). Second, another 102 participants were recruited for an eye-tracking experiment, in which participants read the same passages used in the cloze surveys. When analyzing the eye-tracking data, two information complexity metrics -- character surprisal and entropy reduction -- were computed from the results of the cloze surveys (cf. Lowder et al., 2018) and were used as measures for character predictability. In several analyses, the results showed significant effects for both, suggesting that character predictability plays an important role in eye-movement control. In most cases, informationally more complex characters were viewed longer and skipped less. But in the analysis for characters in content words, higher character surprisal values were associated with higher skipping probabilities, which was in the opposite direction to the effect of word surprisal, indicating that the impacts of character surprisal and word surprisal were different. A post hoc analysis also implied that character surprisal and entropy reduction might function not only as information complexity metrics, but also as word boundary indicators which guided the readers to land their saccades at word-middle positions. Finally, the character-property-based word segmentation hypothesis was supported by the findings that a linear discriminant analysis model, trained by only character variables, was able to predict word boundaries fairly well (M = 84.66%, for word boundaries before characters; M = 70.62%, for word boundaries after); and that when character properties were included in the analyses, explicitly coding whether there was a word boundary before a character did not explain more variance in the reading time and skipping probability, whereas in the baseline model it had significant effects. These results suggest that the word boundary information provided by a character’s properties is sufficient for judging whether a word begins at that character. While word properties had robust effects in all analyses, the findings in the current study suggest that “word” is not the only important processing unit in reading Chinese. Character properties have independent influences on eye-movement control and may be used for online word segmentation.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2023-08-01","The student, You Li, accepted the attached license on 2021-07-07 at 17:09.","The student, You Li, submitted this Dissertation for approval on 2021-07-07 at 17:15.","This Dissertation was approved for publication on 2021-07-12 at 15:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16761 on 2022-01-12 at 13:03:55","Made available in DSpace on 2022-01-12T22:54:09Z (GMT). No. of bitstreams: 3 LI-DISSERTATION-2021.pdf: 4868651 bytes, checksum: e24d6c032ca6df2d947664e6a0d4d7a5 (MD5) LICENSE.txt: 4203 bytes, checksum: d157268a2d445612da61073848211672 (MD5) PROQUEST_LICENSE.txt: 4549 bytes, checksum: 879f754f9aaebe033ee1a6761470bc58 (MD5) Previous issue date: 2021-07-12","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:54:14Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:55:09Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 121200 Lift date: 2024-01-12T22:56:20Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113273"],"dc:language":["en"],"dc:rights":["Copyright 2021 You Li"],"dc:subject":["Mandarin","reading","character","word","predictability","entropy reduction","surprisal","eye-movement"],"dc:title":["The effects of character predictability on eye-movement control in reading Mandarin Chinese texts"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["E Asian Languages & Cultures"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}