Back to results

Universiti Tun Hussein Onn Malaysia

A corpus-based lexical and grammatical error identification: L2 learners academic writing

Abstract

dc:description.abstract

Writing in English has never been an easy task to many second language (L2) learners. Many of them perform poorly in their English academic writing where numerous lexical and grammatical errors are found in their report. Therefore, this thesis investigates the difficulties faced by UTHM learners involved in academic writing by identifying and analyzing errors made by them with the application of error analysis procedures. This research attempts to find out the types and patterns of errors in which it focuses on the frequency of the lexical and grammatical errors of the L2 learners in their writing. Errors were investigated and identified based on students’ 36 progress and final reports which were assembled from first year engineering students; named as the Learner Corpus Universiti Tun Hussein Onn Malaysia(LCUTHM). The LCUTHM was analyzed by means of linguistics Natural Language Processing tools (NLP) such as CLAWS 5 tag set, Markin Version 4 and categorized by MonoConc Pro II in the form of word lists. Data were also analyzed using Statistical Package for the Social Science (SPSS) software to determine the major errors learners committed in learners’ written work. The findings reveal that the major lexical and grammatical error categories made by learners were “Missing Word”, “Repetition”, and “Verb Form”. Finally, the integration of technology and the linguistics Natural Language Processing (NLP) tools can provide a fast and more effective method in assisting teachers in identifying errors, and designing syllabus in improving the language skills and achievement of L2 learners in their academic writing.

Degree

thesis:*
Name dc:type.qualificationname
phd
Level dc:type.qualificationlevel
doctoral
Grantor dc:publisher.institution
Universiti Tun Hussein Onn Malaysia
Year dc:date.issued
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Talib, Salleh

Subjects

dc:subject × 1

Rights

Language dc:language
en

Chain of custody

source
Harvested from
Universiti Tun Hussein Onn Malaysia
Base URL
eprints.uthm.edu.my/cgi/oai2
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Talib, Salleh. A corpus-based lexical and grammatical error identification: L2 learners academic writing. doctoral thesis, Universiti Tun Hussein Onn Malaysia, 2021.