Repository logo
 

Author Identification in Free Texts

Date

Supervisor

Nand, Parma

Item type

Thesis

Degree name

Master of Computer and Information Sciences

Journal Title

Journal ISSN

Volume Title

Publisher

Auckland University of Technology

Abstract

Information Extraction is a popular topic in the Natural Language Processing area. This thesis focuses on author identi cation in free text. This study divided the author identi cation task into two subtask, quotation extraction and speaker attribution. The entire system contains two parts, a rule based model for quotation extraction and a machine learning model for speaker attribution. The resource domain used in this thesis is the literary narrative. There is also a generalisation test on the news domain. The results of the experiment show that the rule based model can achieve a 0.88 F-score on quotation extraction and the best result of a machine learning model is 85.7% accuracy. The overall test on the entire system returns 77.9% accuracy on the literary source domain and 73.6% on the news domain.

Description

Source

DOI

Publisher's version

Rights statement

Collections