A boundary-based tokenization technique for extractive text summarization
Department of Computer Science, School of Information and Communication Technology, Federal University of Technology, P.M.B 1526, Owerri, Nigeria.
Research Article
World Journal of Advanced Research and Reviews, 2021, 11(02), 303–312
Article DOI: 10.30574/wjarr.2021.11.2.0351
Publication history:
Received on 25 June 2021; revised on 10 August 2021; accepted on 12 August 2021
Abstract:
The need to extract and manage vital information contained in copious volumes of text documents has given birth to several automatic text summarization (ATS) approaches. ATS has found application in academic research, medical health records analysis, content creation and search engine optimization, finance and media. This study presents a boundary-based tokenization method for extractive text summarization. The proposed method performs word tokenization by defining word boundaries in place of specific delimiters. An extractive summarization algorithm was further developed based on the proposed boundary-based tokenization method, as well as word length consideration to control redundancy in summary output. Experimental results showed that the proposed approach enhanced word tokenization by enhancing the selection of appropriate keywords from text document to be used for summarization.
Keywords:
Boundary-based; Tokenization; Extractive; Automatic; Text; Summarization
Full text article in PDF:
Copyright information:
Copyright © 2021 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0