Natural Language Processing for Under-resourced Languages: Developing a Welsh Natural Language Toolkit

Daniel Cunliffe, Andreas Vlachidis, Daniel Williams, Douglas Tudhope

Research output: Contribution to journalArticlepeer-review

57 Downloads (Pure)

Abstract

Language technology is becoming increasingly important across a variety of application domains which have become common place in large, well-resourced languages. However, there is a danger that small, under-resourced languages are being increasingly pushed to the technological margins. Under-resourced languages face significant challenges in delivering the underlying language resources necessary to support such applications.
This paper describes the development of a natural language processing toolkit for an under-resourced language, Cymraeg (Welsh). Rather than creating the Welsh Natural Language Toolkit (WNLT) from scratch, the approach involved adapting and enhancing the language processing functionality provided for other languages within an existing framework and making use of external language resources where available.
This paper begins by introducing the GATE NLP framework, which was used as the development platform for the WNLT. It then describes each of the core modules of the WNLT in turn, detailing the extensions and adaptations required for Welsh language processing. An evaluation of the WNLT is then reported. Following this, two demonstration applications are presented. The first is a simple text mining application that analyses wedding announcements. The second describes the development of a Twitter NLP application, which extends the core WNLT pipeline.
As a relatively small-scale project, the WNLT makes use of existing external language resources where possible, rather than creating new resources. This approach of adaptation and reuse can provide a practical and achievable route to developing language resources for under-resourced languages.
Original languageEnglish
Article number101311
Number of pages20
JournalComputer Speech & Language
Volume72
Early online date26 Oct 2021
DOIs
Publication statusPublished - Mar 2022

Keywords

  • natural language processing
  • under-resourced languages
  • Welsh
  • Cymraeg
  • language technology

Fingerprint

Dive into the research topics of 'Natural Language Processing for Under-resourced Languages: Developing a Welsh Natural Language Toolkit'. Together they form a unique fingerprint.

Cite this