COLLECTING TEXTUAL DATA FROM WEBSITES FOR LOW-RESOURCE LANGUAGES

Authors

  • Nietbay Uteulievich Uteuliev

    Nukus State Technical University

  • Gulistan Rafikovna Abdalieva

    Nukus State Technical University

  • Bekbos Kushkinbaevich Kalmuratov

    Nukus State Technical University

  • Dalloris Adilbayevna Asenbaeva

    Tashkent University of Information Technology image/svg+xml

Keywords: Karakalpak language, web scraping, electronic corpus, low-resource languages, NLP

Abstract

This article focuses on the automatic collection of textual data from websites for low-resource languages, using the Karakalpak language as a case study. The study developed a methodology for collecting article texts from kknews.uz, yuz.uz, and joqargikenes.uz using Python, Requests, and BeautifulSoup. As a result, 1 386 navigation pages were analyzed, and a primary raw text database containing 18 784 articles and 6 090 553 words was formed. The collected raw corpus may serve as an initial resource for further preprocessing and for future development of NLP systems, machine translation tools, and Transformer-based models for the Karakalpak language.

References

1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser Ł., Polosukhin I. Attention is All You Need. // Advances in Neural Information Processing Systems, 2017, 30. – P. 5998-6008.

2. Devlin J., Chang M.W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. // arXiv preprint arXiv:1810.04805. 2018.

3. Joshi P., Santy S., Budhiraja R., Bali K., Choudhury M. The State and Fate of Linguistic Diversity in the NLP World. // Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. – P. 6282-6293.

4. Linder L., Jungo M., Hennebert J., Musat C., Fischer A. Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020. – P. 2706-2711.

5. Laippala V., Rönnqvist S., Hellström S., Luotolahti J., Repo L., Salmela A., Skantsi V., Pyysalo S. From Web Crawl to Clean Register-Annotated Corpora. Proceedings of the 12th Web as Corpus Workshop, 2020. – P. 14-22.

6. Kahlon N.K., Singh W. Comparative Analysis of Web Scraping Tools for Low-Resource Language Text. // International Journal of Engineering Trends and Technology, 72(1), 2024. – P. 284-299.