СБОР ТЕКСТОВЫХ ДАННЫХ С ВЕБ-САЙТОВ ДЛЯ ЯЗЫКОВ С ОГРАНИЧЕННЫМИ РЕСУРСАМИ
Авторы
-
Утеулиев Ниетбай Утеулиевич
Нукусский государственный технический университет
-
Абдалиева Гулистан Рафиковна
Нукусский государственный технический университет
-
Калмуратов Бекбос Кушкинбаевич
Нукусский государственный технический университет
-
Асенбаева Даллорис Адилбаевна
Ключевые слова: каракалпакский язык, веб-скрейпинг, электронный корпус, языки с ограниченными ресурсами, NLP
Аннотация
Библиографические ссылки
1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser Ł., Polosukhin I. Attention is All You Need. // Advances in Neural Information Processing Systems, 2017, 30. – P. 5998-6008.
2. Devlin J., Chang M.W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. // arXiv preprint arXiv:1810.04805. 2018.
3. Joshi P., Santy S., Budhiraja R., Bali K., Choudhury M. The State and Fate of Linguistic Diversity in the NLP World. // Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. – P. 6282-6293.
4. Linder L., Jungo M., Hennebert J., Musat C., Fischer A. Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020. – P. 2706-2711.
5. Laippala V., Rönnqvist S., Hellström S., Luotolahti J., Repo L., Salmela A., Skantsi V., Pyysalo S. From Web Crawl to Clean Register-Annotated Corpora. Proceedings of the 12th Web as Corpus Workshop, 2020. – P. 14-22.
6. Kahlon N.K., Singh W. Comparative Analysis of Web Scraping Tools for Low-Resource Language Text. // International Journal of Engineering Trends and Technology, 72(1), 2024. – P. 284-299.