AN INTELLIGENT INFORMATION SYSTEM FOR UZBEK TEXT PROCESSING AND GRAMMATICAL QUESTION-ANSWERING
Avtorlar
Gilt sózler: Uzbek NLP, ByT5, Viterbi, mDeBERTa-v3, CRF, hybrid RAG, BM25
Annotaciya
Paydalanılǵan ádebiyatlar
1. Sharipov M., Avezmatov I., Kuriyozov E., Babayev S. Natural Language Processing for Uzbek and Turkmen: A Systematic Review of Resources, Methods, and Benchmarks. // 2025 IEEE XVII International Scientific and Technical Conference on Actual Problems of Electronic Instrument Engineering (APEIE). IEEE, 2025. – P. 1-6.
2. Abjalova M. Tahrir va tahlil dasturlarining lingvistik modullari. – Tashkent: «Nodirabegim», 2020.
3. Matlatipov G., Vetulani Z. Representation of Uzbek Morphology in Prolog. // Aspects of Natural Language Processing. LNCS 5070. Springer, 2009. – P. 83-110.
4. Sharipov M., Salaev U., Matlatipov G. Morphological Analyzer of Uzbek Language Using Truncating Methods // Ilm Sarchashmalari. – Urgench, 2023. № 5/1. – P. 189-193.
5. Salaev U. UzMorphAnalyser: A Morphological Analysis Model for the Uzbek Language Using Inflectional Endings. // AIP Conference Proceedings. 2024. Vol. 3244(1), 030058. – P. 1-7.
6. Abjalova M. Oʻzbek tilidagi matnlarni avtomatik morfologik tahlil qilishda lemmatizatsiya va stemming jarayoni. // Oʻzbekiston: til va madaniyat. Lingvistika. 2024. № 3. – P. 6-21.
7. Sharipov M., Kuriyozov E., Vičič J. UzbekPOS: A Multi-Domain Dataset for Uzbek Part-of-Speech Tagging. // Data in Brief. 2026. Vol. 66, Article 112640.
8. Matlatipov S.G., Aripov M. UzUDT: Uzbek Universal Dependencies Treebank. // Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). 2026. – P. 11642-11649.
9. Xue L., Barua A., Constant N., et al. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models. // Transactions of the Association for Computational Linguistics. 2022. Vol. 10. – P. 291-306.
10. He P., Gao J., Chen W. DeBERTaV3: Improving DeBERTa Using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. // ICLR. 2023. – P. 1-17.
11. Forney G.D. The Viterbi Algorithm. // Proceedings of the IEEE. 1973. Vol. 61(3). – P. 268-278.
12. Lafferty J., McCallum A., Pereira F.C.N. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data // Proceedings of ICML. 2001. – P. 282-289.
13. Lewis P., Perez E., Piktus A., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. // Advances in Neural Information Processing Systems. 2020. Vol. 33. – P. 9459–9474.
14. Robertson S., Zaragoza H. The Probabilistic Relevance Framework: BM25 and Beyond. // Foundations and Trends in Information Retrieval. 2009. Vol. 3(4). – P. 333-389.
15. Reimers N., Gurevych I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. // Proceedings of EMNLP-IJCNLP. 2019. – P. 3982-3992.
16. Covington M.A., McFall J.D. Cutting the Gordian Knot: The Moving-Average Type–Token Ratio (MATTR). // Journal of Quantitative Linguistics. 2010. Vol. 17(2). – P. 94-100.