AUTOMATED CONSTRUCTION OF A PUNCTUATION-TAGGED CORPUS FOR TRANSFORMER–BILSTM-BASED PUNCTUATION RESTORATION IN UZBEK

Mualliflar

  • Sharipov Maksud Siddiqovich

    Tashkent University of Information Technology image/svg+xml

  • Adinayev Hushnudbek Saylboyevich

    Urganch davlat universiteti image/svg+xml

Калит сўзлар: punctuation restoration, Uzbek language, Transformer, BiLSTM, natural language processing

Аннотация

This study proposes an automated algorithm for constructing a punctuation-annotated dataset for training Transformer–BiLSTM models in Uzbek. The process includes text normalization, punctuation-aware tokenization, token-level labeling, dataset partitioning, and conversion into CoNLL format. A total of 11 labels are used, and the resulting corpus contains 11,888,213 tokens. The proposed method is intended for preparing large-scale annotated Uzbek data for punctuation restoration.

Фойдаланилган адабиётлар

1. Tilk O., Alumäe T. LSTM for punctuation restoration in speech transcripts. Proceedings of Interspeech 2015.

2. Tilk O., Alumäe T. Bidirectional recurrent neural network with attention mechanism for punctuation restoration. // Proceedings of Interspeech 2016. – P. 3047-3051.

3. Courtland M., Faulkner, A., McElvain, G. Efficient automatic punctuation restoration using bidirectional Transformers with robust inference. 2020, Proceedings of the 17th International Conference on Spoken Language Translation. – P. 272-279.

4. Alam T., Khan A., Alam, F. Punctuation restoration using Transformer models for high – and low-resource languages. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020).

5. Nagy A., Bial B., Ács J. Automatic punctuation restoration with BERT models. arxiv preprint arxiv: 2101.07343.

6. Sharipov M.S., Adinaev H.S., Kuriyozov E.R. Rule-Based Punctuation Algorithm for the Uzbek Language. 2024 IEEE 25th International Conference of Young Professionals in Electron Devices and Materials (EDM), 2024. – P. 2410-2414.

7. Sharipov M.S., Adinaev H.S. Development of Models for Punctuation Analysis of Uzbek Language Texts. // Bulletin of TUIT: Management and Communication Technologies, 2025, 1(4).

8. Adinaev H.S. Punctuation Analysis of Uzbek Texts Based on the N-Gram Model. // Electronic Journal of Actual Problems of Modern Science, Education and Training, 2025, № 2. – Р. 104-10.

9. Sharipov M.S., Adinayev H.S., Yusupova M.M. Oʻzbek tili matnlarida punktuatsion tahlil qilish uchun korpus yaratish. // Development of Science, 2025, Vol. 2, №3, 128-133-b.

10. Sharipov M., Adinayev H. Oʻzbek tili matnlarida soʻroq gaplarni aniqlashning qoidaga asoslangan algoritmini ishlab chiqish. // Management and Future Technologies, 2025, Vol. 2, Issue 1, 182-89-b.

11. Sharipov M.S., Adinayev H.S. Shartli tasodifiy maydonlar modeli asosida oʻzbek tili matnlarini punktuatsion tahlil qilish. // Al-Fargʻoniy avlodlari, 2025, Vol. 1, Issue 2.

Sharipov M.S., Adinayev H.S. Parallel hisoblash va Apache Spark asosida katta hajmdagi matnlarni punktuatsion tahlil qilish. // Oʻzbekiston milliy axborot agentligi – Oʻza ilm-fan boʻlimi (elektron jurnal), Vol. 11(73), November 2025, 243-247-b.