SPEAKER DIARIZATION OF KARAKALPAK SPEECH USING A PRETRAINED SORTFORMER MODEL UNDER LOW-RESOURCE CONDITIONS
DOI:
https://doi.org/10.5281/zenodo.22768834Abstract
pretrained Sortformer model was adapted to Karakalpak using a 20-hour multi-speaker speech corpus. The
model was evaluated with DER and JER metrics to determine the effectiveness of language-specific adaptation for low-resource
speaker diarization
Keywords
speaker diarization, Karakalpak language, low-resource speech, Sortformer, multi-speaker speech, overlapping speech, transfer learning, DER, JER.References
Park, T. J., Kanda, N., Dimitriadis, D., Han, K. J., Watanabe, S., & Narayanan, S. (2022). A review of speaker diarization:
Recent advances with deep learning. Computer Speech & Language, 72, 101317. https://doi.org/10.1016/j.csl.2021.101317
Fujita, Y., Kanda, N., Horiguchi, S., Nagamatsu, K., & Watanabe, S. (2019). End-to-End Neural Speaker Diarization
with Permutation-Free Objectives. Proceedings of Interspeech 2019.
Landini, F., Profant, J., Diez, M., & Burget, L. (2022). Bayesian HMM clustering of x-vector sequences in speaker
diarization: Theory, implementation and analysis on standard tasks. Computer Speech & Language, 71, 101254. https://
doi.org/10.1016/j.csl.2021.101254
Horiguchi, S., Fujita, Y., Watanabe, S., Xue, Y., & Nagamatsu, K. (2020). End-to-End Speaker Diarization for an
Unknown Number of Speakers with Encoder-Decoder Based Attractors. Proceedings of Interspeech 2020, 269–273.
https://doi.org/10.21437/Interspeech.2020-1022
Horiguchi, S., et al. (2020). End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-
Decoder Based Attractors. Interspeech 2020.
Medennikov, I., Korenevsky, M., Prisyach, T., et al. (2020). Target-Speaker Voice Activity Detection: A Novel Approach for
Multi-Speaker Diarization in a Dinner Party Scenario. Proceedings of Interspeech 2020, 274–278. https://doi.org/10.21437/
Interspeech.2020-1602
Plaquet, A., & Bredin, H. (2023). Powerset multi-class cross entropy loss for neural speaker diarization. Proceedings of
Interspeech 2023.
Ryant, N., Singh, P., Krishnamohan, V., Varma, R., Church, K., Cieri, C., Du, J., & Ganapathy, S. (2021). The Third
DIHARD Diarization Challenge. Proceedings of Interspeech 2021, 3570–3574. https://doi.org/10.21437/Interspeech.2021-1208
Chung, J. S., Huh, J., Nagrani, A., Afouras, T., & Zisserman, A. (2020). Spot the Conversation: Speaker Diarisation in
the Wild. Proceedings of Interspeech 2020, 299–303. https://doi.org/10.21437/Interspeech.2020-2337
Uteuliev, N. U., Kudaybergenov, Zh. K., & Kudaybergenov, T. K. (2025). Development of an automatic speech
recognition system for the Karakalpak language using Wav2Vec. Science and Society, 1-1, 30–32.
Mamatov, N. S., Jalelov, K. M., Samijonov, B. N., Samijonov, A. N., & Madaminjonov, A. D. (2023). Qaraqalpaq tilindegi
tekstti sóylewge sintezlew sisteması ushın maǵlıwmatlar bazasın qáliplestiriw. Science and Society, 2-extra, 9–11.
Uteuliev, N., Khudaybergenov, K., Kudaybergenov, J., & Kudaybergenov, T. (2026). Karakalpak Speech Corpus: The
First Benchmark Dataset for Automatic Speech Recognition. Zenodo. https://doi.org/10.5281/zenodo.19079670
Park, T., Medennikov, I., Dhawan, K., Wang, W., Huang, H., Koluguri, N. R., Puvvada, K. C., Balam, J., & Ginsburg,
B. (2025). Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.
Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 48153–48169.
Medennikov, I., Park, T., Wang, W., Huang, H., Dhawan, K., Wang, J., Balam, J., & Ginsburg, B. (2025). Streaming
Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering. Proceedings of Interspeech
, 5238–5242. https://doi.org/10.21437/Interspeech.2025-2244
Dawalatabad, N., Ravanelli, M., Grondin, F., Thienpondt, J., Desplanques, B., & Na, H. (2021). ECAPATDNN
Embeddings for Speaker Diarization. Proceedings of Interspeech 2021, 3560–3564. https://doi.org/10.21437/
Interspeech.2021-941
Kadirbergenov, A., Tleumuratov, S., & Nasiratdinov, D. (2026). Karakalpak Speech Corpus: A 100-Hour Crowdsourced
Speech Dataset for Karakalpak ASR. Hugging Face.


