COMPARATIVE EVALUATION OF CNN-LSTM AND SQUEEZEFORMER-TRANSFORMER ARCHITECTURES FOR REAL-TIME UZBEK SIGN LANGUAGE RECOGNITION
DOI:
https://doi.org/10.5281/zenodo.22840569Abstract
Uzbek Sign Language (UzSL) remains underrepresented in machine-learning research, in part because annotated datasets and deployable recognition systems are scarce. This paper develops and evaluates a camera-based, isolated-sign recognition pipeline for UzSL using a proprietary landmark dataset, and compares a CNN-LSTM baseline against a Squeezeformer-style Transformer architecture. Over 500 gesture classes were recorded with HD webcams (1280×720, 30 fps) under varied lighting, background, clothing, and camera conditions, converted to normalized 130-point landmark sequences, and split 70/15/15 into training, validation, and test sets. The Transformer achieved 91% accuracy versus 87% for CNN-LSTM, with lower latency (0.58 s vs. 0.72 s), higher throughput (78 vs. 65 gestures/min), and stronger low-light accuracy (83% vs. 76%). A small controlled usability study (n = 10) returned a mean satisfaction score of 4.6/5. These results support attention-based sequence modeling as a promising direction for real-time UzSL recognition, though the system is currently limited to isolated-sign recognition and has not been evaluated with a signer-independent split. We present the dataset and prototype as a reproducible starting point for future continuous UzSL translation research.Keywords
Uzbek Sign Language, sign language recognition, CNN-LSTM, Transformer, Squeezeformer, computer vision, landmark sequences, assistive technologyReferences
J. Ergashev, M. Rakhmatova and M. Eshimov, “Sign Language Translation Model,” B.S. thesis, Dept. of Computer Science, Central Asian University, Tashkent, Uzbekistan, 2025.
W. Tak, V. Arora, V. Chambyal, V. Kumari, Sumitra, and S. Chug, “A Dual-Path Deep Learning Framework for Static and Dynamic Sign Language Recognition Using CNN and LSTM Architectures,” Procedia Computer Science, vol. 283, pp. 2557–2566, 2026, doi: 10.1016/j.procs.2026.06.325.
H. Luqman and E. Elalfy, “Utilizing motion and spatial features for sign language gesture recognition using cascaded CNN and LSTM models,” Turkish Journal of Electrical Engineering and Computer Sciences, vol. 30, no. 7, pp. 2508– 2525, 2022, doi: 10.55730/1300-0632.3952.
L. Pigou, S. Dieleman, P.-J. Kindermans, and B. Schrauwen, “Sign Language Recognition Using Convolutional Neural Networks,” in Computer Vision – ECCV 2014 Workshops, 2015, pp. 572–578, doi: 10.1007/978-3-319-16178-5_40.
Vaswani, N. Shazeer, N. Parmar, et al., “Attention Is All You Need,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
M. De Coster, M. Van Herreweghe, and J. Dambre, “Sign Language Recognition with Transformer Networks,” in Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC 2020), Marseille, France, 2020, pp. 6018–6024.
M. Mukushev, A. Ubingazhibov, A. Kydyrbekova, A. Imashev, V. Kimmelman, and A. Sandygulova, “FluentSigners-50: A signer independent benchmark dataset for sign language processing,” PLOS ONE, vol. 17, no. 9, e0273649, 2022, doi: 10.1371/journal.pone.0273649.
C. Lugaresi, J. Tang, H. Nash, et al., “MediaPipe: A Framework for Building Perception Pipelines,” arXiv preprint arXiv:1906.08172, 2019.
M. Sokolova and G. Lapalme, “A Systematic Analysis of Performance Measures for Classification Tasks,” Information Processing & Management, vol. 45, no. 4, pp. 427–437, 2009.
F. Pedregosa, G. Varoquaux, A. Gramfort, et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
World Wide Web Consortium (W3C), “Web Content Accessibility Guidelines (WCAG) 2.1,” 2018.
Google Cloud, “Text-to-Speech Documentation,” accessed May 2025.
J. Shen, R. Pang, R. J. Weiss, et al., “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” arXiv preprint arXiv:1712.05884, 2018.
J. Hamari, J. Koivisto, and H. Sarsa, “Does Gamification Work? A Literature Review of Empirical Studies,” in Proc. Hawaii Int. Conf. System Sciences (HICSS), pp. 3025–3034, 2014.
SignAll, “SignAll SDK for Sign Language Recognition,” 2023.
N. C. Camgoz, S. Hadfield, O. Koller, H. Ney, and R. Bowden, “Neural Sign Language Translation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7784–7793.
Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” in Proc. ICML, pp. 369–376, 2006.
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y. Sheikh, “OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 1, pp. 172–186, 2021.
T. Baltrušaitis, P. Robinson, and L.-P. Morency, “OpenFace: An Open Source Facial Behavior Analysis Toolkit,” in Proc. WACV, 2016, pp. 1–10.
TensorFlow, “Post-training quantization,” TensorFlow Lite Documentation. [Online]. Available: TensorFlow documentation.
B. Woll and G. Morgan, Language Learning and Language Teaching in Deaf Education, Cambridge University Press, 2012.

