TY - GEN
T1 - Continuous sign language recognition with iterative spatiotemporal fine-tuning
AU - Koishybay, Kenessary
AU - Mukushev, Medet
AU - Sandygulova, Anara
N1 - Funding Information:
This work was supported by the Nazarbayev University Faculty Development Competitive Research Grant Program 2019-2021 “Kazakh Sign Language Automatic Recognition System (K-SLARS)”. Award number is 110119FD4545.
Publisher Copyright:
© 2021 IEEE
PY - 2020
Y1 - 2020
N2 - This paper aims to develop a deep neural network for Continuous Sign Language Recognition (CSLR) with iterative Gloss Recognition (GR) fine-tuning. CSLR has been a popular research field in the last few years and iterative optimization methods are well established. This paper introduces our proposed architecture involving Spatiotemporal feature-extraction model to segment useful “gloss-unit” features and BiLSTM with CTC as a sequence model. Spatiotemporal Feature Extractor is used for both image features extraction and sequence length reduction. To this end, we compare different architectures for feature extraction and sequence model. In addition, we iteratively fine-tune feature extractor on gloss-unit video segments with alignments from the end2end model. During the iterative training, we use novel alignment correction technique, which is based on minimum transformations of Levenshtein distance. All the experiments are conducted on the RWTH-PHOENIX-Weather-2014 dataset.
AB - This paper aims to develop a deep neural network for Continuous Sign Language Recognition (CSLR) with iterative Gloss Recognition (GR) fine-tuning. CSLR has been a popular research field in the last few years and iterative optimization methods are well established. This paper introduces our proposed architecture involving Spatiotemporal feature-extraction model to segment useful “gloss-unit” features and BiLSTM with CTC as a sequence model. Spatiotemporal Feature Extractor is used for both image features extraction and sequence length reduction. To this end, we compare different architectures for feature extraction and sequence model. In addition, we iteratively fine-tune feature extractor on gloss-unit video segments with alignments from the end2end model. During the iterative training, we use novel alignment correction technique, which is based on minimum transformations of Levenshtein distance. All the experiments are conducted on the RWTH-PHOENIX-Weather-2014 dataset.
KW - Gesture recognition
KW - Language
KW - Sequence modeling
KW - Vision
UR - https://www.scopus.com/pages/publications/85110430862
UR - https://www.scopus.com/pages/publications/85110430862#tab=citedBy
U2 - 10.1109/ICPR48806.2021.9412364
DO - 10.1109/ICPR48806.2021.9412364
M3 - Conference contribution
AN - SCOPUS:85110430862
T3 - Proceedings - International Conference on Pattern Recognition
SP - 10211
EP - 10218
BT - Proceedings of ICPR 2020 - 25th International Conference on Pattern Recognition
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 25th International Conference on Pattern Recognition, ICPR 2020
Y2 - 10 January 2021 through 15 January 2021
ER -