Перейти к основной навигации Перейти к поиску Перейти к основному содержанию

Continuous sign language recognition with iterative spatiotemporal fine-tuning

  • Nazarbayev University

Результат исследований

Аннотация

This paper aims to develop a deep neural network for Continuous Sign Language Recognition (CSLR) with iterative Gloss Recognition (GR) fine-tuning. CSLR has been a popular research field in the last few years and iterative optimization methods are well established. This paper introduces our proposed architecture involving Spatiotemporal feature-extraction model to segment useful “gloss-unit” features and BiLSTM with CTC as a sequence model. Spatiotemporal Feature Extractor is used for both image features extraction and sequence length reduction. To this end, we compare different architectures for feature extraction and sequence model. In addition, we iteratively fine-tune feature extractor on gloss-unit video segments with alignments from the end2end model. During the iterative training, we use novel alignment correction technique, which is based on minimum transformations of Levenshtein distance. All the experiments are conducted on the RWTH-PHOENIX-Weather-2014 dataset.

Язык оригиналаEnglish
Название основной публикацииProceedings of ICPR 2020 - 25th International Conference on Pattern Recognition
ИздательInstitute of Electrical and Electronics Engineers Inc.
Страницы10211-10218
Число страниц8
ISBN (электронное издание)9781728188089
DOI
СостояниеPublished - 2020
Событие25th International Conference on Pattern Recognition, ICPR 2020 - Virtual, Milan
Продолжительность: янв. 10 2021янв. 15 2021

Серия публикаций

НазваниеProceedings - International Conference on Pattern Recognition
ISSN (печатное издание)1051-4651

Conference

Conference25th International Conference on Pattern Recognition, ICPR 2020
Страна/TерриторияItaly
ГородVirtual, Milan
Период1/10/211/15/21

ASJC Scopus subject areas

  • Computer Vision and Pattern Recognition

Fingerprint

Подробные сведения о темах исследования «Continuous sign language recognition with iterative spatiotemporal fine-tuning». Вместе они формируют уникальный семантический отпечаток (fingerprint).

Цитировать