Перейти к основной навигации Перейти к поиску Перейти к основному содержанию

Comparison of Word Embeddings of Unaligned Audio and Text Data Using Persistent Homology

Результат исследований

Аннотация

We have performed preliminary work on topological analysis of audio and text data for unsupervised speech processing. The work is based on the assumption that phoneme frequencies and contextual relationships are similar in the acoustic and text domains for the same language. Accordingly, this allowed the creation of a mapping between these spaces that takes into account their geometric structure. As a first step, generative methods based on variational autoencoders were chosen to map audio and text data into two latent vector spaces. In the next stage, persistent homology methods are used to analyze the topological structure of two spaces. Although the results obtained support the idea of the similarity of the two spaces, further research is needed to correctly map acoustic and text spaces, as well as to evaluate the real effect of including topological information in the autoencoder training process.

Язык оригиналаEnglish
Название основной публикацииSpeech and Computer - 24th International Conference, SPECOM 2022, Proceedings
РедакторыS.R. Mahadeva Prasanna, Alexey Karpov, K. Samudravijaya, Shyam S. Agrawal
ИздательSpringer Science and Business Media Deutschland GmbH
Страницы700-711
Число страниц12
ISBN (печатное издание)9783031209796
DOI
СостояниеPublished - 2022
Событие24th International Conference on Speech and Computer, SPECOM 2022 - Gurugram
Продолжительность: нояб. 14 2022нояб. 16 2022

Серия публикаций

НазваниеLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Том13721 LNAI
ISSN (печатное издание)0302-9743
ISSN (электронное издание)1611-3349

Conference

Conference24th International Conference on Speech and Computer, SPECOM 2022
Страна/TерриторияIndia
ГородGurugram
Период11/14/2211/16/22

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Fingerprint

Подробные сведения о темах исследования «Comparison of Word Embeddings of Unaligned Audio and Text Data Using Persistent Homology». Вместе они формируют уникальный семантический отпечаток (fingerprint).

Цитировать