Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

9 Commits
 
 
 
 
 
 

Repository files navigation

Supplementary Data: A Systematic Review on Audio-Visual Speech-To-Speech Translation Models

DOI

This repository contains the supplementary material and data extraction records for our systematic literature review exploring the transition from cascaded systems to direct end-to-end Audio-Visual Speech-to-Speech Translation (AV-S2ST) models.

Authors

  • Alexandre de Godoy Pereira (IPT - USP)
  • Renato Cordeiro Ferreira (IME - USP)
  • Alfredo Goldman (IME - USP)

Repository Contents

  • parsifal_export.md: Raw data export from Parsifal, containing the search protocol, inclusion/exclusion criteria, and quality assessment of the reviewed literature.
  • README.md: This document, detailing the repository structure and providing the exact DOIs for the 9 selected studies.

Selected Articles

Below is the final list of the 9 articles included in the systematic review:

Primary AV-S2ST Studies

  • AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation (Choi et al., 2024).
  • Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation (Goncalves et al., 2025).
  • AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation (Huang et al., 2023).
  • MixSpeech: Cross-modality Self-learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition (Cheng et al., 2023).
  • TransFace: Unit-based Audio-Visual Speech Synthesizer for Talking Head Translation (Cheng et al., 2024).
  • XLAVS-R: Cross-lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (Han et al., 2024).

Foundational S2ST Studies

License

This dataset and accompanying documentation are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) License.

About

Supplementary data, Parsifal export, and references for the systematic literature review on end-to-end Audio-Visual Speech-to-Speech Translation (AV-S2ST) models.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors