This repository contains the supplementary material and data extraction records for our systematic literature review exploring the transition from cascaded systems to direct end-to-end Audio-Visual Speech-to-Speech Translation (AV-S2ST) models.
- Alexandre de Godoy Pereira (IPT - USP)
- Renato Cordeiro Ferreira (IME - USP)
- Alfredo Goldman (IME - USP)
parsifal_export.md: Raw data export from Parsifal, containing the search protocol, inclusion/exclusion criteria, and quality assessment of the reviewed literature.README.md: This document, detailing the repository structure and providing the exact DOIs for the 9 selected studies.
Below is the final list of the 9 articles included in the systematic review:
- AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation (Choi et al., 2024).
- Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation (Goncalves et al., 2025).
- AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation (Huang et al., 2023).
- MixSpeech: Cross-modality Self-learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition (Cheng et al., 2023).
- DOI: 10.48550/arXiv.2303.05309 (arXiv DOI)
- TransFace: Unit-based Audio-Visual Speech Synthesizer for Talking Head Translation (Cheng et al., 2024).
- XLAVS-R: Cross-lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (Han et al., 2024).
- One-to-many Multilingual End-to-End Speech Translation (Di Gangi et al., 2019).
- Multilingual End-to-End Speech Translation (Inaguma et al., 2019).
- DOI: 10.48550/arXiv.1910.00254 (arXiv DOI)
- Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation (Dong et al., 2022).
- DOI: 10.48550/arXiv.2205.08993 (arXiv DOI)
This dataset and accompanying documentation are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) License.