ViSQA: A benchmark dataset and baseline models for Vietnamese spoken question answering

L Le Trong Minh N Nguyen Duc Thinh N Nguyen Khanh Tho Loc L Le Van Quan N Ngo Duc Tam L Le Hoang Son

Abstract

Spoken Question Answering (SQA) extends machine reading comprehension to spoken content and requires models to handle both automatic speech recognition (ASR) errors and downstream language understanding. Although large-scale SQA benchmarks exist for high-resource languages, Vietnamese remains underexplored due to the lack of standardized datasets. This paper introduces ViSQA, the first benchmark for Vietnamese Spoken Question Answering. ViSQA extends the UIT-ViQuAD corpus using a reproducible text-to-speech and ASR pipeline, resulting in over 13,000 question–answer pairs aligned with spoken inputs. The dataset includes clean and noise-degraded audio variants to enable systematic evaluation under varying transcription quality. Experiments with five transformer-based models show that ASR errors substantially degrade performance (e.g., ViT5 EM: 62.04% → 36.30%), while training on spoken transcriptions improves robustness (ViT5 EM: 36.30% → 50.70%). ViSQA provides a rigorous benchmark for evaluating Vietnamese SQA systems and enables systematic analysis of the impact of ASR errors on downstream reasoning.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 1
Published January 12, 2026
Pages e0340771
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (6)

L

Le Trong Minh

N

Nguyen Duc Thinh

N

Nguyen Khanh Tho Loc

L

Le Van Quan

N

Ngo Duc Tam

L

Le Hoang Son