Proceedings of Interspeech 2024, (Kos Island, Greece, September 1–5, 2024)
AlignNet: Learning Dataset Score Alignment Functions To Enable Better Training of Speech Quality Estimators
doi: 10.21437/Interspeech.2024-74
Cite This Publication
Jaden Pieper and Stephen D. Voran, “AlignNet: Learning Dataset Score Alignment Functions To Enable Better Training of Speech Quality Estimators,” in Proceedings of Interspeech 2024 (Kos Island, Greece, September 1–5, 2024). http://dx.doi.org/10.21437/Interspeech.2024-74
Jaden Pieper and Stephen D. Voran
Abstract:
We develop two complementary advances for training no-reference (NR) speech quality estimators with independent datasets. Multi-dataset finetuning (MDF) pretrains an NR estimator on a single dataset and then finetunes it on multiple datasets at once, including the dataset used for pretraining. AlignNet uses an AudioNet to generate intermediate score estimates before using the Aligner to map intermediate estimates to the appropriate score range. AlignNet is agnostic to the choice of AudioNet so any successful NR speech quality estimator can benefit from its Aligner. The methods can be used in tandem, and we use two studies to show that they improve on current solutions: one study uses nine smaller datasets and the other uses four larger datasets. AlignNet with MDF improves on other solutions because it efficiently and effectively removes misalignments that impair the learning process, and thus enables successful training with larger amounts of more diverse data.
Keywords: speech quality; subjective test; machine learning; corpus effect; listening experiment; no reference (NR) estimator
For technical information concerning this report, contact:
Jaden Pieper
Institute for Telecommunication Sciences
jpieper@ntia.gov
Disclaimer: Certain commercial equipment, components, and software may be identified in this report to specify adequately the technical aspects of the reported results. In no case does such identification imply recommendation or endorsement by the National Telecommunications and Information Administration, nor does it imply that the equipment or software identified is necessarily the best available for the particular application or uses.
For questions or information on this or any other NTIA scientific publication, contact the ITS Publications Office at ITSinfo@ntia.gov or 303-497-3572.