AlignNet: Learning Dataset Score Alignment Functions To Enable Better Training of Speech Quality Estimators

Jaden  Pieper

September 2024 | Conference Paper

AlignNet: Learning Dataset Score Alignment Functions To Enable Better Training of Speech Quality Estimators

doi: 10.21437/Interspeech.2024-74

Cite This Publication

Jaden Pieper and Stephen D. Voran

Abstract:

We develop two complementary advances for training no reference (NR) speech quality estimators with independent datasets. Multi-dataset finetuning (MDF) pretrains an NR estimator on a single dataset and then finetunes it on multiple datasets at once, including the dataset used for pretraining. AlignNet uses an AudioNet to generate intermediate score estimates before using the Aligner to map intermediate estimates to the appropriate score range. AlignNet is agnostic to the choice of AudioNet so any successful NR speech quality estimator canbenefit from its Aligner. The methods can be used in tandem, and we use two studies to show that they improve on current solutions: one study uses nine smaller datasets and the other uses four larger datasets. AlignNet with MDF improves on other solutions because it efficiently and effectively removes misalignments that impair the learning process, and thus enables successful training with larger amounts of more diverse data.

Keywords: speech quality; subjective test; machine learning; corpus effect; listening experiment; no reference (NR) estimator

(PieperInterspeech2024-1.pdf)

For technical information concerning this report, contact:

Jaden Pieper
Institute for Telecommunication Sciences
(202) 236-7516
jpieper@ntia.gov

For funding information concerning this report, click this link.

Disclaimer:

Certain commercial equipment, components, and software may be identified in this report to specify adequately the technical aspects of the reported results. In no case does such identification imply recommendation or endorsement by the National Telecommunications and Information Administration, nor does it imply that the equipment or software identified is necessarily the best available for the particular application or uses.

For questions or information on this or any other NTIA scientific publication, contact the ITS Publications Office at ITSinfo@ntia.gov or 303-497-3572.

Back to Search Results

Search Research Publications

AlignNet: Learning Dataset Score Alignment Functions To Enable Better Training of Speech Quality Estimators

Cite This Publication

Funding Information

Performing Agency

Funding Agency