Catégorie de document |
Contribution à un colloque ou à un congrès |
Titre |
Voice quality transformation using an extended source-filter speech model |
Auteur principal |
Stefan Huber |
Co-auteur |
Axel Roebel |
Colloque / congrès |
12th Sound and Music Computing Conference (SMC). Dublin : Juillet 2015 |
Comité de lecture |
Oui |
Collation |
p.69-76 |
Année |
2015 |
Statut éditorial |
Publié |
Résumé |
In this paper we present a flexible framework for parametric speech analysis and synthesis with high quality. It constitutes an extended source-filter model. The novelty of the proposed speech processing system lies in its extended means to use a Deterministic plus Stochastic Model (DSM) for the estimation of the unvoiced stochastic component from a speech recording. Further contributions are the efficient and robust means to extract the Vocal Tract Filter (VTF) and the modelling of energy variations. The system is evaluated in the context of two voice quality transformations on natural human speech. The voice quality of a speech phrase is altered by means of re-synthesizing the deterministic component with different pulse shapes of the glottal excitation source. A Gaussian Mixture Model (GMM) is used in one test to predict energies for the re-synthesis of the deterministic and the stochastic component. The subjective listening tests suggests that the speech processing system is able to successfully synthesize and arise to a listener the perceptual sensation of different voice quality characteristics. Additionally, improvements of the speech synthesis quality compared to a baseline method are demonstrated. |
Mots-clés |
Glottal source / voice quality / LF model / source-filter model / speech synthesis |
Equipe |
Analyse et synthèse sonores |
Cote |
Huber15a |
Adresse de la version en ligne |
http://architexte.ircam.fr/textes/Huber15a/index.pdf |
|