Bilkent University
Department of Computer Engineering
M.S.THESIS PRESENTATION

 

Timestep-Dependent Conditioning for Emotion Transfer in Flow-Matching Text-to-Speech

 

Ali Azak
Master Student
(Supervisor: Assoc.Prof.Hamdi Dibeklioğlu )

Computer Engineering Department
Bilkent University

Abstract: Emotion transfer in text-to-speech aims to generate speech that expresses a target emotion while preserving the linguistic content and speaker identity of a source recording. Existing systems typically represent emotion as a single static condition obtained from an emotion encoder, an external emotion classifier, or a text prompt, and apply this condition uniformly throughout generation. When it is combined with reference speech in a different emotion, the reference tends to dominate, and the target emotion is only partially expressed. This thesis addresses both limitations by fine-tuning a pretrained F5-TTS flow-matching model. Given a source utterance and its transcription, the model reconstructs a parallel recording of the same sentence in the target emotion. Emotion transfer is therefore learned within the generator. A learned emotion-timestep interaction table injects emotion through the timestep-conditioning path and stores one offset vector for each emotion and timestep bin, allowing control to vary along the generation trajectory. Frozen WavLM speaker embeddings enter through the same path, while factorized classifier-free guidance assigns separate inference-time coefficients to the text-and-emotion, speaker, and reference-audio guidance directions. Across all directed emotion pairs for a held-out speaker from the English subset of the Emotional Speech Database, the method achieves target-emotion expression comparable to that of a specialized system trained from scratch, with better intelligibility and higher predicted quality at a small cost in speaker similarity.

 

DATE: September 10, Thursday @ 10:00

Place: EA 516