Text-to-motivational-speech with adjustable motivational factor to control motivational prosody
Artistic research deconstructing the performative excess of motivational western subcultures
Motivational speech has emerged as a popular audiovisual phenomenon within Western subcultures, conveying strategies for individual success through expressive, high-energy delivery. This paper artistically explores methods for synthesizing its distinctive prosody while critically examining its sociocultural foundations. Drawing on recent advances in emotion-controllable text-to-speech (TTS) and speech emotion recognition (SER), we employ deep learning models to replicate and analyze motivational speech. Our architecture introduces a one-dimensional motivational factor as a representation of the false promise of social mobility through individual effort, while enabling intensity-based control of motivational prosody. Situated within discourses on self-optimization and meritocracy, Motivational Speech Synthesis contributes to emotional speech synthesis while prompting reflection on work ethic.
The following diagram illustrates our EmoKnob based architecture for synthesizing motivational speech. Motivational intensity is controlled via averaged speaker embeddings, which are derived by selecting and averaging speech samples corresponding to different levels of motivational intensity within our dataset. These embeddings are generated in discrete increments along our one-dimensional motivational factor, ranging from 0 (low intensity) to 1 (high intensity). During inference, the closest embedding is selected according to the desired motivational factor, allowing precise emotional adjustment of generated speech.
The visualization below presents a 3-dimensional representation of emotional speech samples drawn from motivational speeches, projected into Valence-Arousal-Dominance (VAD) emotion space. Here, the scales range from negative to positive emotions (valence), calm to stimulated emotions (arousal), and submissive to dominant emotions (dominance). Each point represents a segment of motivational speech collected from YouTube, embedded using the deep learning-based speech emotion recognition model wav2vec 2.0. To distill these complex emotional patterns into a single interpretable value, we apply the UMAP dimensionality reduction algorithm, resulting in our motivational factor—a continuous scale ranging from 0 (low motivational intensity) to 1 (high motivational intensity). Explore the interactive plot to observe how different datapoints align along this motivational intensity spectrum.
Below are synthesized audio samples demonstrating how our motivational factor influences speech prosody. For each of seven motivational speech prompts, we generated audio at varying motivational intensity levels, ranging from 0.00 (low motivational intensity) to 1.00 (high motivational intensity). Listen to these examples to perceive how adjusting the motivational factor affects the expressiveness, tone, and emotional delivery of synthesized motivational speech.