Aolan Sun
YOU?
Author Swipe
View article: Graphpb: Graphical Representations Of Prosody Boundary In Speech Synthesis
Graphpb: Graphical Representations Of Prosody Boundary In Speech Synthesis Open
This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relationship of input sequences in a graphical domain for improv…
View article: FastGraphTTS: An Ultrafast Syntax-Aware Speech Synthesis Framework
FastGraphTTS: An Ultrafast Syntax-Aware Speech Synthesis Framework Open
This paper integrates graph-to-sequence into an end-to-end text-to-speech framework for syntax-aware modelling with syntactic information of input text. Specifically, the input text is parsed by a dependency parsing module to form a syntac…
View article: SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model
SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model Open
In recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there are always distortions existing in the predicted acoustic …
View article: Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking Avatar
Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking Avatar Open
Since the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to l…
View article: Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion
Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion Open
One-shot voice conversion (VC) with only a single target speaker's speech for reference has become a hot research topic. Existing works generally disentangle timbre, while information about pitch, rhythm and content is still mixed together…
View article: GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis
GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis Open
This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relationship of input sequences in a graphical domain for improv…
View article: GraphTTS: graph-to-sequence modelling in neural text-to-speech
GraphTTS: graph-to-sequence modelling in neural text-to-speech Open
This paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs consist of node and edge representations constructed from inp…