Senior Research Scientist at Fish Audio working on generative audio and realtime voice intelligence, with 7+ years of experience from research to production. My work spans full-duplex spoken dialogue, realtime speech and voice-agent models, hyper-realistic and highly expressive TTS, music generation, and the full pipeline from data curation and codec/tokenization design to large-scale multi-GPU training and deployment. I focus on turning advanced speech and audio methods into product-ready AI systems for natural voice interaction and creative audio generation.
Download my resumé .
PhD in Informatics, 2024
National Institute of Informatics & SOKENDAI
MEng in Electrical Engineering and Information System, 2020
The University of Tokyo
BEng in Measurement and Control Technology and Instruments, 2016
Tianjin University
Python, C++, Shell, Git, MySQL
PyTorch, PyTorch Lightning, Hugging Face
SpeechBrain, WeNet, WeSpeaker, Kaldi, ESPnet
Expressive TTS, codec and tokenizer design, voice LLMs
Audio-language modeling, full-duplex speech systems
Chinese, English, Japanese
Research and develop next-generation speech, audio, and voice-agent models for natural realtime interaction and creative audio generation.
Conducted research on generative audio, voice LLMs, multimodal interaction, and intelligent audio understanding.
Led R&D of expressive TTS and full-stack voice-agent technologies for avatar and game products.
Developed voice generation systems for Li Auto smart-space products and contributed to the multimodal foundation model MindGPT-4o.
Focused on high-fidelity 48kHz singing voice generation in collaboration with research and engineering teams.
Developed speech AI systems for Taobao Live compliance and broadcaster-risk control.