4/16/2026
Qlean Dataset Launches Japanese Emotional Dialogue Speech Dataset for SER, LLM, and TTS Development

Visual Bank, Inc. (Minato, Tokyo; CEO: Masayuki Nagai), through its subsidiary amanaimages Inc., has released the Japanese Bi-Speaker Emotional Dialogue Speech Dataset under its AI training data solution Qlean Dataset.
■ What Is an Emotional Speech Dataset?
An emotional speech dataset is a speech corpus with emotion labels — such as joy, anger, and sadness — assigned to audio recordings. It serves as machine learning data for training and evaluating Speech Emotion Recognition (SER) models, improving emotional understanding in large language models (LLMs), and building expressive text-to-speech (TTS) systems. Two-speaker dialogue corpora are especially rare and valuable, as they capture acoustic features — backchanneling, emotional fluctuation, and cross-speaker intonation synchronization — that single-speaker datasets cannot provide.
■ Dataset Specifications
Studio-recorded natural dialogues between 15 pairs of Japanese speakers (ages 20s–70s), expressing four emotional states: excitement, anger, sorrow, and joy. Unlike scripted read-speech, the dyadic (two-speaker) interaction format captures how emotions propagate and synchronize between speakers in real conversation. Audio metadata including bitrate information is included.
Data Type: | Audio(two-speaker dialogue) |
|---|---|
Subject Profile: | 15 pairs of Japanese speakers, ages 20s–70s |
Data Volume: | 10 GB |
Total Files: | 63 files |
Format: | mp3 |
Emotions: | 4 categories (Excitement, Anger, Sorrow, Joy) |
Total Duration: | Approx. 15 hours (Approx. 20 minutes per file) |
Recording Environment: | Studio |
License: | Commercially licensed |
→ Sample data & full details:https://qleandataset.visual-bank.co.jp/en/lineup/ds-051
■ FAQ
Q: How can this dataset be used for SER development?
A: Benchmark WER on emotion-labeled dialogue against models like wav2vec2 or HuBERT. Use for LoRA or full fine-tuning for emotion-adaptive ASR under real conversational conditions.
Q: How does this dataset support LLM and multimodal AI?
A: Training/evaluation data for emotion understanding, style transfer (emotion→neutral), and multimodal models combining audio, transcripts, and emotion labels.
Q: Can this data be used for TTS fine-tuning?
A: Yes. Fine-tune VITS or StyleTTS on 4-emotion studio-quality prosody for expressive AI characters and virtual assistants.
Q: Is this dataset usable for Speaker Diarization?
A: Yes. Two-speaker alternating structure enables clean-condition DER benchmarking and analysis of emotional arousal effects on speaker identification.
Q: Is custom recording available?
A: Yes. Additional emotions, age groups, or scenarios available on request.
■ Use Case Scenarios
SER Model Training & Benchmarking
Train and evaluate emotion inference models using F0, MFCC, and spectral features. Enables quantitative analysis of cross-speaker emotional synchronization — not possible with single-speaker datasets.Emotion-Aware LLM & Multimodal AI Development
Combine audio, transcribed text, and emotion labels for context-dependent emotion understanding. Supports emotion-to-response style transfer and multimodal emotion recognition tasks.Expressive TTS & Conversational AI Fine-Tuning Fine-tune
VITS or StyleTTS on studio-quality 4-emotion prosody to build expressive voice synthesis engines for AI characters or virtual assistants.Contact Center Sentiment Analysis Engine
Ground-truth training data for SER models detecting customer dissatisfaction (anger) or satisfaction (joy). Combine with Google STT or Amazon Transcribe for real-time emotional alert systems.Speaker Diarization Accuracy Benchmarking
Two-speaker turn-taking structure enables clean-condition Diarization benchmarking and quantifies the impact of emotional states on speaker identification using WER and DER metrics.
About Qlean Dataset
Qlean Dataset is a commercially cleared AI training data solution provided by Amana Images, a subsidiary of Visual Bank Group. The platform offers diverse data formats including image, video, audio, 3D, and text, as well as a specialized AI Data Recipe lineup developed through collaborations with major media organizations and data rights holders.
URL:https://qleandataset.visual-bank.co.jp/en
URL:https://qleandataset.visual-bank.co.jp/en/products/japanese-language-corpora
Contact
About Visual Bank Inc.
Visual Bank Group is a technology company developing data infrastructure and AI solutions that support advanced AI development. The company operates THE PEN, an AI tool for manga creators, and its subsidiary, amanaimages Inc., provides commercial digital content and AI training data solutions, including Qlean Dataset. Visual Bank is also a selected participant in GENIAC, a Japanese government initiative supporting the advancement of next generation AI technologies.
CEO: Saneyuki Nagai
Website:https://visual-bank.co.jp/en





