7/14/2025
Qlean Dataset Launches Japanese Emotional Speech Dataset: 9 Emotions × 2 Intensity Levels for SER, TTS, and Multimodal AI

Visual Bank, Inc. (Minato, Tokyo; CEO: Masayuki Nagai), through its subsidiary amanaimages Inc., has released the Emotionally Expressive Japanese Speech Dataset under its AI training data solution Qlean Dataset.
■ What Is a Japanese Emotional Speech Dataset?
A speech corpus with fine-grained emotion labels — 9 emotion categories × 2 intensity levels — recorded by native Japanese speakers. Used as ML data for Speech Emotion Recognition (SER) model training and evaluation, expressive TTS development, and multimodal emotion recognition tasks. The fixed-utterance design enables controlled cross-speaker experiments and intensity-level benchmarking unavailable in free-speech corpora.
■ Dataset Specifications
100 native Japanese speakers (ages 10s–50s, male and female) each recorded two fixed utterances across 9 emotion categories at two intensity levels (normal and strong), yielding 6,800 wav files with fine-grained emotion and intensity labels.
Data Type: | Audio (single-speaker, emotion-labeled utterances) |
|---|---|
Speakers: | 100 native Japanese speakers (ages 10s–50s, gender-balanced) |
Emotions: | 9 categories (Neutral, Calm, Joy, Sadness, Anger, Fear, Disgust, Surprise, Impatience) |
Intensity: | 2 levels (Normal / Strong) |
Utterances: | "I just heard this news on the radio." |
Files: | 6,800 |
Format: | wav |
License: | Commercially licensed |
Sample data & full details:: https://qleandataset.visual-bank.co.jp/en/lineup/ds-001
■ FAQ
Q: How can this dataset be used for SER development?
A: Train and evaluate 9-class × 2-intensity emotion classifiers using F0, MFCC, and spectral features. Fixed utterances enable controlled cross-speaker experiments and intensity-level accuracy analysis.
Q: Can this data be used for TTS fine-tuning?
A: Yes. Fine-tune VITS or StyleTTS on 9-emotion × 2-intensity prosody for intensity-aware expressive voice synthesis in AI characters and virtual assistants.
Q: How does this dataset support multimodal AI?
A: Combine audio, transcribed text, and speaker attributes (age, gender) to train multimodal emotion recognition models for VTubers or conversational agents.
Q: Is this dataset applicable to contact center sentiment analysis?
A: Yes. Use anger/impatience-labeled audio as ground truth for real-time SER models integrated with Google STT or Amazon Transcribe for escalation detection.
Q: Is custom recording available?
A: Yes. Additional emotions, intensity levels, age groups, or recording conditions available on request.
■ Use Cases
SER Model Training & Intensity-Level Benchmarking
Train and evaluate 9-class emotion classifiers using F0, MFCC, and spectral features. Fixed utterances enable controlled cross-speaker comparison and intensity-stratified accuracy evaluation.Expressive TTS & Conversational AI Fine-Tuning
Fine-tune VITS/StyleTTS on 9-emotion × 2-intensity prosody to build voice synthesis engines capable of intensity-aware emotional expression for AI characters or virtual assistants.Multimodal Emotion Recognition Development
Combine audio, utterance text, and speaker attributes to train multimodal models for VTuber/avatar emotion inference or affective computing research.Contact Center Real-Time Sentiment Analysis
Ground-truth SER training data for anger/impatience detection. Integrate with Google STT or Amazon Transcribe custom vocabulary for real-time escalation alert systems.
Model Evaluation & Benchmarking
9-emotion × 2-intensity labeled audio enables multi-class emotion classifier evaluation, intensity-level A/B testing, and WER measurement across emotional conditions.
About Qlean Dataset
Qlean Dataset is a commercially licensed AI training data solution provided by amanaimages Inc., a wholly owned subsidiary of Visual Bank. All datasets are rights-cleared for commercial use, giving AI developers a legally secure environment to source and deploy high-quality training data. The platform covers audio, image, video, 3D, and text modalities, serving foundation model developers and applied AI teams alike. Through partnerships with domestic and international data holders, broadcasters, newspapers, and newswire agencies, Qlean Dataset continuously expands its AI Data Recipe lineup. Existing datasets ship within 2 business days; custom recording and data collection also available.
URL: https://qleandataset.visual-bank.co.jp/en
URL: https://qleandataset.visual-bank.co.jp/en/lineup

About Visual Bank Inc.
Visual Bank Group is a technology company developing data infrastructure and AI solutions that support advanced AI development. The company operates THE PEN, an AI tool for manga creators, and its subsidiary, amanaimages Inc., provides commercial digital content and AI training data solutions, including Qlean Dataset. Visual Bank is also a selected participant in GENIAC, a Japanese government initiative supporting the advancement of next generation AI technologies.
CEO: Saneyuki Nagai
Website:https://visual-bank.co.jp/en





