12/12/2025
Qlean Dataset Releases Japanese 3-Speaker Multi-Party Speech Dataset for ASR and Speaker Diarization

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has released the Japanese 3-Speaker Comedy-Themed Dialogue Speech Corpus Dataset within its AI training data solution Qlean Dataset, provided through its subsidiary Amana Images Inc.
This dataset, newly added to Qlean Dataset’s AI Data Recipe lineup, contains natural three-speaker comedy-style dialogues.
It supports multi-speaker AI development, including ASR, conversational understanding, dialogue generation, and speaker tracking.
The recordings capture key multi-party interaction features—overlapping speech, interruptions, fast turn-taking, and topic transitions—making them effective training and evaluation data for separation, diarization, and dialogue models.
Recorded under natural multi-speaker conditions, the dataset enables validation and generalization testing for real-world scenarios.
It is suitable for applications such as interactive AI, meeting-minutes AI, voice agents, and robotics dialogue systems, and can be used in both research and educational environments.
Dataset Specifications
Data Type | Audio |
Speaker Attributes | Male and female speakers in their 20s to 50s |
File Format | mp3 / wav |
Total Duration | Approximately 100 hours (individual recordings: 20–30 minutes each) |
Sampling Rate | 44.1 kHz |
Scenes | ・Comedy-style casual talk, banter, and episodic exchanges among three speakers |
Topic Examples | Romantic advice, childhood memories (first love, humorous mistakes), personal trends, hobbies, popular topics, favorite snacks, and approximately 200 topics in total. |
Sample Details |
Use Case Examples Research
— Research and Academic Applications
Speaker Separation and Speaker Diarization Research
Natural three-speaker interactions—including simultaneous speech, interruptions, and overlap—enable performance evaluation of diarization models and multi-speaker identification methods.Natural Dialogue Understanding and Conversational Behavior Analysis
Comedy-style pacing, improvisation, and topic shifts make the dataset valuable for studying turn-taking, discourse structure, and topic transition models.Multimodal Dialogue Research Combining NLP and Speech Processing
The multi-speaker audio characteristics can be used to train dialogue generation models, utterance prediction models, and response optimization models.
— Industrial Applications
ASR Engine Development for Multi-Speaker Environments
Three-speaker data with overlapping speech and interruptions supports ASR performance improvements for meeting AI, automated minutes-generation AI, and customer-service dialogue systems.Conversational AI and Voice Assistant Development
Fast-paced banter contributes to natural response generation, improved reaction modeling, and enhanced diversity in conversational AI outputs.Evaluation of Multi-Speaker Audio Processing Technologies
Useful for testing algorithms for speech separation, speaker tracking, volume estimation, and spatial inference under multi-speaker conditions.
— Educational Applications
The dataset can be used in academic settings as training material for speech engineering and dialogue AI, serving as practical multi-speaker audio data for exercises and coursework.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports diverse data types including images, videos, audio, 3D, and text—enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continuously expands its specialized, industry-relevant lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps build legally compliant and risk-free AI development environments.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Full consent obtained from all subjects; compliant with GDPR and CCPA
Existing datasets deliverable within one business day
Custom data collection and recording available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building next-generation data infrastructure to maximize AI development capabilities under the mission, “Unlock the potential of all data.”
The company operates THE PEN, an AI-assisted creative tool for manga artists, and wholly owns Amana Images Inc., which provides the Qlean Dataset service.
CEO: Saneyuki Nagai
Address: C-Cube Minami Aoyama Building 6F, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en/
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





