12/12/2025

Qlean Dataset Releases Japanese 3-Speaker Multi-Party Speech Dataset for ASR and Speaker Diarization

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has released the Japanese 3-Speaker Comedy-Themed Dialogue Speech Corpus Dataset within its AI training data solution Qlean Dataset, provided through its subsidiary Amana Images Inc.


This dataset, newly added to Qlean Dataset’s AI Data Recipe lineup, contains natural three-speaker comedy-style dialogues.
It supports multi-speaker AI development, including ASR, conversational understanding, dialogue generation, and speaker tracking.

The recordings capture key multi-party interaction features—overlapping speech, interruptions, fast turn-taking, and topic transitions—making them effective training and evaluation data for separation, diarization, and dialogue models.

Recorded under natural multi-speaker conditions, the dataset enables validation and generalization testing for real-world scenarios.
It is suitable for applications such as interactive AI, meeting-minutes AI, voice agents, and robotics dialogue systems, and can be used in both research and educational environments.

Dataset Specifications  

Data Type

Audio

Speaker Attributes

Male and female speakers in their 20s to 50s

File Format

mp3 / wav

Total Duration

Approximately 100 hours (individual recordings: 20–30 minutes each)

Sampling Rate

44.1 kHz

Scenes

・Comedy-style casual talk, banter, and episodic exchanges among three speakers
・Fast-paced responses, improvised remarks, and natural timing
・Multi-speaker topic shifts with overlapping speech and interruptions
・Unscripted dialogue with spontaneous topics and emotional variation

Topic Examples

Romantic advice, childhood memories (first love, humorous mistakes), personal trends, hobbies, popular topics, favorite snacks, and approximately 200 topics in total.

Sample Details

https://qleandataset.visual-bank.co.jp/en/lineup/pn-035

Use Case Examples Research

— Research and Academic Applications

  •  Speaker Separation and Speaker Diarization Research
    Natural three-speaker interactions—including simultaneous speech, interruptions, and overlap—enable performance evaluation of diarization models and multi-speaker identification methods.

  • Natural Dialogue Understanding and Conversational Behavior Analysis
    Comedy-style pacing, improvisation, and topic shifts make the dataset valuable for studying turn-taking, discourse structure, and topic transition models.

  • Multimodal Dialogue Research Combining NLP and Speech Processing
    The multi-speaker audio characteristics can be used to train dialogue generation models, utterance prediction models, and response optimization models.

— Industrial Applications 

  • ASR Engine Development for Multi-Speaker Environments
    Three-speaker data with overlapping speech and interruptions supports ASR performance improvements for meeting AI, automated minutes-generation AI, and customer-service dialogue systems.

  • Conversational AI and Voice Assistant Development
    Fast-paced banter contributes to natural response generation, improved reaction modeling, and enhanced diversity in conversational AI outputs.

  • Evaluation of Multi-Speaker Audio Processing Technologies
    Useful for testing algorithms for speech separation, speaker tracking, volume estimation, and spatial inference under multi-speaker conditions.

— Educational Applications

The dataset can be used in academic settings as training material for speech engineering and dialogue AI, serving as practical multi-speaker audio data for exercises and coursework.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports diverse data types including images, videos, audio, 3D, and text—enabling both research and commercial AI development in a legally safe environment.

Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continuously expands its specialized, industry-relevant lineup known as the “AI Data Recipe.”

By reducing the operational burden of data collection and preparation, Qlean Dataset helps build legally compliant and risk-free AI development environments.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Full consent obtained from all subjects; compliant with GDPR and CCPA

  • Existing datasets deliverable within one business day

  • Custom data collection and recording available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building next-generation data infrastructure to maximize AI development capabilities under the mission, “Unlock the potential of all data.”
The company operates THE PEN, an AI-assisted creative tool for manga artists, and wholly owns Amana Images Inc., which provides the Qlean Dataset service.

CEO: Saneyuki Nagai
Address: C-Cube Minami Aoyama Building 6F, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en/
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.