11/12/2025
Qlean Dataset Launches Japanese Two-Speaker Daily Conversation Audio Corpus

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) has announced that its subsidiary, Amana Images Inc., has released a new dataset within its AI training data solution Qlean Dataset: the “Japanese Two-Speaker Daily Conversation Audio Corpus.”
This dataset contains recordings of real Japanese daily conversations between two speakers—such as family members, friends, and colleagues—in various natural settings. Each conversation is recorded in high-quality WAV format, using stereo L/R channels to separate the speakers, and includes natural pacing, backchanneling, and overlapping speech.
The dataset features diverse topics such as relationships, pets, local culture, and food. It is ideal for training models in Automatic Speech Recognition (ASR), Natural Language Processing (NLP), and Conversational AI applications.
Additionally, it can be effectively used as training and validation data for multimodal AI and voice-based large language models (LLMs) that incorporate speech input.
Overview of the “Japanese Two-Speaker Daily Conversation Audio Corpus”
Speaker Attributes:Men and women in their 20s–40s
Data Format:WAV
Total Duration:Several hundred hours
Audio Specifications:Stereo (L/R channel separation)
Scenes Covered:Natural daily conversations between family, friends, and colleagues (unscripted)
Main Topics:Love advice, pet care, food, and regional culture
Sample:https://qleandataset.visual-bank.co.jp/en/lineup/pn-022
Use Case Examples of the Dataset
1: Development of ASR and Conversational Understanding AI
Speech Recognition for Natural Conversations
By utilizing unscripted Japanese dialogue featuring backchanneling, overlaps, and intonation variations, this dataset helps improve ASR model accuracy under realistic conditions.
It is also suitable for evaluating conversational ASR systems used in smart devices and voice assistants.Context and Intent Understanding
Using spontaneous dialogue data enables testing models that interpret topic shifts and ellipsis.
It is valuable for dialogue summarization, utterance classification, and intent estimation in the NLP domain.
Use Case 2: Emotional and Communication AI Research
Emotion Recognition and Behavioral Analysis
The dataset preserves features such as speech rate, prosody, and pauses, enabling emotion estimation and psychological state classification.
It also supports communication analysis by studying laughter, timing of responses, and conversational rhythm.Conversational Skill Evaluation and Educational Applications
By analyzing speech flow and response structures, the dataset can be used in AI systems for dialogue ability assessment, language learning, and speaking education.
Use Case 3: Applied AI and Real-World Implementations
Conversation Summarization and Meeting Minutes Generation
Since it includes conversations from home, work, and daily life, it is suitable for testing dialogue summarization, structuring, and information extraction models.
It can also be applied in customer support automation and call log analysis.Human Interaction Research
By analyzing speech timing and response tendencies, researchers can develop interaction models that simulate human-like dialogue.
It is useful for social robotics, educational AI, and other human-AI interaction design studies.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports diverse data types including images, videos, audio, 3D, and text—enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continuously expands its specialized, industry-relevant lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps build legally compliant and risk-free AI development environments.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en/
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Full consent obtained from all subjects; compliant with GDPR and CCPA
Existing datasets deliverable within one business day
Custom data collection and recording available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building next-generation data infrastructure to maximize AI development capabilities under the mission, “Unlock the potential of all data.”
The company operates THE PEN, an AI-assisted creative tool for manga artists, and wholly owns Amana Images Inc., which provides the Qlean Dataset service.
CEO: Saneyuki Nagai
Address: C-Cube Minami Aoyama Building 6F, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en/
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





