1/15/2026
Qlean Dataset Launches a Japanese Two-Speaker Conversational Speech Corpus with Transcripts

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) has announced the release of a new dataset, “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts,” through its AI training data solution, Qlean Dataset, operated by its subsidiary, Amanai Images Inc.
This dataset is designed for the development of voice and language AI technologies, including ASR (Automatic Speech Recognition), NLP (Natural Language Processing), and large language models (LLMs).
This dataset is a new addition to AI Data Recipe, the machine learning dataset lineup provided by Qlean Dataset.
It consists of Japanese conversational audio recorded between two speakers, paired with corresponding transcripts.
The conversations focus on leisure, hobbies, and entertainment, covering everyday topics such as impressions of TV dramas and anime, reviews of games and gadgets, and personal experiences related to travel and outings.
All recordings were conducted without scripts, allowing participants to exchange opinions and impressions in a natural conversational flow.
By capturing spontaneous dialogue rather than controlled utterances, the dataset is suitable for research and development scenarios that assume real-world conversational settings, including speech recognition and dialogue processing tasks.
Qlean Dataset provides AI development data prepared with clearly defined usage conditions and rights clearance, supporting both research and commercial development.
This dataset is offered as part of that initiative, with the aim of enabling robust evaluation environments based on Japanese conversational data reflecting everyday communication.
Dataset Overview: “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts”
Data Types | Audio, Text |
|---|---|
Speakers | Male and female participants in their 20s to 50s |
Data Format | Audio: mp3 / wav |
Total Recording Time | Approximately 400 hours |
Sampling Rate | 44.1 kHz |
Included Conversation Scenarios | The dataset includes conversations in which two speakers continuously explain, reflect on, and discuss leisure-related topics. |
Sample Details |
Use Case Examples for the “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts”
Use Case Examples — Research
Evaluation of Japanese Conversational ASR Models
This dataset can be used to evaluate recognition accuracy in ASR models that process multi-speaker dialogue, including speaker switching and turn-taking behavior.Context-Aware Language Model Research
By leveraging conversational Japanese text that includes topic transitions and cross-references, researchers can analyze context understanding and response generation behavior in LLMs and dialogue models.
Use Case Examples — Industry
Validation of Voice UI and Conversational AI Systems
For the development of voice assistants and conversational interfaces, this dataset supports PoC-level validation of speech input processing and dialogue control using natural Japanese conversations.Evaluation and Fine-Tuning of Japanese LLMs
The dataset can be applied to evaluate and fine-tune Japanese LLMs for natural response generation and dialogue continuity using conversational text not limited to business-specific scenarios.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Existing datasets deliverable within one business day
Custom data collection and recording services available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





