12/17/2025
Qlean Dataset Launches a Japanese Two-Speaker Science-Themed Conversational Speech Corpus

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has launched the “Japanese Two-Speaker Science-Themed Conversational Speech Corpus Dataset” as part of its AI training data solution, Qlean Dataset, operated through its subsidiary Amana Images Inc.
This dataset is part of the “AI Data Recipe” lineup offered by Qlean Dataset and is intended for research and development in speech-based AI, including automatic speech recognition (ASR), dialogue understanding, natural language processing (NLP), and generative AI models.
It contains Japanese conversational speech in which two speakers discuss scientific concepts and phenomena through explanations, questions, comparisons, and examples. The dialogues feature natural turn-taking and explanatory exchanges that extend beyond basic question-and-answer formats.
All recordings are unscripted and reflect real conversational flow, including overlapping speech, paraphrasing, and in-depth explanations. The dataset also includes long-form dialogues covering multiple scientific topics in sequence.
These characteristics make the dataset suitable for training and evaluating models under conditions close to real-world use, supporting applications such as scientific and technical conversational AI, explanation-focused AI systems, and speech-input-based generative AI.
The dataset is applicable to a broad range of AI development settings, from academic research to commercial implementation, where Japanese domain-specific conversational speech is required.
Overview of the “Japanese Two-Speaker Science-Themed Conversational Speech Corpus Dataset”
Data Type | Audio |
Speaker Attributes | Japanese male and female speakers in their 20s to 50s |
Data Format | MP3 / WAV |
Total Duration | Approximately 400 hours |
Sampling Rate | 44.1 kHz |
Target Scenarios | ・Two-speaker conversations on scientific concepts and topics |
Sample Details |
Use Case Examples
[Research Use]
Research on dialogue understanding models in scientific domains
Using two-speaker conversational speech on scientific and technical topics, the dataset can be applied to training and evaluating dialogue understanding models that incorporate speaker turn-taking and explanatory structures.Specialized speech and language processing research
By leveraging Japanese conversational speech that includes technical terminology and conceptual explanations, the dataset can be used to evaluate ASR and NLP model performance in specialized domains.
[Industry Use]
Advancement of conversational AI and voice assistants
The dataset can be used as training data for developing speech-based conversational AI systems designed for question-answering and explanatory dialogue in scientific and technical fields.Development of speech-based interfaces for generative AI
By utilizing conversational speech that includes domain-specific knowledge, the dataset contributes to improving dialogue accuracy in speech-input-based generative AI systems and knowledge delivery applications.
[Other Practical Applications]
Development of educational voice-based dialogue systems
The dataset can be used to develop educational support systems and teaching materials that incorporate conversational speech featuring scientific explanations and question-and-answer interactions.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Existing datasets deliverable within one business day
Custom data collection and recording services available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





