12/17/2025

Qlean Dataset Launches a Japanese Two-Speaker Science-Themed Conversational Speech Corpus

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has launched the “Japanese Two-Speaker Science-Themed Conversational Speech Corpus Dataset” as part of its AI training data solution, Qlean Dataset, operated through its subsidiary Amana Images Inc.

This dataset is part of the “AI Data Recipe” lineup offered by Qlean Dataset and is intended for research and development in speech-based AI, including automatic speech recognition (ASR), dialogue understanding, natural language processing (NLP), and generative AI models.

It contains Japanese conversational speech in which two speakers discuss scientific concepts and phenomena through explanations, questions, comparisons, and examples. The dialogues feature natural turn-taking and explanatory exchanges that extend beyond basic question-and-answer formats.

All recordings are unscripted and reflect real conversational flow, including overlapping speech, paraphrasing, and in-depth explanations. The dataset also includes long-form dialogues covering multiple scientific topics in sequence.

These characteristics make the dataset suitable for training and evaluating models under conditions close to real-world use, supporting applications such as scientific and technical conversational AI, explanation-focused AI systems, and speech-input-based generative AI.

The dataset is applicable to a broad range of AI development settings, from academic research to commercial implementation, where Japanese domain-specific conversational speech is required.

Overview of the “Japanese Two-Speaker Science-Themed Conversational Speech Corpus Dataset”

Data Type

Audio

Speaker Attributes

Japanese male and female speakers in their 20s to 50s

Data Format

MP3 / WAV

Total Duration

Approximately 400 hours
(Each recording ranges from approximately 5 to 60 minutes)

Sampling Rate

44.1 kHz

Target Scenarios

・Two-speaker conversations on scientific concepts and topics
・Dialogues involving explanations and questions on specialized subjects
・Unscripted, naturally progressing conversations
・Dialogues with examples, comparisons, and explanations
・Long-form conversations covering multiple scientific topics

Sample Details

https://qleandataset.visual-bank.co.jp/en/lineup/pn-019

Use Case Examples 

[Research Use]

  • Research on dialogue understanding models in scientific domains
    Using two-speaker conversational speech on scientific and technical topics, the dataset can be applied to training and evaluating dialogue understanding models that incorporate speaker turn-taking and explanatory structures.

  • Specialized speech and language processing research
    By leveraging Japanese conversational speech that includes technical terminology and conceptual explanations, the dataset can be used to evaluate ASR and NLP model performance in specialized domains.

[Industry Use]

  • Advancement of conversational AI and voice assistants
    The dataset can be used as training data for developing speech-based conversational AI systems designed for question-answering and explanatory dialogue in scientific and technical fields.

  • Development of speech-based interfaces for generative AI
    By utilizing conversational speech that includes domain-specific knowledge, the dataset contributes to improving dialogue accuracy in speech-input-based generative AI systems and knowledge delivery applications.

[Other Practical Applications]

  • Development of educational voice-based dialogue systems
    The dataset can be used to develop educational support systems and teaching materials that incorporate conversational speech featuring scientific explanations and question-and-answer interactions.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Existing datasets deliverable within one business day

  • Custom data collection and recording services available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.

CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.