1/15/2026

Qlean Dataset Launches a Japanese Two-Speaker Conversational Speech Corpus with Transcripts

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) has announced the release of a new dataset, “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts,” through its AI training data solution, Qlean Dataset, operated by its subsidiary, Amanai Images Inc.

This dataset is designed for the development of voice and language AI technologies, including ASR (Automatic Speech Recognition), NLP (Natural Language Processing), and large language models (LLMs).

This dataset is a new addition to AI Data Recipe, the machine learning dataset lineup provided by Qlean Dataset.
It consists of Japanese conversational audio recorded between two speakers, paired with corresponding transcripts.
The conversations focus on leisure, hobbies, and entertainment, covering everyday topics such as impressions of TV dramas and anime, reviews of games and gadgets, and personal experiences related to travel and outings.

 All recordings were conducted without scripts, allowing participants to exchange opinions and impressions in a natural conversational flow.
By capturing spontaneous dialogue rather than controlled utterances, the dataset is suitable for research and development scenarios that assume real-world conversational settings, including speech recognition and dialogue processing tasks.

 Qlean Dataset provides AI development data prepared with clearly defined usage conditions and rights clearance, supporting both research and commercial development.
This dataset is offered as part of that initiative, with the aim of enabling robust evaluation environments based on Japanese conversational data reflecting everyday communication.

Dataset Overview: “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts”  

Data Types

Audio, Text

Speakers

Male and female participants in their 20s to 50s

Data Format

Audio: mp3 / wav
Text: txt

Total Recording Time

Approximately 400 hours

Sampling Rate

44.1 kHz

Included Conversation Scenarios

The dataset includes conversations in which two speakers continuously explain, reflect on, and discuss leisure-related topics.
Examples include commentary and analysis of entertainment content such as TV dramas and anime, reviews of games and gadgets, and personal stories related to travel and recreational activities.
Free-flowing conversations that incorporate personal experiences and subjective impressions are also included.

Sample Details

https://qleandataset.visual-bank.co.jp/en/lineup/pn-018

Use Case Examples for the “Japanese Two-Speaker Leisure-Themed Conversational Speech Corpus with Transcripts”

Use Case Examples — Research 

  • Evaluation of Japanese Conversational ASR Models
    This dataset can be used to evaluate recognition accuracy in ASR models that process multi-speaker dialogue, including speaker switching and turn-taking behavior.

  • Context-Aware Language Model Research
    By leveraging conversational Japanese text that includes topic transitions and cross-references, researchers can analyze context understanding and response generation behavior in LLMs and dialogue models.

Use Case Examples — Industry 

  • Validation of Voice UI and Conversational AI Systems
    For the development of voice assistants and conversational interfaces, this dataset supports PoC-level validation of speech input processing and dialogue control using natural Japanese conversations.

  • Evaluation and Fine-Tuning of Japanese LLMs
    The dataset can be applied to evaluate and fine-tune Japanese LLMs for natural response generation and dialogue continuity using conversational text not limited to business-specific scenarios.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Existing datasets deliverable within one business day

  • Custom data collection and recording services available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.

CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.