1/21/2026

Qlean Dataset Launches a Japanese Two-Speaker TV & Film Dialogue Audio Dataset with Transcripts

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai), through its subsidiary Amana Images Inc., has begun offering a new dataset under its AI training data solution, Qlean Dataset.
The newly released dataset, titled “Japanese Two-Speaker TV and Film-Themed Conversation Audio Corpus with Transcripts,” is designed for the development of speech- and language-based AI systems, including ASR (Automatic Speech Recognition), NLP (Natural Language Processing), and LLMs.

This dataset is a new addition to Qlean Dataset’s machine learning dataset lineup, AI Data Recipe.
It consists of Japanese audio recordings in which two native Japanese speakers—one male and one female—engage in conversational discussions centered on television programs, drama series, variety shows, and films. Each recording is paired with transcripts that accurately reflect the spoken content.

The conversations primarily focus on exchanging opinions based on shared content experiences, such as impressions of storylines, evaluations of characters, and reflections on specific scenes. As a result, the dataset captures natural dialogue that assumes common background knowledge of entertainment content, reflecting realistic conversational contexts.

All recordings are conducted without scripted control. Speakers freely share impressions and interpretations at a natural pace, resulting in conversations that include agreement and disagreement, follow-up explanations, and topic development.

The dataset captures authentic conversational structures, including backchannel responses, speaker turn-taking, and topic shifts. This makes it well suited for evaluating AI systems that must handle real-world conversational speech rather than isolated or monologic utterances.

Dataset Overview: Japanese Two-Speaker TV and Movie Theme Talk Speech Corpus and Transcrip

Data Types

Audio, Text

Speaker Attributes

Japanese speakers, male and female, aged 20s to 50s

Data Formats

Audio: mp3 / wav

Text: txt /json /csv

Total Duration

Approximately 220 hours (each recording ranges from approximately 5 to 60 minutes)

Audio Sampling Rate

44.1kHz / 48kHz

Recorded Scenarios

Two-speaker conversations exchanging opinions on TV programs, drama series, and filmsNaturally occurring, unscripted conversational speech

Sample Details

https://qleandataset.visual-bank.co.jp/en/lineup/pn-026



Use Case Examples for the Japanese Two-Speaker TV and Film-Themed Conversation Dataset

Research Use Cases

  • Evaluation of Dialogue-Based Speech Recognition Models
    The dataset can be used to compare recognition accuracy in Japanese ASR research using natural conversational speech that includes overlapping utterances and backchannel responses. It is particularly effective for analyzing recognition error patterns specific to dialogue, which are difficult to assess using monologue-only data.

  • Japanese Language Model Research Incorporating Dialogue Structure
    By leveraging dialogue text grounded in shared knowledge of TV and film content, researchers can analyze and evaluate language models with respect to topic transitions, response relationships, and conversational coherence.

Industrial Use Cases

  • Conversation Understanding Validation for Dialogue AI and Chatbots
    Natural dialogue data that includes entertainment-related topics can be used to validate the understanding and response generation performance of conversational AI systems designed for multi-user interaction scenarios.

  • Operational Testing of Voice-Input Applications
    By using speech data in which multiple speakers converse freely, developers can evaluate and improve ASR performance in real-world voice-enabled services and applications that assume natural conversational input.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Existing datasets deliverable within one business day

  • Custom data collection and recording services available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.

CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.