1/16/2026
Qlean Dataset Launches a Japanese Two-Speaker Sports Dialogue Audio Dataset

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) has launched a new dataset through Qlean Dataset, its AI training data solution operated via its subsidiary Amana Images Inc. The newly released dataset, titled Japanese Two-Speaker Sports Dialogue Audio Corpus with Transcripts, is designed for the development and evaluation of speech- and language-based AI technologies, including ASR (Automatic Speech Recognition), NLP (Natural Language Processing), and LLM-driven applications.
This dataset is part of Qlean Dataset’s machine learning lineup, AI Data Recipe. It features Japanese audio recordings of two speakers engaging in natural, unscripted conversations focused on sports topics, with each recording paired with accurate transcripts.
The conversations include discussions of sports experiences, match reviews, and opinions on tactics and performance, reflecting how sports-related dialogue occurs in real-world settings. All recordings are conducted without scripts, capturing natural conversational patterns such as speaker turn-taking and overlapping speech.
These characteristics make the dataset suitable for research and development in speech recognition, dialogue processing, and spoken language understanding. As with all Qlean Dataset offerings, the data is provided for both research and commercial use, with rights clearance and usage conditions carefully organized.
Dataset Overview: “Japanese Two-Speaker Sports Dialogue Audio Corpus with Transcripts”
Subject Attributes | Japanese men and women in their 20s to 50s |
|---|---|
Data Format | Audio data: wav / mp3 |
Recording Time | Total: Approximately 200 hours (approximately 5-60 minutes per audio segment) |
Audio Rate | 44.1kHz |
Target Scenes | - Scenes in which two people share their sports experiences, competition analysis, and impressions of watching a game |
sample page |
Use Case Examples for the "Japanese Two-Speaker Sports Dialogue Audio Corpus with Transcripts"
Research Applications
Evaluation and Analysis of Conversational ASR Models
In Japanese ASR research, this dataset can be used to analyze recognition accuracy and error patterns under conditions that include speaker turn-taking and overlapping speech, using natural two-speaker dialogue audio.Dialogue Understanding and Discourse Structure Research
The dataset supports research on intent estimation, discourse structure analysis, and dialogue segmentation by providing continuous conversational exchanges involving explanations and opinions about sports.
Industrial Applications
Development of Voice-Based Conversational AI and Assistants
For voice interfaces designed to deliver sports information or interact with users, the dataset enables validation of recognition and response models using dialogue audio that closely reflects real conversational behavior.Validation of Conversation Log Analysis Technologies
By leveraging naturally progressing two-speaker conversations, the dataset can be used for preliminary validation of technologies such as speech separation and speaker turn detection in dialogue analysis systems.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Existing datasets deliverable within one business day
Custom data collection and recording services available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





