1/7/2026
Qlean Dataset Launches a Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has announced the release of the “Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts” as part of Qlean Dataset, its AI training data solution operated through its subsidiary, Amana Images Inc.
This dataset is newly released as part of Qlean Dataset’s machine learning dataset lineup, “AI Data Recipe.”
The dataset consists of Japanese monologue speech on leisure-related topics such as hobbies and entertainment, paired with verbatim transcripts.
By providing aligned audio and text data, it supports research and development in speech- and language-based AI, including ASR and NLP.
The recordings feature unscripted speech in which speakers naturally describe personal experiences, creative works, reviews, and reflections.
Continuous topic development and frequent evaluative expressions make the data closely representative of real-world user speech.
These characteristics make the dataset suitable for evaluating long-form speech recognition and context-aware language models.
It can also be used in the development of voice-driven AI services and language processing features such as review analysis and summarization, across both research and commercial development stages.
Dataset Overview: Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts
Data Type | Voice, text | |
Subject attributes | Japanese men and women in their 20s to 50s | |
Data Format | Audio data: wav / mp3 | Text data: txt |
Recording Time | Total: Approximately 600 hours (approximately 5-40 minutes per audio segment) | |
Audio Rate | 44.1kHz | |
Target Scenes | ・Japanese speech in which a single speaker continuously talks about personal experiences and thoughts on leisure-related topics | ・Unscripted monologue-style natural speech |
Sample |
Use Case Examples for the Japanese Single-Speaker Leisure Talk Speech Corpus
Research Applications
Evaluating Long-Form Speech Recognition Accuracy
The dataset can be used to evaluate ASR accuracy for extended monologue speech in which topics are discussed continuously, such as personal experiences at theme parks, explanations of walking on snowy roads, or impressions of dramas and video games.
It enables analysis of recognition errors that arise as context accumulates over long utterances.Discourse Structure and Pragmatics Research
Based on speech in which a single speaker reflects on experiences or evaluates creative works, the dataset can be used for linguistic research analyzing transitions from topic introduction to development, evaluation, and supplementary explanation, as well as the manifestation of pragmatic functions such as opinions, comparisons, and cautions.
Industrial Applications
ASR Development for Voice-Input Applications
By leveraging monologue speech about leisure experiences, travel tips, and reviews of dramas or games, the dataset can support the development of speech recognition features for applications with voice search, voice memo, and spoken review input functions.Fine-Tuning Natural Language Processing Models
The transcribed text data can be used to train NLP models for extracting key points from experiences, organizing evaluative aspects, and classifying content by topic or perspective in leisure-related narratives.Validation of Audio–Text Integrated AI Systems
Because the speech data is aligned with corresponding transcripts, the dataset can be used to validate and evaluate AI systems that interpret spoken input as text, enabling coordinated assessment of speech understanding and text processing performance.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Existing datasets deliverable within one business day
Custom data collection and recording services available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





