1/7/2026

Qlean Dataset Launches a Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has announced the release of the “Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts” as part of Qlean Dataset, its AI training data solution operated through its subsidiary, Amana Images Inc.


This dataset is newly released as part of Qlean Dataset’s machine learning dataset lineup, “AI Data Recipe.”
The dataset consists of Japanese monologue speech on leisure-related topics such as hobbies and entertainment, paired with verbatim transcripts.
By providing aligned audio and text data, it supports research and development in speech- and language-based AI, including ASR and NLP.

The recordings feature unscripted speech in which speakers naturally describe personal experiences, creative works, reviews, and reflections.
Continuous topic development and frequent evaluative expressions make the data closely representative of real-world user speech.

These characteristics make the dataset suitable for evaluating long-form speech recognition and context-aware language models.
It can also be used in the development of voice-driven AI services and language processing features such as review analysis and summarization, across both research and commercial development stages.

Dataset Overview: Japanese Single-Speaker Leisure Talk Speech Corpus with Transcripts

Data Type

Voice, text

Subject attributes

Japanese men and women in their 20s to 50s

Data Format

Audio data: wav / mp3

Text data: txt

Recording Time

Total: Approximately 600 hours (approximately 5-40 minutes per audio segment)

Audio Rate

44.1kHz

Target Scenes

・Japanese speech in which a single speaker continuously talks about personal experiences and thoughts on leisure-related topics

・Unscripted monologue-style natural speech
— Including topic development and shifts, episodic storytelling, and emotional expression

Sample

https://qleandataset.visual-bank.co.jp/en/lineup/pn-006

Use Case Examples for the Japanese Single-Speaker Leisure Talk Speech Corpus

Research Applications

  • Evaluating Long-Form Speech Recognition Accuracy
    The dataset can be used to evaluate ASR accuracy for extended monologue speech in which topics are discussed continuously, such as personal experiences at theme parks, explanations of walking on snowy roads, or impressions of dramas and video games.
    It enables analysis of recognition errors that arise as context accumulates over long utterances.

  • Discourse Structure and Pragmatics Research
    Based on speech in which a single speaker reflects on experiences or evaluates creative works, the dataset can be used for linguistic research analyzing transitions from topic introduction to development, evaluation, and supplementary explanation, as well as the manifestation of pragmatic functions such as opinions, comparisons, and cautions.

Industrial Applications

  • ASR Development for Voice-Input Applications
    By leveraging monologue speech about leisure experiences, travel tips, and reviews of dramas or games, the dataset can support the development of speech recognition features for applications with voice search, voice memo, and spoken review input functions.

  • Fine-Tuning Natural Language Processing Models
    The transcribed text data can be used to train NLP models for extracting key points from experiences, organizing evaluative aspects, and classifying content by topic or perspective in leisure-related narratives.

  • Validation of Audio–Text Integrated AI Systems
    Because the speech data is aligned with corresponding transcripts, the dataset can be used to validate and evaluate AI systems that interpret spoken input as text, enabling coordinated assessment of speech understanding and text processing performance.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Existing datasets deliverable within one business day

  • Custom data collection and recording services available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.

CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.