12/10/2025

Qlean Dataset Releases Japanese Single-Speaker Social & Cultural Speech Corpus for ASR and Language Modeling

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) has released a new dataset within its AI training data solution “Qlean Dataset”, operated through its subsidiary Amana Images Inc. The newly launched Japanese Single-Speaker Social/Cultural Themed Speech Corpus Dataset contains monologue recordings in which speakers talk freely about familiar topics such as daily experiences, memories from family or school life, personal values, and reflections.
This dataset can be used for training and evaluating automatic speech recognition (ASR), natural language processing (NLP), and foundational generative AI models.

 
The dataset features unscripted monologue speech, with speakers talking freely based on their own recollections and experiences. The recordings naturally include diverse linguistic expressions such as retrospection, explanation, topic shifts, and episodic storytelling, making the dataset suitable for training and evaluating models that handle continuous monologue structures.

 
The storytelling style—where speakers describe their personal experiences, daily observations, and reflections—supports accuracy evaluation for speech recognition systems that must operate within real-life conversational or narrative contexts.
It is also well suited for natural language processing tasks such as summarization, topic extraction, and semantic understanding.
Additionally, the long-form monologue structure is valuable for improving generative AI and conversational AI models, and can be applied across domains including educational AI, speech analytics, and content creation workflows.

Overview of the “Japanese Single-Speaker Social/Cultural Themed Speech Corpus”

Overview

This is an audio dataset of monologues on a wide range of sociocultural topics, including society, culture, lifestyle, and education.

Data Type

Audio

Speaker Attributes

Male and female speakers, ages 20s–50s

Format

MP3 / wav

Recording Time

5–60 min per file

Audio Rate

44.1kHz

Scenes

- Scenes in which the speaker continuously explains or explains social or cultural themes
- Long monologues and spontaneous interlocutor-style speech
— Includes everyday topic development, argument organization, and anecdotes
- Unscripted monologues that reflect the speaker's natural rhythm and pauses
— Includes context-dependent narration, topic changes, and emotional inflections

Sample

https://qleandataset.visual-bank.co.jp/en/lineup/pn-011

Use Case Examples: Japanese Single-Speaker Social/Cultural Themed Speech Corpus Dataset

Academic Research

  • ASR evaluation involving monologue structures
    Because the speech includes natural topic shifts, retrospection, and emotional expressions, it is suitable for testing ASR robustness beyond what can be evaluated using conventional scripted speech.

  • Research on long-form semantic understanding and summarization
    Long monologues based on personal narratives are ideal for studies involving temporal reasoning, key-point extraction, and topic segmentation.

Industry Applications

  • Improving speech-driven generative AI systems
    Using natural monologue speech enhances the accuracy of long-form processing pipelines such as speech-to-text, text summarization, and explanation generation.

  • Voice-based lifelog and diary AI analysis
    The dataset supports evaluation of services that process personal reflections, emotional expressions, and episodic speech.

  • Enhancing context understanding in customer-support AI
    Natural monologue data includes redundancies and digressions that resemble real user behavior, making it valuable for evaluating contextual tracking.

Other Practical Uses

  • Analysis of explanatory speech in educational and learning-support AI
    Personal, experience-based narratives provide suitable material for long-form summarization, comprehension, and keyword extraction in educational AI systems.

About Qlean Dataset

Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports diverse data types including images, videos, audio, 3D, and text—enabling both research and commercial AI development in a legally safe environment.

Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continuously expands its specialized, industry-relevant lineup known as the “AI Data Recipe.”

By reducing the operational burden of data collection and preparation, Qlean Dataset helps build legally compliant and risk-free AI development environments.

▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup

Key Features of Qlean Dataset

  • Full consent obtained from all subjects; compliant with GDPR and CCPA

  • Existing datasets deliverable within one business day

  • Custom data collection and recording available

▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact

About Visual Bank Inc.

Visual Bank Inc. is a Tokyo-based startup building next-generation data infrastructure to maximize AI development capabilities under the mission, “Unlock the potential of all data.”
The company operates THE PEN, an AI-assisted creative tool for manga artists, and wholly owns Amana Images Inc., which provides the Qlean Dataset service.

CEO: Saneyuki Nagai
Address: C-Cube Minami Aoyama Building 6F, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en/
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.