8/27/2025

Qlean Dataset Launches 70,000+ Hours of Japanese Speech Data for ASR, LLM, and Speaker Diarization Development

Visual Bank, Inc. (Minato, Tokyo; CEO: Masayuki Nagai), through its subsidiary amanaimages Inc., has added a large-scale Japanese speech dataset exceeding 70,000 hours to its AI training data solution Qlean Dataset, available for immediate purchase.

■ What Is a Large-Scale Japanese Speech Dataset?

A Japanese speech corpus spanning single-speaker, two-speaker, and multi-speaker formats across multiple domains — education, healthcare, business, and cultural performance. Used as ML data for foundation model pretraining, ASR WER evaluation and domain adaptation, Speaker Diarization benchmarking, and conversational AI training. All data is rights-cleared for commercial use.

■ Dataset Overview

Over 70,000 hours of Japanese audio across three recording formats, covering a wide range of real-world speech styles from monologue readings to medical consultations, group discussions, and TV/film audio.

▷ Single-Speaker Audio 

Monologue & Reading: novels, personal storytelling Education & Lectures: university classes, instructional materials Cultural Performance: koudan Other: text-to-speech readings, entertainment talks (lifestyle, sports, business, music, beauty, etc.)

▷ Two-Speaker Audio

 Business dialogue (in-person or phone) Simulated calls (business and private) Child-to-child natural conversations Medical dialogues: doctor–nurse–patient scenarios

▷ Multi-Speaker Audio (3+) 

Group conversations: private and business settings Media audio: TV programs, film scenes, comedy
→ Inquiries: https://qleandataset.visual-bank.co.jp/en/contact

■ FAQ

Q: How can this dataset be used for ASR development? 
A: Benchmark WER/CER across monologue, lecture, and medical audio. Use for LoRA fine-tuning of Whisper or ESPnet with domain corpora, or mix with standard corpora to optimize generalization.

Q: How does this dataset support Speaker Diarization?
A: Overlapping two-speaker and multi-speaker interruption audio enables DER benchmarking. Media audio with BGM supports ASR robustness evaluation under noisy conditions.

Q: How can this dataset be used for LLM and conversational AI?
 A: Multi-domain dialogue audio serves as training data for context-dependent response generation and domain-specific chatbot development.

Q: Is this dataset applicable to TTS? 
A: Yes. Rakugo and koudan audio supports prosody control training for VITS or StyleTTS. Media audio enables robustness evaluation under noisy conditions.

Q: Is custom data collection available? 
A: Yes. Additional domains, speaker profiles, or scenarios available on request.

■ Use Cases

  • ASR Robustness Benchmarking & Domain Adaptation
    Benchmark WER/CER across monologue, lecture, and medical audio. Fine-tune Whisper/ESPnet via LoRA or experiment with standard/domain corpus mixing ratios.

  • Speaker Diarization Benchmarking

    Overlapping two-speaker and multi-speaker interruption audio enables DER evaluation; media audio with BGM supports noisy-condition ASR robustness testing.

  • Healthcare & Education Domain AI

    Medical dialogue for medical ASR benchmarking and EMR automation pretraining; lecture audio for academic ASR evaluation and dataset augmentation.

  • LLM & Conversational AI Training

    Multi-domain dialogue audio for context-dependent response generation and domain-specific chatbot training.

  • TTS & Voice Synthesis Development

    Koudan audio for VITS/StyleTTS prosody fine-tuning; TV/film audio for generative AI response naturalness evaluation.

Contact form: https://qleandataset.visual-bank.co.jp/en/contact
Official site: https://qleandataset.visual-bank.co.jp/en/

About Qlean Dataset

Qlean Dataset is a commercially licensed AI training data solution provided by amanaimages Inc., a wholly owned subsidiary of Visual Bank. All datasets are rights-cleared for commercial use, giving AI developers a legally secure environment to source and deploy high-quality training data. The platform covers audio, image, video, 3D, and text modalities, serving foundation model developers and applied AI teams alike. Through partnerships with domestic and international data holders, broadcasters, newspapers, and newswire agencies, Qlean Dataset continuously expands its AI Data Recipe lineup. Existing datasets ship within 2 business days; custom recording and data collection also available. 
URL: https://qleandataset.visual-bank.co.jp/en 

About Visual Bank Inc.

Visual Bank Group is a technology company developing data infrastructure and AI solutions that support advanced AI development. The company operates THE PEN, an AI tool for manga creators, and its subsidiary, amanaimages Inc., provides commercial digital content and AI training data solutions, including Qlean Dataset. Visual Bank is also a selected participant in GENIAC, a Japanese government initiative supporting the advancement of next generation AI technologies.

CEO: Saneyuki Nagai
Website:https://visual-bank.co.jp/en

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.