Audio
/
Text
Japanese Hobby & Leisure Talk Speech Dataset (Monologue & Dialogue)

This dataset provides Japanese speech recordings on hobby and leisure themes. It includes both single-speaker monologues and two-speaker dialogues, each paired with a transcript of the spoken content. It captures natural conversational flow, including comments and reviews of media, and travel or outing experiences. Suitable for training AI models for speech recognition, natural language processing, and dialogue AI.
Number of Files
1,151
Subject Attributes
Subject: Japanese talk speech on hobby and leisure themes by speakers in their 20s-50s, composed of two groups:
① Monologue format (1 speaker): approx. 1,112 hours
② Dialogue format (2 speakers): approx. 39 hours
Collection method: Audio recording (unscripted, natural speech)
Attributes: Transcript included (paired audio/text) / audio rate 44.1kHz
License
We warrant, under our provision agreement, that all data has been legally and properly obtained and that we either hold the copyright or have obtained licensing from the rights holder. The intended use and scope, including AI training, can be confirmed in advance via our license terms.
Notes
A transcript (written text of the spoken content) is provided paired with the audio data.
FAQ
Q. Can this dataset be used for commercial AI model training?
A. Permitted use and scope can be confirmed in advance based on our license terms. Please see https://qleandataset.visual-bank.co.jp/en/legal/terms-and-conditions for details.
A. Yes, each format is organized as a separate dataset. Let us know your requirements and we'll guide you to the closest match.
Q. Are transcripts included?
A. Yes, both formats include a transcript of the spoken content.
Q. What format is the data delivered in?
A. Delivered as paired audio data and transcript. Please contact us for specific file format details.
Q. Can I request a sample?
A. Please request a sample using the form on this page.















