Audio
Japanese Regional Dialect Speech Dataset (Dialogue & Monologue)

This dataset provides audio recordings of regional Japanese dialects. It includes both two-speaker dialogue in a natural conversational format and single-speaker monologue recordings, covering dialects such as Hiroshima-ben, Osaka-ben, Kansai-ben, Okayama-ben, Iyo-ben, and Tosa-ben, with new dialects added on an ongoing basis. Suitable for training AI models for speech recognition, dialect classification, and speech synthesis.
Number of Files
20
Subject Attributes
Subject: Regional dialect speech by Japanese speakers (males and females in their 20s-60s), composed of two groups:
① Dialogue format (2 speakers): approx. 10 hours, no transcript
② Monologue format (1 speaker): approx. 10 hours, with transcript
Collection method: Studio recording (script-based with natural speech rhythm)
Attributes: Dialect type (Hiroshima-ben, Osaka-ben, Kansai-ben, Okayama-ben, Iyo-ben, Tosa-ben, etc., expanding) / audio rate (44.1kHz/48kHz, 16/24bit)
License
We warrant, under our provision agreement, that all data has been legally and properly obtained and that we either hold the copyright or have obtained licensing from the rights holder. The intended use and scope, including AI training, can be confirmed in advance via our license terms.
Notes
Monologue recordings include a transcript, while dialogue recordings are provided as audio only. Metadata such as audio bit rate can also be provided.
FAQ
Q. Can this dataset be used for commercial AI model training?
A. Permitted use and scope can be confirmed in advance based on our license terms. Please see https://qleandataset.visual-bank.co.jp/en/legal/terms-and-conditions for details.
A. The dataset includes Hiroshima-ben, Osaka-ben, Kansai-ben, Okayama-ben, Iyo-ben, Tosa-ben, and others, with new dialects added on an ongoing basis. Please contact us if you need a specific dialect.
Q. Are transcripts included?
A. Monologue (single-speaker) recordings include a transcript. Dialogue (two-speaker) recordings are provided as audio only.
Q. Can I get only the dialogue or only the monologue format?
A. Yes, each format is organized as a separate dataset. Let us know your requirements and we'll guide you to the closest match.
Q. Can I request a sample?
A. Please request a sample using the form on this page.















