Audio
Japanese 2-Speaker Customer Service Phone Reservation Dialogue Speech Dataset (Professional Voice Actors, L/R Separated)

This dataset provides Japanese dialogue speech in which 20 professional Japanese voice actors play the roles of operators and customers at reservation desks and similar services. Natural dialogues simulating phone reservations across six scenes (hotel, restaurant, taxi, train, hospital, and tourist site) were recorded in a studio with each speaker in a separate booth, and separated into left and right channels. It also includes utterances containing numbers such as postal codes and phone numbers, as well as audio conditions with noise added to one side of the conversation. Suitable for training AI models for speech recognition, speaker separation, call center AI, and voice agents.
Number of Files
77
Data Overview
Subject: Phone reservation dialogue speech by 20 professional Japanese voice actors (ages 20s-70s, 10 male / 10 female, 10 pairs) playing operator and customer roles: 930 files, approx. 76.8 hours (approx. 0.7-7.1 min per file, 4.7 min on average)
Scenarios: 6 scenes (hotel, restaurant, taxi, train, hospital, tourist site) x 5 patterns each + personal information confirmation, 33 themes in total; includes utterances with numbers such as postal codes and phone numbers
Collection method: Studio recording (directed, each speaker recorded in a separate booth)
Attributes: 3 audio conditions (no noise: approx. 25.9 hours / noise added to operator side: approx. 25.4 hours / noise added to customer side: approx. 25.4 hours); noise was added by editing the clean recordings
Audio specs: wav (linear PCM), 16kHz, 16bit, 2ch L/R separated (ch1: customer / ch2: operator)
License
We warrant, under our provision agreement, that all data has been legally and properly obtained and that we either hold the copyright or have obtained licensing from the rights holder. The intended use and scope, including AI training, can be confirmed in advance via our license terms.
Notes
Metadata such as recording equipment is provided.
FAQ
Q. Can this dataset be used for commercial AI model training?
A. Permitted use and scope can be confirmed in advance based on our license terms. Please see https://qleandataset.visual-bank.co.jp/en/legal/terms-and-conditions for details.
A. It includes phone reservation dialogues across six scenes (hotel, restaurant, taxi, train, hospital, and tourist site), with five patterns each. It also includes personal information confirmation exchanges containing numbers such as postal codes and phone numbers.
Q. Can it be used to train or evaluate speech recognition in noisy conditions?
A. In addition to clean audio, it includes versions with noise added to the operator side and to the customer side (noise was added by editing the clean recordings).
Q. Are transcripts included?
A. Transcripts are not currently included. Please let us know your requirements, including whether you need transcription.
Q. Can I request a sample?
A. Please request a sample using the form on this page.
Preview
















