1/5/2026
Qlean Dataset Launches Japanese Two-Speaker Technology Dialogue Speech & Transcription Corpus

Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai; hereinafter “Visual Bank”) has launched a new dataset titled “Japanese Two-Speaker Technology-Themed Speech Transcripts” as part of Qlean Dataset, its AI training data solution operated through its subsidiary, Amana Images Inc.
This dataset is a new addition to AI Data Recipe, Qlean Dataset’s machine learning dataset lineup. It consists of Japanese two-speaker conversational speech centered on technology and IT topics, along with corresponding speech transcripts.
The conversations cover multiple contextual layers, including references to recent technological developments such as generative AI, related industry news, and practical usage ideas from everyday perspectives. The dialogues are unscripted and progress naturally through questions, explanations, opinion exchanges, comparisons, and real-world examples, closely reflecting authentic technical discussions.
The dataset is suitable for research and development across AI models that handle both speech and text, including automatic speech recognition (ASR), natural language processing (NLP), and conversational speech AI systems.
Overview of the “Japanese Two-Speaker Technology-Themed Speech Transcripts”
Data Types | Audio, Text |
Speaker Attributes | Japanese speakers, male and female, aged 20s to 50s |
Data Formats | Audio: wav / mp3 |
Total Recording Duration | Approximately 200 hours in total (each recording ranges from approximately 5 to 60 minutes) |
Sampling Rate | 44.1 kHz |
Covered Dialogue Scenes | ・This dataset features two-speaker conversations focused on technology and IT topics. |
Sample Details |
Use Cases for the Japanese Two-Speaker Technology Dialogue Dataset
Research Applications
Understanding Speaker Roles in Technical Dialogue
Two-speaker conversations on generative AI and IT topics support analysis of how questions, explanations, agreement, and disagreement unfold in natural technical discussions.Evaluating ASR Performance on Technical Speech
Dialogue speech and transcripts containing technical terms enable evaluation of ASR accuracy and error patterns beyond everyday conversation.Dialogue Understanding in Technology News Contexts
Conversations referencing recent technologies and news can be used to test topic tracking, contextual understanding, and key information extraction models.
Industrial Applications
Training Conversational AI for Technical Domains
Dialogue data discussing generative AI and IT services can be used to train speech-based conversational AI and chatbots that require technical context understanding.Speech-to-Text and Summarization for Technical Content
Long-form technology conversations support the development of transcription, summarization, and highlight extraction models for technical audio content.Validating Models for Technical Support and Knowledge Sharing
Practical conversations about IT tools and workflows are suitable for evaluating speech recognition and dialogue understanding models for internal use.
Other Practical Applications
Dialogue-Based Learning and Education
Clear, accessible explanations of technical topics make the dataset useful for developing dialogue-based learning and explanation support models in AI and IT education.
About Qlean Dataset
Qlean Dataset is a commercial-use-ready AI training data solution provided by Amana Images Inc., a subsidiary of Visual Bank Inc.
It supports a wide range of data types, including images, videos, audio, 3D assets, and text, enabling both research and commercial AI development in a legally safe environment.
Through collaborations with data partners such as Chiba Lotte Marines Co., Ltd. and Toyo Keizai Inc., Qlean Dataset continues to expand its specialized, industry-focused lineup known as the “AI Data Recipe.”
By reducing the operational burden of data collection and preparation, Qlean Dataset helps organizations establish AI development environments that are both legally compliant and risk-free.
▶ Qlean Dataset: https://qleandataset.visual-bank.co.jp/en
▶ AI Data Recipe: https://qleandataset.visual-bank.co.jp/en/lineup




Key Features of Qlean Dataset
Existing datasets deliverable within one business day
Custom data collection and recording services available
▶ Contact: https://qleandataset.visual-bank.co.jp/en/contact
About Visual Bank Inc.
Visual Bank Inc. is a Tokyo-based startup building Next-Generation Data infrastructure to enhance AI development capabilities under the mission “Unlocking Data Accessibility.”
The company operates THE PEN, an AI-assisted creative tool for manga artists and the Qlean Dataset service.
Its subsidiaries include Amana Images Inc., one of Japan’s largest photostock providers; Qlean Dataset, which leads research and development in AI data; and THE PEN Inc., an AI-assisted creative tool for manga artists.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Corporate Site: https://visual-bank.co.jp/en
Amana Images: https://qleandataset.visual-bank.co.jp/en/company-overview





