3/12/2026
Qlean Dataset Launches Compliant Safety-Aligned Datasets for Japanese Language Corpora and Multimodal AI

TOKYO, March 12, 2026 — Visual Bank Inc. (Minato-ku, Tokyo; CEO: Saneyuki Nagai) today announced that its AI training data solution, Qlean Dataset, operated through its subsidiary amanaimages Inc., has launched a new service suite offering specialized safety-aligned training datasets for multimodal AI systems, including Vision-Language Models (VLMs) that process digital assets such as images and videos.
As an initial initiative, Qlean Dataset positions itself as a leading provider of Japanese language data infrastructure for high-performance AI.
The new service covers the design, collection, creation, and provision of safety-aligned datasets that support responsible model training across Large Language Models (LLMs), image generation models, and multimodal AI systems. The initiative addresses the growing need to mitigate inappropriate outputs and manage compliance risks in generative AI development.
Background: The Need for Safety-by-Design in Foundation Models
As generative AI evolves from text-only systems to multimodal models capable of processing images, video, and audio, foundation models are rapidly increasing in capability. At the same time, these advancements introduce new risks, including harmful content generation, disinformation, and the creation of illegal or inappropriate material. These risks are becoming major concerns from technical, societal, and regulatory perspectives.
Traditional approaches, such as rule-based filtering or post-deployment guardrails, often struggle to ensure safety without negatively affecting model capability or performance. In multimodal AI systems, risks frequently emerge from interactions between different modalities—such as text, images, and video—making risk identification and mitigation significantly more complex.
As global regulatory frameworks and industry standards continue to evolve, generative AI systems are increasingly expected to adopt “Safety-by-Design” and Responsible AI principles, integrating safety considerations directly into model architecture, training data, and development workflows.
Within the Japanese market, several challenges have become particularly evident:
Limited Japanese Cultural Context
Models trained primarily on global datasets often lack the contextual nuance required to understand Japanese cultural taboos, gestures, and local legal frameworks such as copyright and personality rights, leading to unintended outputs.
Composite Multimodal Risks
Content that appears harmless in isolation may become problematic when combined with specific prompts or instructions.
Leveraging amanaimages’ decades of experience in rights management and visual content production, Qlean Dataset provides structured, rights-cleared datasets engineered for linguistic precision, contextual intelligence, and enterprise-grade AI reliability.
Service Overview: Safety-Aware Training Data for Multimodal AI
Qlean Dataset provides comprehensive dataset design, creation, and provisioning services for LLMs, VLMs, and image generation models, focusing on:
Dataset Design & Collection
Text prompts and training instructions aligned with Japan-specific cultural and regulatory contexts.Multimodal Risk Datasets
Training data addressing composite risks across image, video, audio, and text modalities.Evaluation & Annotation
Specialized labeling frameworks for intellectual property risks (copyright and trademarks) and demographic fairness.
Core Challenges in Safety Data Preparation
Alignment with Local Ethical Standards
Global datasets often lack the granularity required to reflect Japanese legal frameworks and social norms.Technical and Ethical Data Collection Constraints
Sensitive content must be collected under strict legal compliance and carefully managed annotation workflows.Composite Risk Identification
Risks created by combinations of modalities require advanced domain expertise.Annotator Well-being
Ensuring the mental well-being of annotators handling sensitive content while maintaining annotation consistency through rigorous guidelines.
Data Solutions by Modality
1.Text (LLM): Alignment with Japanese Ethical Context
Localization of Global Safety Benchmarks:
Redefining international safety metrics such as hate speech and harassment to align with Japanese legal frameworks and cultural norms.Safety Focused Instruction Tuning:
Response pairs designed to guide models in safely rejecting or redirecting jailbreak prompts in accordance with Japanese social norms and cultural context.
2.Image Generation: Intellectual Property Protection and Standards in Japan
IP Risk Evaluation:
A multi stage evaluation process that assesses potential infringement risk where prompts may generate outputs resembling specific characters, copyrighted works, or distinctive artistic styles.Japan Standard NSFW Detection:
Detailed tagging and classification of violent and sexual content in accordance with Japanese legal standards and domestic platform moderation policies.
3.Vision-Language Models (VLM): Multimodal Risk Detection
Cross Modal Risk Understanding:
Datasets designed to capture risks that emerge only through the combination of image and text. For example, an image of a landmark paired with a prompt asking how to sabotage or damage the location.Fairness and Bias Mitigation:
Diverse and balanced datasets designed to reduce recognition bias and model bias related to attributes such as race, gender, or age.
About Qlean Dataset
Qlean Dataset is a commercially cleared AI training data solution provided by Amana Images, a subsidiary of Visual Bank Group. The platform offers diverse data formats including image, video, audio, 3D, and text, as well as a specialized AI Data Recipe lineup developed through collaborations with major media organizations and data rights holders.
URL:
https://qleandataset.visual-bank.co.jp/en




About Visual Bank Inc.
Visual Bank Group is a technology company developing data infrastructure and AI solutions that support advanced AI development. The company operates THE PEN, an AI tool for manga creators, and its subsidiary, amanaimages Inc., provides commercial digital content and AI training data solutions, including Qlean Dataset. Visual Bank is also a selected participant in GENIAC, a Japanese government initiative supporting the advancement of next generation AI technologies.
CEO: Saneyuki Nagai
Website
https://visual-bank.co.jp/en





