8/6/2026
Qlean Dataset of Visual Bank Group and Japan’s National Institute of Informatics (NII-LLMC) Release “msts-japanese,” the First Multi-modal AI Safety Dataset in Japanese

August 6, 2026, Tokyo (Japan) ー Visual Bank Inc. (Head Office: Minato-ku, Tokyo; CEO Saneyuki Nagai; hereinafter “Visual Bank”) announces that the rights-cleared photographic images provided through ‘Qlean Dataset’, an AI training data solution for foundation models operated by its subsidiary amanimages inc., have been incorporated into “msts-japanese”, a safety evaluation dataset for vision-language models (VLMs) developed by the Center for Large Language Model Research and Development at the National Institute of Informatics (“NII-LLMC”). The dataset was released on Hugging Face on July 7, 2026 (*1).
Companies and research institutions developing foundation models and VLMs can use “msts-japanese” to conduct multimodal safety evaluations (*2) tailored to the Japanese language and the cultural and everyday contexts of Japan, using rights-cleared photographic images. The dataset is available exclusively for purposes related to improving the safety of LLMs, including commercial applications (*1).
Research conducted by the National Institute of Informatics found that simply changing the text prompt from English to Japanese while presenting the same image resulted in a two- to threefold increase in the rate of harmful responses across all three VLMs evaluated (*3). As AI safety can vary depending on language and cultural context, establishing a robust Japanese-language evaluation framework has become a shared challenge for companies and institutions developing and providing foundation models in Japan.
■ Background: Addressing the Gap in Japan-Specific AI Safety Evaluation
Since the emergence of ChatGPT, many leading AI models have evolved into vision language models, or VLMs, capable of processing image inputs, accelerating their deployment across a wide range of real world applications. While companies have made progress in developing safety evaluation frameworks for text only systems, publicly available Japanese language evaluation datasets specifically designed to assess multimodal risks have remained limited. These risks include cases in which an image and a text prompt may each appear harmless in isolation but produce harmful content when interpreted together (*3).
The Multimodal Safety Test Suite, or MSTS (*4), serves as an international evaluation framework for multimodal safety and contains prompts in 11 languages. However, it does not include Japanese, and its images primarily reflect Western social and everyday life contexts.
As mentioned above, changing the language of a prompt to Japanese can substantially affect the rate at which a model produces harmful responses. Therefore, the safe use and deployment of AI in Japan requires evaluation using both Japanese language prompts and images grounded in Japanese cultural and everyday life contexts.
Images used for safety evaluation must also satisfy requirements specific to this type of research. First, because safety evaluations are intended to measure how models respond to real world situations, the images should be authentic photographs rather than AI generated content. Second, because such evaluations may involve sensitive subjects, including hazardous materials or identifiable individuals, the necessary rights and permissions relating to both the depicted subjects and the photographers must be appropriately cleared.
Finally, it is not easy to source Japanese photographic images from the open web that satisfy both requirements and are also suitable for inclusion in a publicly released evaluation dataset.
■ Overview of msts-japanese and scope
Name | msts-japanese |
|---|---|
Published by | LLM Study Group, a research community led by llm-jp and the NII Large Language Model Research and Development Center (NII-LLMC) |
Release date | July 7, 2026, on Hugging Face |
Content | Of the 400 items included in MSTS (*4), 330 have been translated into Japanese and provided with Japanese reference responses. Among these, 82 items additionally include “Japan-localized images” reflecting everyday life and cultural contexts in Japan, together with dedicated prompts and reference responses. |
Terms of Use | Registration is required. Use is limited to purposes related to improving the safety of large language models. Commercial use is permitted, but redistribution is prohibited. |
URL |
Amanaimages Inc., (hereinafter “amana images”) contributed the photographic assets used as the “Japan-localized images” in the dataset. All contributed assets consist exclusively of authentic photographs for which the necessary rights clearances and permissions from the respective photographers and rights holders have been secured. No AI-generated or synthetically produced images are included in this image set.
■ The Role of Qlean Dataset: Providing Data Across the Foundation Model Development Lifecycle, from Training to Evaluation
While the original MSTS benchmark uses images sourced from publicly available web resources, including public-domain materials (*4), its Japanese-language adaptation incorporates rights-cleared, authentic photographs that amana images has lawfully obtained from and managed on behalf of the respective rights holders.
The provenance, licensing status, and rights-clearance process for each data asset are clearly documented. This enables research institutions to incorporate and publish the images as evaluation data without having to conduct separate rights investigations. Such traceability is particularly important in AI safety research, where datasets may address sensitive or potentially harmful contexts and therefore require a high degree of legal, ethical, and methodological accountability.
Safety-evaluation datasets should meet the same fundamental data-governance standards as datasets used to train foundation models. These include verified rights clearance, transparent data provenance, and the use of authentic real-world data that accurately reflects the environments and cultural contexts being evaluated.
amana images supports AI developers and research institutions in Japan and internationally not only through the provision of training data, but also as a data partner across the broader foundation model development lifecycle, including benchmarking, model evaluation, safety assessment, and risk mitigation.
■ Comment from Hisami Suzuki, Center for Large Model Research and Development, National Institute of Informatics
“AI safety can vary significantly depending on linguistic and cultural context. In our research, we found that simply changing a question about the same image from English to Japanese increased the frequency of harmful responses generated by the model. This finding demonstrated the urgent need for evaluation data that reflects both the Japanese language and everyday contexts in Japan.
At the same time, images used for safety evaluation must be authentic photographs that depict real world situations, and their rights clearance status must be clearly established. The provision of rights cleared photographic images by amana images enabled us to develop an evaluation dataset that we can release publicly with confidence.
We will continue to expand the evaluation data available for Japanese language contexts and contribute to improving AI safety across the Japanese speaking community.”
■ Future Outlook
amana images will continue to support the development of multimodal safety evaluation data led by NII LLMC, while expanding its provision of evaluation data for applications including safety assessment and red teaming (*5).
The company also plans to exhibit at MIRU 2026, one of Japan’s largest academic conferences on image recognition and understanding, to be held in Nagasaki from August 3 to 6, 2026, where it will present this initiative as a case study.
About Qlean Dataset
Qlean Dataset is an AI training-data solution for foundation-model development, provided by amana images Inc., a subsidiary of Visual Bank.
For more than 40 years, amana images has worked with content and data entrusted to it by rights holders, including broadcasters and publishers, properly managing the associated rights and making those assets available for legitimate use. With the principle that ‘protecting the rights associated with data is what enables rights holders to entrust their valuable content to us’. This long-standing approach to rights management forms the foundation of Qlean Dataset.
Drawing on its experience in clearing rights for original, primary-source data that is not readily available on the open web, Qlean Dataset provides data in a form that is ready for AI training. This expertise has led to large-scale data deliveries to foundation-model developers in Japan and internationally.
Qlean Dataset supports a wide range of data modalities, including audio, images, video, 3D data, and text. Its dataset portfolio continues to expand through partnerships with data holders and media organizations in Japan and overseas. Custom data recording and collection services are also available to meet specific development and research requirements.
Qlean Dataset Website: https://qleandataset.visual-bank.co.jp/
AI Data Recipe: https://qleandataset.visual-bank.co.jp/lineup
Inquiries: https://qleandataset.visual-bank.co.jp/contact
About Visual Bank Inc.
Visual Bank Inc. is a next-generation startup focused on building advanced data infrastructure to accelerate AI development. Guided by its mission to “unlock the full potential of all data,” the company develops solutions that enable high-quality data to be used more effectively across the AI ecosystem.
Its wholly owned subsidiary, amana images Inc., provides ‘THE PEN,’ an AI-powered support tool designed to help manga artists “draw more!,” as well as ‘Qlean Dataset,’ a service for developing AI training datasets.
The company has also been selected for GENIAC, a national research and development program in Japan, and is accelerating efforts to bring its technologies and services into real-world use.
CEO: Saneyuki Nagai
Address: 6F, C-Cube Minami-Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062
Visual Bank corporate URL:
https://visual-bank.co.jp/
amana images corporate URL: https://amanaimages.com/about/
*1. llm-jp. “msts-japanese.” Hugging Face, 7 July 2026, https://huggingface.co/datasets/llm-jp/msts-japanese. Accessed 4 Aug. 2026. Use of the dataset is governed by the AnswerCarefully Dataset Terms of Use. Its use is limited to purposes related to improving the safety of large language models. Commercial use is permitted, but redistribution is prohibited.
*2. The evaluation of safety in generative AI models that process multiple types of data, such as text and images, simultaneously.
*3. Suzuki, Kumi, Tetsuro Takahashi, and Su Myat Noe. “Extending the AnswerCarefully Dataset: Adding Regionally Sensitive Issues and Multimodal Questions.” Proceedings of the 32nd Annual Meeting of the Association for Natural Language Processing, Mar. 2026. The harmful response rate is measured as the Violation Rate, defined as the proportion of responses receiving a score of 1 or 2 on a five-point safety evaluation scale. The figures are based on the preliminary study reported in the paper, which evaluated three vision-language models using 40 images and two types of prompts.
*4. Röttger, Paul, et al. “MSTS: A Multimodal Safety Test Suite for Vision-Language Models.” arXiv, 2025, arXiv:2501.10057. The test suite comprises 200 images and prompts in English, together with translations into 10 additional languages. According to the paper, the images were sourced from publicly available web content in the public domain or under CC BY or CC0 licenses.
*5. Red teaming is a method of systematically assessing system vulnerabilities by intentionally conducting adversarial tests from the perspective of a potential threats.





