8/13/2026

Visual Bank Makes Rights-Cleared “Qlean Dataset” Data Available Free to Academia via Hugging Face Hub

July 30, 2026, Tokyo (Japan) ー Visual Bank Inc. (Head Office: Minato-ku, Tokyo; CEO Saneyuki Nagai; hereinafter “Visual Bank”) has begun offering datasets from Qlean Dataset, an AI training-data solution for foundation models operated by its subsidiary, amana images inc., free of charge to academic researchers through Hugging Face Hub. 

Under a dedicated “Academic Research License,” researchers affiliated with universities and research institutions may use the datasets not only for viewing and quality assessment, but also for training, fine-tuning, and evaluating AI/ML models for academic research purposes. Researchers may also cite and present the datasets in academic papers at no cost. The publication of trained models for non-commercial academic purposes is also permitted (*1). 

This initiative expands Visual Bank’s ‘Academia Support Program’, launched in 2025, which already supports more than 10 research laboratories, by making its resources available through Hugging Face Hub, a global platform widely used by AI researchers and developers. 

*1 Commercial use requires a separate commercial license agreement.

Background and Objectives: Addressing the Structural Shortage of Japan-Specific Research Data

As competition in the development of foundation models continues to intensify, access to high-quality training data has become a factor that determines the success or failure of AI research. For academia institutions, however, limited research budgets can make the costs associated with data collection, rights clearance and dataset preparation a substantial barrier. As a result, independently securing sufficient and appropriate data for research remains challenging. 

This has led researchers to rely heavily on publicly available datasets originating outside Japan. On Hugging Face Hub, which is used by AI researchers and developers worldwide, approximately 3,000 Japanese-language datasets are available, representing only around 0.3% of the total and approximately one-thirtieth of the number of English-language datasets, which stands at roughly 85,000 (*2). 

Furthermore, many Japanese-language datasets are derived from the Japanese portions of multilingual corpora developed overseas, while datasets specifically designed and collected within Japan remain even more limited. Japanese content also accounts for only approximately 5% of Common Crawl, one of the principal data sources used for the pre-training of large language models (*3). 

As researchers around the world increasingly work with the same limited pool of publicly available data, it is becoming more difficult to differentiate research outcomes (*4). At the same time, computing resources, large-scale datasets, and specialized AI talent have become increasingly concentrated within the private sector. Reflecting this trend, approximately 90% of major AI models announced in 2024 were led by industry (*5). 

Japan-specific data—including conversational Japanese featuring regional dialects, images depicting Japanese landscapes and culture, and video capturing everyday life and industrial environments—remains particularly scarce on open platforms. This lack of accessible data has limited opportunities to pursue research themes that reflect Japan’s distinctive linguistic, cultural, social, and industrial characteristics. 

To address this structural challenge, Visual Bank, through amana images inc., has operated the “Academia Support Program” since 2025. Through the program, more than 500,000 samples spanning 80 dataset categories have been provided free of charge for research and development purposes. Participating institutions include the University of Tokyo, RIKEN, and the National Institute of Advanced Industrial Science and Technology (AIST), with more than 10 research laboratories supported to date. 

The latest initiative extends the Academia Support Program to Hugging Face Hub, providing academic researchers with easier access to rights-cleared, commercial-grade datasets in a format suitable for research use . 

By expanding access to high-quality, Japan-specific data, we aim to foster an environment in which researchers can pursue distinctive research questions and develop new AI technologies grounded in data originating from Japan. 

*2 Based on public datasets on Hugging Face Hub for which the creator applied a language tag (as of July 2026, Visual Bank’s own count). Datasets tagged Japanese number 3,942 and datasets tagged English number 85,069. Language tags are optional metadata, and most of the total public datasets on the Hub (approximately 1.05 million, including non-text data such as images and audio) are untagged and therefore excluded from the count.  
*3 Common Crawl. “Distribution of Language”. Common Crawl Official Statistics.
*4 Ahmed, Wahed & Thompson, "The growing influence of industry in AI research", Science, 2023 
*5 Stanford Institute for Human-Centered Artificial Intelligence. AI Index Report 2025. Stanford University, 2025. According to the report, approximately 90% of notable AI models released in 2024 were industry-led, up from approximately 60% in 2023.

Release Details

■ Initial Dataset Lineup  (Examples from the launch lineup)

Japanese Two-Speaker Emotional Conversational Speech Dataset

Two-person conversational speech incorporates four emotional expressions: excitement, anger, sadness, and joy. Each item contains approximately 20 minutes of audio. 

Japanese Two-Speaker L/R-Separated Private Conversation and Transcript Dataset

Facial images of 200 Japanese individuals across a range of ages and genders, captured under various conditions, including different facial angles and with or without accessories. 

Japanese Facial Image Dataset Across Age Groups and Genders, With and Without Accessories 

Video data recreating the self-promotion (“self-PR”) stage commonly used in Japanese new-graduate recruitment. A total of 72 students speak directly to the camera about their strengths, skills, and personal experiences.

■ Release Format

Datasets for which Qlean Dataset holds the necessary rights will be released in phases through the Qlean Dataset Organization on Hugging Face Hub. The releases will also include the rich metadata that distinguishes Qlean Dataset, provided in a fully rights-cleared format for academic use.

■ Terms of Use: Academic Research License

Eligible users  

・ Faculty members, staff, and students affiliated with universities, other higher-education institutions, research organizations, and government-affiliated research institutions, when using the data for academic research purposes. 
・Independent researchers whose primary purpose is to publish research findings in academic journals, conferences, or other scholarly media. 
・ Individuals and organizations evaluating the datasets for the purpose of considering a separate commercial license agreement. Use in production environments is not permitted during such evaluation.

Permitted Uses Free of Charge

Viewing, inspecting and assessing of the data; quality evaluation and non-commercial benchmarking; training, fine-tuning, and evaluation of AI/ML models for academic research purposes; citation and publication in academic papers and conference presentations; and publication of trained models for non-commercial, academic purposes.  

Excluded and prohibited uses

Model training by organizations that generate revenue from AI/ML products or services (including those with a research division); use in commercial products or services; and redistribution, mirroring, or resale.

Credit

When using the datasets, the following credit should be included: "Data provided by: Visual Bank / Qlean Dataset (amana images Inc.)" 

Commercial use

Provided under a separate paid license agreement. For inquiries: https://qleandataset.visual-bank.co.jp/contact 

Governing law and jurisdiction

The license is governed by the laws of Japan. The Tokyo District court shall have exclusive jurisdiction as the court of first instance for any disputes arising in connection with the license. 

* For complete terms and conditions, please refer to the full license text published on each dataset page on Hugging Face Hub. 

■ Hugging Face Hub page 
https://huggingface.co/qleandataset 

Comment from a Researcher Supported by the Program 

Professor Yutaka Matsushita, Department of Media Informatics, College of Information and Frontier Sciences, Kanazawa Institute of Technology 

“Rights-cleared Japanese speech data is an extremely valuable research resource that we can use with confidence, and it has contributed significantly to the advancement of our research and development. 

Our research focuses on generating speech-bubble-style subtitles for video content by predicting emotions, such as joy and anger, from the voices of people appearing in the video. Looking ahead, if datasets with more detailed annotations indicating the intensity of each emotion become available, we expect it will become possible to predict not only the type of emotion, but also its degree of intensity. 

Access to rights-cleared Japanese-language data also contributes to improving the accuracy of emotion prediction. By changing the shape and presentation of speech bubbles according to the emotion being expressed, we believe this research could ultimately lead to more accessible and intuitive subtitles for people with hearing impairments. 

As one of our research outcomes, we analyzed five categories of emotion using a convolutional neural network (CNN) and achieved a very high level of prediction accuracy. We are extremely grateful for the valuable data provided, which played an important role in achieving these results. 

Going forward, we plan to further improve prediction accuracy by expanding the network architecture with additional hidden layers and conducting more detailed analysis.” 

Future Outlook 

Building on this initial release, Visual Bank plans to gradually expand the range of multimodal datasets available on Hugging Face Hub, including audio, image, and video data. 

We will continue providing individualized support to research institutions in Japan through its Academia Support Program. This will include responding to requests for specific datasets as well as providing custom recording and data collection services tailored to individual research themes. 

Through these initiatives, we aim to establish an environment in which rights-cleared, Japan-specific data can serve as a foundation for research conducted around the world, while contributing to greater diversity and originality in AI research originating from Japan. 

About Qlean Dataset 

Qlean Dataset is an AI training-data solution for foundation-model development, provided by amana images Inc., a subsidiary of Visual Bank. 

For more than 40 years, amana images has worked with content and data entrusted to it by rights holders, including broadcasters and publishers, properly managing the associated rights and making those assets available for legitimate use. With the principle that ‘protecting the rights associated with data is what enables rights holders to entrust their valuable content to us’. This long-standing approach to rights management forms the foundation of Qlean Dataset. 

Drawing on its experience in clearing rights for original, primary-source data that is not readily available on the open web, Qlean Dataset provides data in a form that is ready for AI training. This expertise has led to large-scale data deliveries to foundation-model developers in Japan and internationally. 

Qlean Dataset supports a wide range of data modalities, including audio, images, video, 3D data, and text. Its dataset portfolio continues to expand through partnerships with data holders and media organizations in Japan and overseas. Custom data recording and collection services are also available to meet specific development and research requirements.  

Qlean Dataset Website: https://qleandataset.visual-bank.co.jp/
AI Data Recipe: https://qleandataset.visual-bank.co.jp/lineup
Inquiries: https://qleandataset.visual-bank.co.jp/contact 

About Visual Bank Inc.  

Visual Bank Inc. is a next-generation startup focused on building advanced data infrastructure to accelerate AI development. Guided by its mission to “unlock the full potential of all data,” the company develops solutions that enable high-quality data to be used more effectively across the AI ecosystem. 

Its wholly owned subsidiary, amana images Inc., provides ‘THE PEN,’ an AI-powered support tool designed to help manga artists “draw more!,” as well as ‘Qlean Dataset,’ a service for developing AI training datasets. 

The company has also been selected for GENIAC, a national research and development program in Japan, and is accelerating efforts to bring its technologies and services into real-world use. 

CEO: Saneyuki Nagai  
Address: 6F, C-Cube Minami-Aoyama Building, 7-1-7 Minami-Aoyama, Minato-ku, Tokyo 107-0062  
Visual Bank corporate URL: 
https://visual-bank.co.jp/  
amana images corporate URL: https://amanaimages.com/about/

    amana images inc.

    Visual Bank Inc.


    © amanaimages inc.