ICCV2025 Conference Participation Report

Our Paper and Workshop Accepted at ICCV 2025, the Top International Conference on Computer Vision (Participation Report)

Hello, I’m Shunsuke Kitada (@shunk031), and I work at LY Corporation on research and development for image generation and design generation.
From October 19 to 23, 2025, I attended and presented at the International Conference on Computer Vision, ICCV 2025, held in Hawaii, USA.

In this article, I’ll share insights gained from ICCV workshops and the main conference. First, I’ll report on the workshop we organized, where I and other LY employees led international discussions, and introduce the insights gained from a workshop specialized in advertising and design generation. Then, I’ll analyze the latest trends at the main conference—such as acceleration and controllability of diffusion models, and hierarchical / structured design generation—alongside an overview of our own research team’s presentation.

Entrance of the venue
Entrance of the venue

This article is an English translation of the following Japanese article:

Report on the FOUND Workshop

Workshop Overview and Activities as Organizers

Our workshop, titled Foundation Data for Vision: Challenges and Opportunities (FOUND), was organized by LY employees including myself (Kitada) and Dr. Komatsu, and was accepted at ICCV 2025.

Our company’s news page about the workshop acceptance
Our company’s news page about the workshop acceptance
FOUND workshop homepage
FOUND workshop homepage

The theme centers on “data,” which is indispensable for recent advances in foundation models, and the workshop aimed to provide an international forum to discuss the challenges and opportunities around it.

This acceptance represents one of the rare cases where a workshop at ICCV, one of the world’s top-tier conferences, is led by Japanese researchers. It demonstrates that research and industrial efforts originating in Japan are being recognized on a global stage.

LY’s Contribution and On-Site Highlights

LY Corporation took the lead in organizing this workshop. Through this, we were able to demonstrate, at least to some extent, that we are in a position to drive international discussions on foundation data, an urgent topic that will shape the future of the computer vision field.

FOUND workshop opening.
FOUND workshop opening. Opening session by lead organizer Yoshihiro Fukuhara.
FOUND workshop sponsors.
FOUND workshop sponsors. LY Corporation (LINEヤフー) was also listed as a sponsor.

The workshop attracted many researchers and practitioners from both academia and industry, leading to lively technical exchanges. In particular, the joint social event we held together with the LIMIT Workshop, also at ICCV 2025, was extremely successful. It served as a valuable opportunity to foster deeper networking among participants and to expand the possibilities for future international collaborations.

LIMIT & FOUND workshop banquet
LIMIT & FOUND workshop banquet. . Researchers from Google, OpenAI, Salesforce Research, who gave invited talks at the workshops, as well as many members from universities and research institutions joined us.

Based on this experience as organizers, if you are considering hosting a workshop at an international conference, or if you feel uncertain about how to write a proposal or how to manage operations, please feel free to reach out. In particular, Mr. Fukuhara, who served as the lead organizer, and Dr. Kataoka from AIST, who led the co-hosted LIMIT Workshop, both have extensive experience and can provide concrete advice. I (Kitada) would also be happy to support where I can. I sincerely hope that more researchers and engineers from Japan will take the lead in discussions at international conferences.

Insights from Attended Workshops

From the perspective of my domain of responsibility and expertise—image generation, design generation, and business impact (especially in advertising and marketing), I attended and will report on two workshops.

Report on “Computer Vision in Advertising and Marketing (CVAM)”

The CVAM Workshop focused on the latest applications of computer vision (CV) in the fields of digital advertising and marketing.

CVAM workshop page
CVAM workshop page

From a business impact standpoint, the workshop’s theme, applications of CV technologies in advertising and marketing, specifically covered creative generation, optimization of marketing systems, brand intelligence, and more.

In our business, which aims to maximize advertising effectiveness, understanding and generating visual content is an urgent challenge. By participating in CVAM, I was able to reconfirm the importance of discussions on how to connect the latest image generation technologies not only to creative production, but also to ad effectiveness measurement and prediction of consumer behavior.

Report on “Workshop on Graphic Design Understanding and Generation 2025 (GDUG)”

The GDUG Workshop aimed to discuss key concepts, technical constraints, and ethical aspects in the recognition and generation of graphic design and documents.

GDUG workshop page
GDUG workshop page

From the perspective of design generation, a shared concern was that, while many research efforts focus on pixel-based image generation, real-world design workflows—such as creating posters, online ads, and websites—are based on structured documents (e.g., layered object representations, style attributes, typography) rather than raw pixels. This gap between research and practice was repeatedly highlighted.

In the invited talks, cutting-edge methods for layer decomposition were introduced, where existing designs are decomposed into native layers such as text, foreground elements, and background (e.g., the Accordion pipeline).
This clearly showed that design generation is evolving from simply producing a “flat, single-layer output” toward handling layered structures similar to those used in real-world design processes. It also provided important hints for rethinking the direction of our own design generation technology development.

Main Conference Report and Key Trend Analysis

Overall Statistics and Scale of ICCV 2025

ICCV 2025 was held at the Hawaii Convention Center, with nearly 7,000 participants. Over 11,000 papers were submitted, of which around 2,600 were accepted (an acceptance rate of about 24%).

The top three research trends were: generative AI (images and videos), 3D processing from multi-view / multi-sensor data, and multimodal learning. In particular, in the generative model area, I felt that a major trend is the shift from diffusion models to flow models.
For a more detailed overview and analysis of trends, I recommend the ICCV 2025 Report compiled by volunteers from cvpaper.challenge.

ICCV 2025 venue guide posted inside the Hawaii Convention Center
ICCV 2025 venue guide posted inside the Hawaii Convention Center

Introduction of Our Research Team’s Presentation “PINO”

Our research presentation PINO (Person-Interaction Noise Optimization) is a technique that enables long-duration, customizable, arbitrary-size group motion generation. This research is the result of our summer internship program, led primarily by Mr. Ota from Tokyo Institute of Technology.

PINO poster session
Our research scientists Dr. Yu (left) and Dr. Fujiwara (right) presenting the poster. Yu was also selected as an Outstanding Reviewer for the main conference.

In terms of the overview and contributions of PINO, the method employs a unique approach that optimizes the noise input when denoising motion sequences using a base diffusion model (e.g., InterGen).

PINO 1
Figure from PINO [Ota+ ICCV’25]

By this optimization, the generated motions are not only aligned with the text prompts, but also enforced to be physically plausible through cost terms that reduce physical artifacts such as body intersections and penetrations.

PINO 2
Figure from PINO [Ota+ ICCV’25]

Furthermore, users can control root positions, regions, and directions via penalty terms, enabling customizable motion generation.

From the standpoint of my work in image and design generation, I’ll analyze the key trends observed at the main conference, including their potential for business impact.

Efficiency and Reliability of Diffusion Models: The Superiority of Flow Models

The evolution of diffusion models (DMs) is shifting focus from mere realism to speed, controllability, and reliability.

Regarding flow-based generation and fast sampling, Flow Matching models are gaining prominence as a training paradigm for generative models. In particular, Contrastive Flow Matching maximizes the dissimilarity between the estimated flow and an independently sampled flow, thereby consistently outperforming previous Flow Matching methods (better FID scores) and enabling high-quality, fast generation.

Figure from [Stoica+ ICCV’25]
Figure from Contrastive Flow Matching [Stoica+ ICCV’25]

For inversion-free editing, FlowEdit utilizes pretrained flow models (SD3, FLUX.1, etc.) and removes the need for an “inversion” step during editing, instead tracing a shorter, more direct path from the source image distribution to the target distribution. This achieves high text alignment and strong structural preservation (excellent LPIPS), and FlowEdit received the Best Student Paper Award.

Figure from FlowEdit [Kulikov+ ICCV’25]
Figure from FlowEdit [Kulikov+ ICCV’25]

In controllable generation (FlowChef), Rectified Flow Models (RFMs) are leveraged, where sampling trajectories become nearly straight and the nonlinear error term approaches zero. By skipping gradient computation (Gradient Skipping), FlowChef achieves deterministic and efficient controllable generation for tasks such as inpainting and super-resolution—without additional training or large-scale backpropagation.

Figure from FlowChef [Patel+ ICCV’25]
Figure from FlowChef [Patel+ ICCV’25]

Evaluating and Improving Generation Quality (Human-Centered Approaches)

As a method for optimizing image generation based on human preferences, an evaluation model (HPSv3) was proposed and built using a large and diverse dataset (HPDv3: 1.08 million text–image pairs) combined with an uncertainty-aware ranking loss. This enables Model-wise Preference (selecting the optimal model for a given prompt) and Sample-wise Preference (selecting the best sample among multiple generations).

Figure from HPSv3 [Ma+ ICCV’25]
Figure from HPSv3 [Ma+ ICCV’25]

The self-reflection–driven iterative improvement approach (Reflection Tuning) proposes a paradigm in which a reward model or multimodal LLM (MLLM) generates “reflection” prompts that describe the shortcomings of a generated image in text, and the image is then iteratively improved according to these instructions. This allows targeted corrections such as “remove the sunlight” or “change the clothes.”

Figure from Reflection Tuning [Zhou+ ICCV’25]
Figure from Reflection Tuning [Zhou+ ICCV’25]

Structured Generation to Accelerate Design and Advertising Applications

For layer-based structured generation, DreamLayer addresses the long-standing challenge of “layer consistency” in design generation by generating multiple transparent image layers simultaneously in a coherent manner. It uses mechanisms such as Layer-Shared Self-Attention to resolve inconsistencies in occlusion relationships and spatial layout between foreground and background.

Figure from DreamLayer [Huang+ ICCV’25]
Figure from DreamLayer [Huang+ ICCV’25]

Regarding advances in layout control, CreatiLayout (SiamLayout) is a Transformer-based model that generates images from layout (placement information) and demonstrates high performance in spatial consistency, color, shape, and more. Its LayoutDesigner component achieves state-of-the-art accuracy in layout planning tasks, surpassing GPT-4 Turbo.

Figure from CreatiLayout [Zhang+ ICCV’25]
Figure from CreatiLayout [Zhang+ ICCV’25]

For high-fidelity image compositing, DreamFuse proposes a new method called Localized DPO (Localized Direct Preference Optimization) to better fuse foreground and background images in a way that humans prefer. By learning to avoid trivial “copy-and-paste” images (negative samples), the model more appropriately handles fusion-related transformations—such as perspective and affine transformations—leading to improved background consistency and harmony with the foreground.

Figure from DreamFuse [Huang+ ICCV’25]
Figure from DreamFuse [Huang+ ICCV’25]

For accurate visual text synthesis, UniGlyph uses segmentation masks at the pixel level as conditions in a diffusion-based framework, addressing problems such as blurry glyphs and style inconsistencies. It shows strong performance especially for rendering small text and complex layouts.

Figure from UniGlyph [Wang+ ICCV’25]
Figure from UniGlyph [Wang+ ICCV’25]

TextMaster is a unified framework that controls both text glyphs and styles. By integrating OCR techniques to compute L2 losses on character features, it enables realistic text editing.

Figure from TextMaster [Yan+ ICCV’25]
Figure from TextMaster [Yan+ ICCV’25]

For deployment in advertising and marketing, the gaze prediction method ScanDiff proposes using diffusion models to predict scanpaths (gaze trajectories) in response to visual stimuli such as ads. It achieves high performance on datasets like COCO-Search18. This is an important technology for modeling human attention, crucial to optimizing the visibility and effectiveness of ad creatives.

Figure from ScanDiff [Cartella+ ICCV’25]
Figure from ScanDiff [Cartella+ ICCV’25]

A new research theme, understanding advertising videos, has also been proposed with AdsQA, a QA benchmark for video understanding that incorporates challenges unique to ad videos. In this area, issues such as the sensitivity of reinforcement learning methods to data quality and the effects of prompt templates on performance are also analyzed.

Figure from AdsQA [Long+ ICCV’25]
Figure from AdsQA [Long+ ICCV’25]

Conclusion

Through my participation in ICCV 2025, I strongly realized that the computer vision field, especially image and design generation, has entered a phase where the focus has shifted from merely pursuing realism to emphasizing the “practicality and controllability” of the technology.

The trends observed in the workshops and main conference clearly show that generative AI is evolving from pixel-based outputs to the generation of hierarchical and structured design elements that align closely with real design workflows. At the same time, advances in flow models are enabling fast and accurate generation and editing. This will be a key driver in revolutionizing the PDCA cycle for creatives in advertising and marketing applications.

LY Corporation will continue to fulfill its responsibility in governing foundation data and leading international discussions, while proactively applying these cutting-edge generation and control technologies to our business. Through this, we aim to establish a strong technological competitive advantage and contribute to society.

Shunsuke Kitada, Ph.D.
Shunsuke Kitada, Ph.D.
Research Scientist working on Vision & Language with Deep Learning

My research interests include deep learning-based natural language processing, computer vision, medical image processing, and computational advertising.