); (c) the performance gap between major models before and after filtering out shortcut questions. (Image: Naver Cloud)
[Edaily Reporter Han Kwangbeom ] As Team Naver had a total of 23 research papers accepted at “ECCV (European Conference on Computer Vision) 2026,” one of the world’s three major computer vision conferences, Naver (NAVER(035420)) Cloud unveiled four papers that focused on AI’s “process of understanding” and “evaluation system.”
Naver Cloud announced on the 28th that at ECCV 2026, it diagnosed the AI’s tricks and judgment errors hidden behind surface-level benchmark scores and proposed learning methods to correct them.
First, Naver Cloud analyzed the shortcomings of existing video benchmarks used to evaluate AI’s video understanding capabilities (paper: Video-Oasis: Rethinking Evaluation of Video Understanding). After conducting six types of tests—including video removal and scene reordering—on 14 existing benchmarks, the results showed that 55% of all questions could be answered correctly without the AI actually viewing the video. When re-evaluated after excluding these questions, the accuracy of major models ranged from only 26% to 37%. Rather than creating new evaluation criteria, Naver Cloud proposed a tool to diagnose the existing test sets themselves.
The study also addressed the issue of AI making habitual judgments about actions based on objects (Paper: “Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition”). For example, after learning from training data that “drawers” are primarily associated with “closing,” the AI would infer “closing” simply by seeing a drawer, without observing the actual action. The research team improved the precision of action recognition by introducing video synthesis techniques and reverse playback learning—achieving this solely through improvements to the training method, without the need for additional data collection or model expansion.
than existing methods (center) (right, green). (Image: Naver Cloud)
The company also announced “Phoenix,” a mask refinement technology that exploits AI vulnerabilities to improve accuracy (Paper: Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation). By utilizing “adversarial attack” techniques—which are designed to disrupt AI—the team automatically generated training data for areas where the AI actually becomes confused. This boosted the accuracy of object contour extraction and fine-grained segmentation by up to 21 percentage points.
They also presented “SpatialBoost,” a study that trains language models to learn the 3D spatial structure within images (paper: SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning). The approach involved extracting depth and location information from 2D images and converting it into sentences, then building spatial understanding step-by-step—starting with points, then objects, and finally the entire scene—before feeding this information into the visual module via a language model. This simultaneously reduced the cost of building 3D data while improving performance in robot manipulation and image classification.
A Naver Cloud representative stated, “These accepted papers are significant in that they verify whether AI is actually processing information correctly and suggest directions for improvement,” adding, “We will continue to research the very process by which AI understands and makes judgments to build a trustworthy AI ecosystem.”