- Research team led by professor Kim Eui-hwan of AI Convergence Department

- Performance verified through real robot experiments

Kim Eui-hwan (left), professor in the AI Convergence Department at GIST, and Yoon Ji-ae, an integrated master's-doctoral student. [GIST]
Kim Eui-hwan (left), professor in the AI Convergence Department at GIST, and Yoon Ji-ae, an integrated master's-doctoral student. [GIST]

Home service robots may soon be able to determine on their own whether a frequently used item has gone missing, whether something new has been placed in a room, or whether a potentially hazardous change has occurred.

A research team led by Kim Eui-hwan, a professor in the AI Convergence Department at the Gwangju Institute of Science and Technology (GIST), has developed an AI model called VSCDNet that automatically detects real-world object changes by comparing video footage of the same space recorded at different times and along different paths.

For autonomous robots to navigate indoor spaces over extended periods, they must be able to continuously track changes in their surroundings.

Indoor environments change frequently — objects appear or disappear, positions shift, and the states of doors and furniture vary. Such changes can serve as critical cues for safe robot navigation, indoor surveillance, object management and continuous learning. Most existing scene-change detection research, however, has assumed fixed cameras or image pairs captured from similar viewpoints. Because real robots cannot always retrace the exact same path and angle, technology capable of identifying what has changed between videos shot at different times and from freely varying positions is needed.

To overcome this limitation, the research team shifted focus from comparing individual images to analyzing the overall flow of video sequences.

The AI model compares a reference video of a space recorded in the past with a current video of the same space, identifies corresponding scenes across the full footage, and precisely pinpoints only the areas where actual object changes have occurred.

Based on this analysis, the model generates a "change mask" that visually marks the altered regions and presents the final set of detected changes — enabling it to automatically identify real-world object changes such as a laptop going missing or an item moved to a new location.

To systematically validate the model's performance, the team built a large-scale dataset from scratch that includes both virtual-environment and real indoor-environment data, comprising a total of 1,090 video sequences.

In experiments, VSCDNet outperformed existing change-detection methods on both the virtual and real indoor datasets. It also maintained stable detection performance across varying conditions, including different video lengths, image quality levels and numbers of changed objects.

An example of VSCD in use in a real robot environment. A mobile robot repeatedly visits the same space and compares footage from each visit to detect changes such as a door opening or objects appearing and disappearing. The detected change areas can be used for visual surveillance and incremental object learning. [GIST]
An example of VSCD in use in a real robot environment. A mobile robot repeatedly visits the same space and compares footage from each visit to detect changes such as a door opening or objects appearing and disappearing. The detected change areas can be used for visual surveillance and incremental object learning. [GIST]

In experiments using an actual mobile robot, the model automatically detected situations such as a door opening or an object disappearing as the robot moved along different routes, and also demonstrated the ability to register and learn newly appeared objects.

The technology is expected to find applications across a range of fields, including indoor patrol robots, smart security surveillance, facility management and IoT-based smart indoor systems.

"VSCDNet is an AI model that goes beyond recognizing the current scene — it identifies on its own what has changed compared to the past," Kim said. "Because it can compare videos taken along different routes without requiring separate location data or spatial maps, we expect it to be used across a wide range of applications, including indoor patrol robots, smart security surveillance and facility management."

The findings are set to be presented at ICML 2026 (International Conference on Machine Learning), an international AI and machine learning conference, on July 6.


nbgkoo@heraldcorp.com