
[About the Annual Conference of the Japanese Society for Artificial Intelligence]
“This annual conference brings together researchers from across Japan working on artificial intelligence to present their research findings. Through presentations on the latest technology trends, new research results, and ideas related to AI, it provides a forum for exchanging opinions and building connections.”
(Excerpt from the official website: https://www.ai-gakkai.or.jp/jsai2021/)
Event announcement: 2024 Annual Conference of the Japanese Society for Artificial Intelligence (38th)
Dates: Tuesday, May 28 – Friday, May 31, 2024
Venue: Act City Hamamatsu (Hamamatsu, Shizuoka Prefecture) + Online
[About the presentations]
We will present two joint research results with Honda R&D Co., Ltd.
◾Explaining Traffic Risk Using LLMs with GIS Data and Street-View Images
Session time: 17:40–18:00, May 28, 2024
Venue: Room D
Presentation link: https://confit.atlas.jp/guide/event/jsai2024/subject/1D5-GS-10-03/advanced
Abstract
Considering traffic risk in driver-assistance systems and autonomous driving technology is important for preventing traffic accidents, and traffic risk is thought to be heavily embedded in image information. However, explaining the traffic risk contained in a driving scene from image information alone is difficult, and research in this field has not yet progressed sufficiently. This study proposes a multimodal framework that combines GIS data and street-view images to explain traffic risk. The framework identifies the coordinates of high-risk areas from a traffic accident risk map created using GIS data, and trains a multimodal network using street-view images associated with those areas. This builds a framework that effectively explains traffic risk in any given scene. Experimental results confirmed that the proposed framework can generate captions that explain traffic risk for high-risk areas identified from GIS data.
◾Generating Caption Data Using Prompt Engineering for Road Environment Risk Analysis
Session time: 18:00–18:20, May 28, 2024
Venue: Room D
Presentation link: https://confit.atlas.jp/guide/event/jsai2024/subject/1D5-GS-10-04/advanced
Abstract
As driver-assistance systems and autonomous driving technology have become more widespread, they have shown a certain effect in reducing traffic accidents, but interpreting traffic accident risk and analyzing its mechanisms is important for further reducing accidents. In research on multimodal networks that explain driving scenes, methods have been attempted that generate captions considering recognizable objects using metadata. Such methods typically generate captions focused on dynamic objects such as people. However, to interpret the traffic accident risk contained in a driving scene, static risks arising from factors such as road signs and road structures should also be considered when generating captions. Existing large-scale multimodal networks struggle to generate captions that address this type of road environment risk. To address this challenge, we propose a caption generation method that leverages prompt engineering to encompass both dynamic objects and static latent risks. Experiments using the generated captions also confirmed that captions considering both dynamic objects and static latent risks can be generated.
