Machine learning training methods using image generation for detecting illegal sidewalk riding by slow vehicles
Tetsuya Manabe & Shunsuke Katayama
What the paper says
This study proposes and evaluates efficient methods for creating training data using image generation and viewpoint transformation. We aim to prevent illegal sidewalk riding by slow vehicles, including small motorized bicycles. The study demonstrates that introducing image completion of vehicle objects to the conventional method using viewpoint transformation enhances its performance, improving the true positive rate by approximately 6.5% compared to the baseline. In addition, the training data created by image completion of the background area demonstrated a 7.3% higher performance in identifying the riding environment and reduced the training data creation time by 64% compared to the conventional method using viewpoint transformation. These results demonstrate the effectiveness of efficient training data creation using image completion. Because a large amount of training data is necessary to achieve high performance in riding environment identification across various environments, this study provides relevant insights to facilitate safe riding support for slow vehicles. Ultimately, the proposed method contributes to the development of robust monitoring systems that can rapidly adapt to new traffic regulations and diverse road conditions, thereby enhancing the safety of pedestrians and riders. • Efficient methods for creating training data using image generation and viewpoint transformation are proposed. • Introducing image completion of vehicle objects to the conventional method enhanced its performance. • The training data created by image completion of the background area demonstrated higher identification performance. • This study provides relevant insights to facilitate safe riding support for slow vehicles.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.