Bridging Data Gaps in Open-Pit Mining: A Factor-Isolated Framework for Long-Tail Scenario Generation

Abstract:
The development of reliable autonomous driving systems for mining is critically limited by the scarcity of long-tail scenario data. While generative models provide a promising direction, they often lack the fine-grained control and domain-specific realism required in safety-critical applications. To address this challenge, we propose a decoupled training pipeline that models a scene's fundamental components, namely style, layout, and motion, as architecturally separated and controllable factors. Firstly, style and layout are jointly fine-tuned to capture the unique textures and spatial structures of mining environments, ensuring high-fidelity visuals and precise spatial control. Subsequently, we selectively fine-tune the temporal layers of the video generation model. This captures empirically plausible visual motion patterns of heavy machinery, creating short, high-fidelity videos. Experiments show that our pipeline not only outperforms general video baselines but also improves downstream task performance. By bridging the gap between static and dynamic data, our framework supports robust perception and supplies synthetic dynamic scenarios for vision-language-model (VLM)-based training and evaluation in autonomous mining.
Index Terms: Autonomous driving, image generation, video generation, diffusion model
Published in:The International Journal of Intelligent Control and Systems (Volume: 31, Issue: 2, 2026-06-25)
Page(s):205 - 213