Abstract
Multi-sensor perception is central to autonomous driving because cameras, LiDAR, and radar fail in different ways. Recent bird's-eye-view (BEV) fusion models have improved 3D detection and semantic scene understanding by projecting heterogeneous observations into a shared spatial representation, but many fusion pipelines still treat each modality as uniformly reliable within a frame. This paper proposes a lightweight reliability-aware BEV fusion module that estimates per-modality and per-region confidence before feature aggregation. The module can be attached to existing BEV perception backbones and is designed to down-weight degraded camera, LiDAR, or radar evidence without requiring a new end-to-end driving stack. We formulate the method as a three-stage pipeline: BEV feature extraction, reliability estimation from measurement consistency and feature uncertainty, and gated fusion for occupancy and object-level outputs. The intended evaluation uses nuScenes, Waymo Open Dataset, and Argoverse-style sensor logs under ordinary and degraded sensing conditions. The expected contribution is not a new perception foundation model, but a compact safety-facing fusion layer that makes downstream planning inputs more interpretable and less brittle when one modality becomes unreliable.



![Author ORCID: We display the ORCID iD icon alongside authors names on our website to acknowledge that the ORCiD has been authenticated when entered by the user. To view the users ORCiD record click the icon. [opens in a new tab]](https://www.cambridge.org/engage/assets/public/coe/logo/orcid.png)