Structured prompt-guided zero-shot cross-modal reasoning for SME financial risk identification
摘要
Small and medium-sized enterprises (SMEs) often face financial risk assessment difficulties due to limited historical records, irregular information disclosure, and fragmented multimodal evidence. To address this problem, this study proposes a cross-modal zero-shot diagnosis framework, named CM-ZSD, for SME financial risk identification in the digital economy. The framework combines a lightweight dual-encoder architecture, an adaptive fusion module with gated attention, and structured risk prompts that convert abstract risk categories into semantically enriched textual descriptions. In this way, the diagnosis task is reformulated as cross-modal semantic matching rather than conventional supervised classification. To evaluate the method, we construct the SME-RiskZero benchmark, a simulated dataset covering eight previously unseen risk categories. Experimental results show that CM-ZSD achieves 73.1% average accuracy and 72.2% macro F1, outperforming the strongest baseline by 14.9 and 15.2% points, respectively. Ablation studies further confirm the importance of structured prompts and adaptive multimodal fusion. In addition, lightweight backbone experiments show that the framework retains competitive performance under constrained computing resources. These findings suggest that integrating structured domain knowledge with cross-modal reasoning is a promising direction for zero-shot financial risk diagnosis in SMEs.