Augmenting X-ray astronomical representations with text
摘要
Astronomers have accumulated a large amount of multimodal data of astrophysical sources, such as images, spectra, and time series, alongside decades of literature describing their physical properties. However, these complementary data sources have not yet been extensively integrated within unified representation learning frameworks tailored to astrophysical inference. This work introduces a contrastive learning framework that aligns X-ray spectra with textual descriptions of astronomical sources, enabling the creation of multimodal representations. This connection is non-trivial: X-ray spectra capture only part of the characteristics of astronomical objects, while texts capture diverse physical information. Our best model achieves 20% Recall@1%, indicating statistically non-random cross-modal alignment under strict retrieval evaluation and potentially accelerating the interpretation of rare or poorly understood sources. In addition, the latent dimensions exhibit measurable correlation with selected physical observables, suggesting partial physical interpretability. By combining spectral and textual information, we reduce Mean Absolute Error (MAE) for selected catalog parameters by approximately 16-18% relative to spectra-only baselines. Using a Mixture of Experts architecture that leverages both unimodal and shared representations yields the best performance. Finally, the outlier analysis on the multimodal latent space reveals objects worthy of follow-up investigation, such as a candidate pulsating ULX (PULX) and a gravitational lens system.