Description
Galaxy spectra encode detailed information including redshift, ionisation conditions, metallicity and stellar content, but the features needed to recover these quantities are often faint, blended or unmeasurable in faint and distant sources, and labelled training data remain limited. We present SpecML, a self supervised foundation model that learns general representations of galaxy spectra from unlabelled data and transfers them to tasks where labels are scarce. SpecML adapts the OmniSpectra architecture for use with JWST NIRSpec prism spectra drawn from the DAWN archive. Each spectrum is tokenised into overlapping flux patches, and a sinusoidal encoding of mean patch wavelength supplies global positional information. The model is pretrained with a masked reconstruction objective using a validity aware transformer that adds local structure through a depthwise convolution, allowing it to process variable length spectra at native resolution and to reconstruct masked regions from surrounding context. We assess the learned representations with downstream probes. SpecML recovers redshift with a coefficient of determination of 0.936 and a NMAD of 0.035 across 10746 sources, with the largest scatter confined to the low redshift population. We extend the same representations to source classification and galaxy property estimation, and test whether SpecML can predict the evolution of star forming galaxies and active galactic nuclei on the BPT diagram beyond z = 3, where the diagnostic lines typically become difficult to measure directly. These results show that self supervised pretraining recovers physically meaningful information that would otherwise require strong individual line detections.
| I am the presenting author | Yes |
|---|