FIM-Stereo: Mono-to-Stereo Audio Upmixing via Neural-Codec Token Prediction with a Text-Conditioned Language Model | Synapse