What are the training data for the pre-trained model release? Is speech and instruments optimized other than the singing voice?
What are the training data for the pre-trained model release? Is speech and instruments optimized other than the singing voice?