We are excited to announce the XACLE Challenge 2027, following the success of last year's challenge! Many x-to-audio generation models have recently been developed. However, existing automatic evaluation methods still do not reliably reflect human perception. In this challenge, we will provide a shared training dataset consisting of synthetic audio and corresponding human subjective scores. Participants will develop automatic evaluation systems that predict audio–text semantic alignment scores with high correlation to human subjective evaluations.
Last year's challenge was conducted under a seen-generation-system setting, where the generation systems used to create the test data were also included in the training data. This year's challenge will be held under an unseen-generation-system setting. Participants will develop systems that predict audio–text semantic alignment scores for test data generated by previously unseen generation systems while maintaining high correlation with human subjective evaluations. Further details of the task will be announced when the challenge officially opens. We warmly welcome researchers from both academia and industry to join the challenge.