A large-scale music understanding model trained with iterative multi-scale distillation for music audio analysis.