LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection

Chou, Benjamin Shiue-Hal; Jajal, Purvish; Eliopoulos, Nicholas John; Davis, James C.; Thiruvathukal, George K; Yun, Kristen Yeon-Ji; Lu, Yung-Hsiang

LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection

Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, James C. Davis, George K Thiruvathukal, Kristen Yeon-Ji Yun, Yung-Hsiang Lu

ICLR 2026

/iclr/2026/chou2026iclr-laddersym/

Abstract

Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristics or learnable models. This paper introduces \textit{LadderSym}, a novel Transformer-based method for music error detection. \textit{LadderSym} is guided by two key observations about the state-of-the-art approaches: (1) late fusion limits inter-stream alignment and cross-modality comparison capability; and (2) reliance on score audio introduces ambiguity in the frequency spectrum, degrading performance in music with concurrent notes. To address these limitations, \textit{LadderSym} introduces (1) a two-stream encoder with inter-stream alignment modules to improve audio comparison capabilities and error detection F1 scores, and (2) a multimodal strategy that leverages both audio and symbolic scores by incorporating symbolic representations as decoder prompts, reducing ambiguity and improving F1 scores. We evaluate our method on the \textit{MAESTRO-E} and \textit{CocoChorales-E} datasets by measuring the F1 score for each note category. Compared to the previous state of the art, \textit{LadderSym} more than doubles F1 for missed notes on \textit{MAESTRO-E} (26.8\%~$\rightarrow$~56.3\%) and improves extra note detection by 14.4 points (72.0\%~$\rightarrow$~86.4\%). Similar gains are observed on \textit{CocoChorales-E}. Furthermore, we also evaluate our models on real data we curated. This work introduces insights about comparison models that could inform sequence evaluation tasks for reinforcement learning, human skill assessment, and model evaluation.

PDF ICLR OpenReview Semantic Scholar

Cite

Text

Chou et al. "LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection." International Conference on Learning Representations, 2026.

Markdown

[Chou et al. "LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection." International Conference on Learning Representations, 2026.](https://mlanthology.org/iclr/2026/chou2026iclr-laddersym/)

BibTeX

@inproceedings{chou2026iclr-laddersym,
  title     = {{LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection}},
  author    = {Chou, Benjamin Shiue-Hal and Jajal, Purvish and Eliopoulos, Nicholas John and Davis, James C. and Thiruvathukal, George K and Yun, Kristen Yeon-Ji and Lu, Yung-Hsiang},
  booktitle = {International Conference on Learning Representations},
  year      = {2026},
  url       = {https://mlanthology.org/iclr/2026/chou2026iclr-laddersym/}
}