Abstract
This article presents a deep-learning-based Automatic Music Transcription (AMT) framework for the Begena, a traditional Ethiopian ten-stringed musical instrument. Manual transcription is time-consuming, inconsistent, and costly; moreover, most existing AMT systems have been developed primarily for Western instruments and therefore generalise poorly to Begena recordings. In addition, existing approaches often lack explicit scale identification and must accommodate substantial differences between Western staff notation and the Begena's numeric notation. Accordingly, the proposed framework integrates scale/kiñit classification with AMT to strengthen musical analysis and transcription, and further evaluates the extent to which predicted scale labels enable scale-aware refinement of AMT predictions. Specifically, for scale/kiñit classification, Constant-Q Transform (CQT), Mel-spectrogram, and Mel-frequency cepstral coefficient (MFCC) representations are assessed using multiple deep learning architectures, with MFCC and CQT-based models yielding the best classification performance. In parallel, three transcription models are developed and compared: Convolutional Neural Networks (CNN), Convolutional Recurrent Neural Networks (CRNN), and Generative Adversarial Networks (GAN). The GAN-based approach attains the highest transcription accuracy among the evaluated methods. Nevertheless, although both components produce competitive performance when evaluated independently, using predicted scale labels for scale-aware refinement yields only minimal additional improvement, suggesting that residual errors are driven primarily by note-event detection and temporal localisation rather than scale identification alone. On the AMT side, remaining challenges include occasional note omissions and note-duration expansion errors, indicating the need for further refinement. Taken together, these results provide a specialised AMT solution for Begena songs; underscore the promise of unifying scale/kiñit classification and transcription within a computational music analysis pipeline for under-represented musical traditions; and establish a foundation both for the digital preservation of Ethiopian musical heritage and for future computational creativity applications – including symbolic music generation and interactive composition.
Keywords
Automatic Music Transcription (AMT), Begena, scale/kiñit classification, deep learning, Convolutional Neural Network (CNN), Convolutional Recurrent Neural Network (CRNN), Generative Adversarial Network (GAN), Constant-Q Transform (CQT), Mel scale, Mel-Frequency Cepstral Coefficients (MFCC)
How to Cite
Yemarshet, B. G. & Asfaw, T. T., (2026) “Automatic Music Transcription for Ethiopian Begena using Deep Learning and Scale/Kiñit-Aware Note Correction”, Journal of Creative Music Systems 10(1). doi: https://doi.org/10.5920/jcms.1553
147
Views
54
Downloads