Evaluating Whisper Model on Indonesian Educational Videos Transcription with Varying Audio Conditions

Fathia Sabrina, Fitria Nilamsari, Rinny Asasunnaja

Sari


The increasing use of video-based learning in digital education highlights the importance of accurate automatic speech recognition (ASR) systems to support accessibility, subtitle generation, and inclusive learning environments. This study evaluates the performance of the Whisper base ASR model on Indonesian educational videos with diverse audio characteristics and production conditions. Several categories of educational videos were analyzed, including classroom lectures, podcasts, interviews, and animated educational content. Audio recordings were converted into WAV format and evaluated using Word Error Rate (WER) and Character Error Rate (CER). Long-duration recordings were additionally segmented into approximately 30-minute chunks to analyze transcription consistency over time. The results showed that transcription performance varied across recording conditions, with WER values ranging from 18% to 47% and CER values ranging from 5% to 23%. The analysis also identified recurring substitution patterns influenced by phonetic similarity, conversational expressions, and culturally contextual phrases. The findings indicate that Whisper base provides reasonably effective transcription capability for Indonesian educational multimedia under realistic recording conditions.

Teks Lengkap:

PDF (English)

Referensi


M. Noetel et al., “Video Improves Learning in Higher Education: A Systematic Review,” Rev. Educ. Res., vol. 91, no. 2, pp. 204–236, Apr. 2021, doi: 10.3102/0034654321990713.

E. Navarrete, A. Nehring, S. Schanze, R. Ewerth, and A. Hoppe, “A Closer Look into Recent Video-based Learning Research: A Comprehensive Review of Video Characteristics, Tools, Technologies, and Learning Effectiveness,” Int. J. Artif. Intell. Educ., vol. 35, no. 4, pp. 1631–1694, Dec. 2025, doi: 10.1007/s40593-025-00481-x.

P. Khong, D. Holmes, B. Masoudian, G. C. Lund, and S. Garwood, “Lecture Capture, Transcripts, and Captioning in US Colleges of Osteopathic Medicine: Descriptive Cross-Sectional Survey,” Med. Sci. Educ., vol. 35, no. 2, pp. 625–628, Nov. 2024, doi: 10.1007/s40670-024-02224-4.

S. Malakul and I. Park, “The effects of using an auto-subtitle system in educational videos to facilitate learning for secondary school students: learning comprehension, cognitive load, and satisfaction,” Smart Learn. Environ., vol. 10, no. 1, p. 4, Jan. 2023, doi: 10.1186/s40561-023-00224-2.

H. Gonzalez et al., “Automatically Generated Summaries of Video Lectures May Enhance Students’ Learning Experience,” in Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Toronto, Canada: Association for Computational Linguistics, 2023, pp. 382–393. doi: 10.18653/v1/2023.bea-1.31.

A. Kotevski, “USING AUTOMATIC SPEECH RECOGNITION TO SUPPORT STUDENTS WITH DISABILITIES,” UKLO Proc., vol. 1, no. 1, pp. 187–193, May 2025, doi: 10.20544/AISC.1.1.25.P18.

Hasbullah Azis, E. P. Indreswari, and R. Wisudawanto, “PEMANFAATAN TEKNOLOGI AUTOMATIC SPEECH RECOGNITION DALAM MENCIPTAKAN PEMBELAJARAN INKLUSIF: IMPLEMENTASI, EFEKTIFITAS DAN TANTANGAN,” GANESHA J. Pengabdi. Masy., vol. 5, no. 1, pp. 27–33, Jan. 2025, doi: 10.36728/ganesha.v5i1.4171.

E. William and A. Zahra, “Speech Recognition Dengan Whisper Dalam Bahasa Indonesia,” Action Res. Lit., vol. 9, no. 2, pp. 386–397, Feb. 2025, doi: 10.46799/arl.v9i2.2573.

K. Kuhn, V. Kersken, B. Reuter, N. Egger, and G. Zimmermann, “Measuring the Accuracy of Automatic Speech Recognition Solutions,” ACM Trans. Access. Comput., vol. 16, no. 4, pp. 1–23, Dec. 2023, doi: 10.1145/3636513.

A. Adila, D. Lestari, A. Purwarianti, D. Tanaya, K. Azizah, and S. Sakti, “Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities,” in 2024 27th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA), Hsinchu City, Taiwan: IEEE, Oct. 2024, pp. 1–6. doi: 10.1109/O-COCOSDA64382.2024.10800336.

K. Zaman, M. Sah, C. Direkoglu, and M. Unoki, “A Survey of Audio Classification Using Deep Learning,” IEEE Access, vol. 11, pp. 106620–106649, 2023, doi: 10.1109/ACCESS.2023.3318015.

L. Fredianelli, F. Artuso, G. Pompei, G. Licitra, G. Iannace, and A. Akbaba, “Environmental Noise Dataset for Sound Event Classification and Detection,” Sci. Data, vol. 12, no. 1, p. 1712, Oct. 2025, doi: 10.1038/s41597-025-05991-w.

O. E. Tumewu, S. Saprudin, U. Sambiri, R. Achmad, M. H. Rahman, and F. Hamid, “The implementation of gamification with video media variations to improve students’ concept mastery in science learning,” JPPI J. Penelit. Pendidik. Indones., vol. 11, no. 4, pp. 72–81, Dec. 2025, doi: 10.29210/020256620.

S. Alley, “Do Higher Production Value Videos Lead to Improved Engagement and Learning Outcomes? A Field Experiment,” J. Educ. Online, vol. 22, no. 1, Jan. 2025, doi: 10.9743/JEO.2025.22.1.15.

S. Z. Pranida, R. A. Genadi, M. C. Airlangga, and S. Shehata, “ASR Under Noise: Exploring Robustness for Sundanese and Javanese,” in Proceedings of the 9th Widening NLP Workshop, Suzhou, China: Association for Computational Linguistics, 2025, pp. 87–99. doi: 10.18653/v1/2025.winlp-main.16.

M. Borsky, P. Mizera, P. Pollak, and J. Nouza, “Dithering techniques in automatic recognition of speech corrupted by MP3 compression: Analysis, solutions and experiments,” Speech Commun., vol. 86, pp. 75–84, Feb. 2017, doi: 10.1016/j.specom.2016.11.007.

OpenWhispr, “Whisper Model Sizes Explained,” OpenWhispr Blog. Accessed: May 24, 2026. [Online]. Available: https://openwhispr.com/blog/whisper-model-sizes-explained

D. Moser, N. Stanic, and M. Sariyar, “Benchmarking speech-to-text robustness in noisy emergency medical dialogues: an evaluation of models under realistic acoustic conditions,” JAMIA Open, vol. 8, no. 6, p. ooaf147, Nov. 2025, doi: 10.1093/jamiaopen/ooaf147.

T. D. K, J. James, D. P. Gopinath, and M. A. K, “Advocating Character Error Rate for Multilingual ASR Evaluation,” in Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico: Association for Computational Linguistics, 2025, pp. 4926–4935. doi: 10.18653/v1/2025.findings-naacl.277.




DOI: http://dx.doi.org/10.30811/jaise.v6i2.9201

Refbacks

  • Saat ini tidak ada refbacks.


Indexing :

Creative Commons License
Journal of Artificial Intelligence and Software Engineering (JAISE) licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.