SSL-NL dataset

Creators
Publication date 29-05-2025
Description
SSL-NL is a curated dataset of Dutch speech recordings and accompanying forced alignments, designed to test the encoding of Dutch phonetic and lexical features in SSL speech representations while allowing for comparisons across different analysis methods.  It consists of two subsets from different domains: audiobook (MLS) and face-to-face conversations (IFADV). The MLS recordings were extracted from the Dutch part of Multilingual LibriSpeech, and the IFADV recordings were extracted from the IFA Dialog Video corpus and split by speaker turn. All audio recordings were downsampled to 16 kHz, and forced alignments were generated using the available transcripts and the WebMAUS API. The SSL-NL evaluation dataset was released as part of: de Heer Kloots, M., Mohebbi, H., Pouw, C., Shen, G., Zuidema, W., Bentum, M. (2025) What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training. Proc. Interspeech 2025, 256-260, doi: 10.21437/Interspeech.2025-1526 Analysis code accompanying the SSL-NL dataset (to replicate results from the Interspeech paper) is available at https://github.com/mdhk/SSL-NL-eval.
Publisher Zenodo
Organisations
  • Interfacultary Research - Institute for Logic, Language and Computation (ILLC)
Document type Dataset
Related publication What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
DOI https://doi.org/10.5281/zenodo.15548946
Permalink to this page
Back