A standardized naturalistic audio stimulus dataset with unsupervised labeling.
The 3 matches
- [1] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 122–160 · score 0.71 · correctSpelling, stop words, MATLAB, built, autocorrection, voice
- [2] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 163–249 · score 0.65 · random seed, unique guesses, silhouette score, classify, algorithm, cluster
- [3] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 77–118 · score 0.55 · embedding space, GloVe, Twitter, words, clustered
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
MATLAB · 265 lines · 10 KB · MIT · 3 matches
- function [varargout] = textAnalysis(data,varargin)
- % textAnalysis Perform text cleaning and first-step clustering to compute a similarity matrix
- %
- % similarityAll = textAnalysis(data) cleans participant-generated text
- % responses and runs the first step of the two-stage semantic
- % categorization analysis. The function returns a SoundNum-by-SoundNum
- % similarity matrix describing how consistently two sounds are placed in
- % the same semantic cluster across iterations.
- %
- % [similarityAll, allSilChoice, silhouetteAvg] = textAnalysis(data, ...)
- % also returns the selected number of clusters for each iteration and the
- % silhouette scores for all tested cluster numbers.
- %
- % INPUTS
- %
- % data - A matrix of strings of size ParticipantNum-by-SoundNum.
- % Each entry contains the participant’s textual label or
- % guess for that sound. Empty or missing responses should be
- % represented as "".
- %
- % numiter - (Optional) Number of iterations for the first-step
- % clustering procedure. In each iteration, the function
- % embeds all unique cleaned guesses into a semantic space
- % and runs k-means clustering to identify the optimal
- % cluster solution.
- % Default: 1000
- %
- % clusterRange - (Optional) Two-element vector specifying the range of
- % cluster numbers to test during each iteration. The
- % function computes silhouette scores for each cluster
- % solution in this range and selects the cluster number
- % with the highest silhouette value.
- % Default: [5 25]
- %
- % randseed - (Optional) Random seed for reproducibility.
- % Default: 1
- %
- % stopwords - (Optional) List of stop words to remove from text
- % responses. These strings will be replaced by empty
- % entries before embedding.
- % Default:
- % [stopWords('Language','en'), "voice","noise","sound", ...
- % "noises","sounds","voices","answer"]
- %
- % OUTPUTS
- %
- % similarityAll - A SoundNum-by-SoundNum matrix. Element (i,j) represents
- % the normalized number of times sound i and sound j
- % were assigned to the same k-means cluster across all
- % iterations.
- %
- % allSilChoice - A numiter-by-1 vector containing the chosen number of
- % clusters for each iteration (i.e., the cluster count
- % with maximum silhouette score in that iteration).
- %
- % silhouetteAvg - A matrix storing the silhouette score curves for all
- % tested cluster numbers across all iterations.
- %
- % DESCRIPTION
- %
- % This function implements the first step of the semantic categorization
- % pipeline described in the study. All participant guesses are cleaned
- % (stop-word removal, deletion of nonsense entries, autocorrection)
- % and then embedded into a 200-dimensional GloVe semantic space. In each
- % iteration, unique embedded guesses are clustered using k-means across a
- % specified range of cluster numbers. The optimal cluster solution is
- % selected via silhouette analysis. The full set of guesses is then used
- % to assign each sound to a cluster for that iteration. Repeating this
- % procedure numiter times yields a consensus-based similarity matrix
- % measuring how often each pair of sounds co-clusters.
- %
- % This matrix serves as input to the second-stage clustering procedure
- % described in the main analysis pipeline.
- %
- % See also: kmeans, silhouette, stopWords
- numiter = 1000;
- clusterRange = [5,25];
- randseed = 1;
- if nargin >= 2
- if ~isempty(varargin{1})
- numiter = varargin{1};
- end
- end
- disp(['Chosen number of iterations: ' num2str(numiter)])
- if nargin >= 3
- if ~isempty(varargin{2})
- clusterRange = varargin{2};
- end
- end
- if ~numel(clusterRange) == 2
- error('Range of clusters argument must contain a lower and upper bound')
- end
- if clusterRange(1) > clusterRange(2)
- error('Cluster range lower bound should be the first entry')
- end
- if nargin >= 4
- if ~isempty(varargin{3})
- randseed = varargin{3};
- end
- end
- disp('Loading embedding space')
- % Loading the glove twitter word embedding
- try
- emb = readWordEmbedding('glove.twitter.27B.200d.txt');
- catch ME
- error(['Failed to load the GloVe embedding file. ', ...
- 'Please ensure that the file exists and that the path is correct.\n', ...
- 'Specified path: %s\nOriginal error: %s'], 'glove.twitter.27B.200d.txt', ME.message);
- end
- % Define stopwords list (MATLAB's built-in stopwords list)
- stopwords = stopWords('Language','en');
- % Adding our own list of stop words
- % Perhaps add an option for the user to enter their own stop words here?
- stopwords = [stopwords "voice" "noise" "sound" "noises" "sounds" "voices" "answer"];
- if nargin >= 5
- stopwords = [varargin{5}];
- end
- % Tokenize the data
- documents = tokenizedDocument(data);
- % Turn every word into lower case
- documents = lower(documents);
- % Apply MATLAB's autocorrect function
- documents = correctSpelling(documents);
- % Remove stop words
- documents = removeWords(documents, stopwords);
- % Initialize an empty cell to store embeddings
- docVectors = cell(numel(documents), 1);
- disp('Pre-processing text data')
- % Pre-processing of the text data
- for i = 1:numel(documents)
- if strcmp(data(i), "No answer") % Answers that were left empty were automatically filled with No answer
- docVectors{i} = single(zeros(1, emb.Dimension)); % Handle missing data
- continue
- end
- words = string(documents(i));
- wordVectors = word2vec(emb, words); % Retrieve word vectors
- wordVectors = wordVectors(~any(isnan(wordVectors), 2), :); % Remove missing words
- if ~isempty(wordVectors)
- % Average the word vectors for each document
- % This is a quick way to handle guesses with multiple words
- docVectors{i} = mean(wordVectors, 1);
- else
- docVectors{i} = single(zeros(1, emb.Dimension)); % Handle documents with no match
- end
- end
- % Selecting unique and non-empty guesses words only
- docMatrix = cell2mat(docVectors); % Convert to matrix format
- % Selecting the non-empty guesses first:
- % Turning matrix to logical with 1 when any value in the matrix is not empty
- nonEmptyData = docMatrix ~= 0;
- % Turning matrix into vector of 1 iff all values in that row is empty (i.e. No Answer)
- emptyVec = zeros([1,length(nonEmptyData)]);
- for i = 1:length(nonEmptyData)
- emptyVec(i) = all(nonEmptyData(i,:) == 0);
- end
- % Only 1 iff all values in that row is not empty
- nonEmptyDataRow = ~emptyVec;
- % Selecting non-empty guess
- fullMat = docMatrix(nonEmptyDataRow,:);
- % Saving the index of the empty guesses
- emptyIdx = find(emptyVec);
- % Creating a function that will insert these empty guesses back
- % This will be used after running the k-means
- insert = @(n,x,a) [x(1:a-1), n, x(a:end)];
- % Creating the matrix with the unique guesses only
- [X, ~, ic] = unique(fullMat , 'rows');
- % Performing the first step of the algorithm
- disp('Running the first step of the k-means')
- sumSim = zeros([length(data) length(data)]);
- silhouetteAvg = zeros([numiter,clusterRange(2)]);
- for itr = 1:numiter
- % Perform K-means clustering
- disp(['Iteration Number: ' num2str(itr)])
- % Choosing the number of clusters for this iteration
- avg = zeros([1,clusterRange(2)]);
- % Running the algorithm multiple times to check which number of
- % clusters results in the maximum number of clusters
- for r = clusterRange(1):clusterRange(2)
- rng(itr * randseed) % Set random seed
- idx = kmeans(X, r, 'Replicates', 5,'Distance','cosine');
- silhouetteScores = (silhouette(gather(X), gather(idx)));
- avg(r) = mean(silhouetteScores);
- end
- % Save the average silhouette score for every run of the k-means
- silhouetteAvg(itr,:) = avg;
- % Find the run that produced the highest avg silhouette score
- numClusters = find(avg == max(avg(clusterRange(1):clusterRange(2))), 1);
- disp(['Chosen Number of Clusters: ' num2str(numClusters)])
- % Run the k-means again with the chosen number of clusters
- rng(itr * randseed)
- idxMain = kmeans(X, numClusters, 'Replicates', 5, 'Distance','cosine');
- % Insert the non-unique guesses
- OGIDX = zeros([1,length(ic)]);
- for ii = 1:length(ic)
- OGIDX(ii) = idxMain(ic(ii));
- end
- idxMain = OGIDX;
- % Insert the 'No Answer' guesses as zeros
- for ii = 1:length(find(emptyVec))
- idxMain = insert(0,idxMain,emptyIdx(ii));
- end
- % Reshape the classification of the guesses into the original format
- % (Participant x Guess)
- idxMain = idxMain';
- idxReshaped = reshape(idxMain,size(data,1),size(data,2));
- % Count the number of times a sound had a guess in a certain cluster
- sumsOfCluster = zeros([numClusters, size(data,2)]);
- for i = 1:numClusters
- for ii = 1:size(data,2)
- sumsOfCluster(i,ii) = gather(sum(idxReshaped(:,ii) == i));
- end
- end
- % Find the cluster that the guesses of each of the sounds was
- % classified to the most and take that as the classification for that
- % sounds
- [~, idxSum] = max(sumsOfCluster, [], 1);
- % Create a sound x sound matrix which contains information about which
- % sound was classified with another one. For example: if the value at
- % row 2 and column 3 is 1, then sound 2 and sound 3 were classified in
- % the same category
- similarity = zeros(size(data,2));
- for i = 1:length(idxSum)
- similarity(i,:) = idxSum == idxSum(i);
- end
- % Save that matrix for each iteration
- sumSim = sumSim + similarity;
- end
- allSilChoice = zeros([1,numiter]);
- % Save the chosen number of clusters for each iteration
- for i = 1:length(silhouetteAvg)
- allSilChoice(i) = find(silhouetteAvg(i,:) == max(silhouetteAvg(i,clusterRange(1):clusterRange(2))), 1);
- end
- % Create a similarity matrix by dividing the resulting matrix from last
- % step with the number of iterations. This is easier to use with
- % classification algorithms
- similarityAll = sumSim/numiter;
- varargout{1} = similarityAll;
- varargout{2} = allSilChoice;
- varargout{3} = silhouetteAvg;
- end
textAnalysis.m at commit 7955dde, under MIT · at the source
Overview
- Institute of Psychology, University of Muenster,Fliednerstr. 21, 48149 Muenster, Germany
- Otto Creutzfeldt-Center for Cognitive and Behavioral Neuroscience, University of Muenster,Fliednerstr. 21, 48149 Muenster, Germany
Abstract
This study presents a standardized naturalistic audio stimulus dataset designed for use in trial-wise cognitive neuroscience, neuroimaging, and behavioral research. The dataset provides short, recognizable auditory stimuli that are normed for emotional valence and startlingness. To create such a dataset, the current study collected 291 audio files from a range of sources and standardized them to a duration of 1.5 s. A final sample of 361 participants rated the audio clips on emotional valence, startlingness, and recognizability, and subsequently freely described the audios by typing what they believed the sound to be. The text responses of the participants were embedded and clustered using an unsupervised machine-learning algorithm to derive a participant-grounded organization of auditory object categories. The results indicate that the audio clips were generally recognizable, while emotional valence and startlingness ratings varied across stimuli.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.
ajn3333/Sound-Standardization
7955ddee6b30262dafbeb6c596540e4a48ce41cd, 10 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
5 files
- Functions/
ICC2K.m , MATLAB, 35 lines - Functions/
textAnalysis.m , MATLAB, 265 lines, 3 matches - main.m, MATLAB, 157 lines
- LICENSE, License, 21 lines
- README.md, Text, 121 lines
Code availability
The code for the project can also be found on the following GitHub link: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 3 scripts, each with its path and the digest of its content;
- 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
Data availability
The data and the code used for the analysis are available from the Open Science Framework22.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 5 MeSH terms, 30 references.
Cite
This paper
Al-Naji, A., Schubotz, R. I., & Zahedi, A. (2026). A standardized naturalistic audio stimulus dataset with unsupervised labeling. Scientific data, 13(1), 1134. https://
BibTeX
@article{alnaji2026stand
author = {Al-Naji, Anas and Schubotz, Ricarda I. and Zahedi, Anoushiravan},
title = {{A standardized naturalistic audio stimulus dataset with unsupervised labeling}},
journal = {Scientific data},
year = {2026},
month = aug,
volume = {13},
number = {1},
pages = {1134},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/
url = {https://
pmid = {42557253},
pmcid = {PMC13443867}
}
RIS
TY - JOUR
AU - Al-Naji, Anas
AU - Schubotz, Ricarda I.
AU - Zahedi, Anoushiravan
TI - A standardized naturalistic audio stimulus dataset with unsupervised labeling
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/
VL - 13
IS - 1
SP - 1134
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "A standardized naturalistic audio stimulus dataset with unsupervised labeling",
"container-title": "Scientific data",
"author": [
{
"family": "Al-Naji",
"given": "Anas"
},
{
"family": "Schubotz",
"given": "Ricarda I."
},
{
"family": "Zahedi",
"given": "Anoushiravan"
}
],
"container-title-short":
"volume": "13",
"issue": "1",
"page": "1134",
"DOI": "10.1038/
"PMID": "42557253",
"PMCID": "PMC13443867",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
5
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41598-026-48067-4
- Semantic processing and individual suggestibility modulate motor preparation and perceived distance for looming sounds entering the peripersonal space.Journal: Scientific reportsIn common: 3 references
- [2] doi:10.1162/imag.a.1331 [code]
- Beyond the canonical HRF: Flexible temporal modeling reveals unconstrained BOLD profiles during naturalistic viewing.Journal: Imaging neuroscience (Cambridge, Mass.)In common: 3 references
- [3] doi:10.3389/fnins.2026.1804069 [code]
- Neural and autonomic regulation during brief mindfulness and relaxation interventions in clinical populations: a multimodal MEG study protocol.Journal: Frontiers in neuroscienceIn common: 3 references
- [4] doi:10.1186/s12915-026-02630-7 [code]
- Phasic modulation of attentional rhythmic sampling according to task demands.Journal: BMC biologyIn common: Statistics and Machine Learning Toolbox, 2 references
- [5] doi:10.1007/s00429-026-03111-x [code]
- Direction selectivity in naturalistic action observation: distributed representations across the action observation network.Journal: Brain structure & functionIn common: Statistics and Machine Learning Toolbox, 2 references
- [6] doi:10.1523/eneuro.0346-25.2026 [code]
- Effects of TMS on the Decoding and Electrophysiology of Priority in Working Memory.Journal: eNeuroIn common: Statistics and Machine Learning Toolbox, 2 references
- [7] doi:10.1016/j.ynirp.2026.100363
- Narrative coherence shapes functional connectivity in default mode and frontoparietal networks.Journal: Neuroimage. ReportsIn common: 2 references
- [8] doi:10.1038/s41597-026-06616-6 [code]
- Sustained Attention Task (gradCPT) Dataset using simultaneous EEG-fMRI and DTI.Journal: Scientific dataIn common: methods / tools, 2 references
- [9] doi:10.1016/j.neuroimage.2026.122171 [code]
- A conserved node degree-based backbone and flexible hub organization of brain connectome during naturalistic movie watching.Journal: NeuroImageIn common: Statistics and Machine Learning Toolbox, methods / tools, 1 reference
- [10] doi:10.1038/s41597-026-07676-4 [code]
- A neuroimaging dataset combining movie-watching, eye-tracking, sensorimotor mapping, and cognitive tasks.Journal: Scientific dataIn common: Statistics and Machine Learning Toolbox, methods / tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 3 scripts, and 3 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:93ac1c412e28f993…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
