OSCR

A standardized naturalistic audio stimulus dataset with unsupervised labeling.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 122–160 · score 0.71 · correctSpelling, stop words, MATLAB, built, autocorrection, voice
  2. [2] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 163–249 · score 0.65 · random seed, unique guesses, silhouette score, classify, algorithm, cluster
  3. [3] § Technical Validation › Text processing pipeline ↔ Functions/textAnalysis.m, lines 77–118 · score 0.55 · embedding space, GloVe, Twitter, words, clustered

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

MATLAB · 265 lines · 10 KB · MIT · 3 matches

  1. function [varargout] = textAnalysis(data,varargin)
  2. % textAnalysis Perform text cleaning and first-step clustering to compute a similarity matrix
  3. %
  4. % similarityAll = textAnalysis(data) cleans participant-generated text
  5. % responses and runs the first step of the two-stage semantic
  6. % categorization analysis. The function returns a SoundNum-by-SoundNum
  7. % similarity matrix describing how consistently two sounds are placed in
  8. % the same semantic cluster across iterations.
  9. %
  10. % [similarityAll, allSilChoice, silhouetteAvg] = textAnalysis(data, ...)
  11. % also returns the selected number of clusters for each iteration and the
  12. % silhouette scores for all tested cluster numbers.
  13. %
  14. % INPUTS
  15. %
  16. % data - A matrix of strings of size ParticipantNum-by-SoundNum.
  17. % Each entry contains the participant’s textual label or
  18. % guess for that sound. Empty or missing responses should be
  19. % represented as "".
  20. %
  21. % numiter - (Optional) Number of iterations for the first-step
  22. % clustering procedure. In each iteration, the function
  23. % embeds all unique cleaned guesses into a semantic space
  24. % and runs k-means clustering to identify the optimal
  25. % cluster solution.
  26. % Default: 1000
  27. %
  28. % clusterRange - (Optional) Two-element vector specifying the range of
  29. % cluster numbers to test during each iteration. The
  30. % function computes silhouette scores for each cluster
  31. % solution in this range and selects the cluster number
  32. % with the highest silhouette value.
  33. % Default: [5 25]
  34. %
  35. % randseed - (Optional) Random seed for reproducibility.
  36. % Default: 1
  37. %
  38. % stopwords - (Optional) List of stop words to remove from text
  39. % responses. These strings will be replaced by empty
  40. % entries before embedding.
  41. % Default:
  42. % [stopWords('Language','en'), "voice","noise","sound", ...
  43. % "noises","sounds","voices","answer"]
  44. %
  45. % OUTPUTS
  46. %
  47. % similarityAll - A SoundNum-by-SoundNum matrix. Element (i,j) represents
  48. % the normalized number of times sound i and sound j
  49. % were assigned to the same k-means cluster across all
  50. % iterations.
  51. %
  52. % allSilChoice - A numiter-by-1 vector containing the chosen number of
  53. % clusters for each iteration (i.e., the cluster count
  54. % with maximum silhouette score in that iteration).
  55. %
  56. % silhouetteAvg - A matrix storing the silhouette score curves for all
  57. % tested cluster numbers across all iterations.
  58. %
  59. % DESCRIPTION
  60. %
  61. % This function implements the first step of the semantic categorization
  62. % pipeline described in the study. All participant guesses are cleaned
  63. % (stop-word removal, deletion of nonsense entries, autocorrection)
  64. % and then embedded into a 200-dimensional GloVe semantic space. In each
  65. % iteration, unique embedded guesses are clustered using k-means across a
  66. % specified range of cluster numbers. The optimal cluster solution is
  67. % selected via silhouette analysis. The full set of guesses is then used
  68. % to assign each sound to a cluster for that iteration. Repeating this
  69. % procedure numiter times yields a consensus-based similarity matrix
  70. % measuring how often each pair of sounds co-clusters.
  71. %
  72. % This matrix serves as input to the second-stage clustering procedure
  73. % described in the main analysis pipeline.
  74. %
  75. % See also: kmeans, silhouette, stopWords
  76. numiter = 1000;
  77. clusterRange = [5,25];
  78. randseed = 1;
  79. if nargin >= 2
  80. if ~isempty(varargin{1})
  81. numiter = varargin{1};
  82. end
  83. end
  84. disp(['Chosen number of iterations: ' num2str(numiter)])
  85. if nargin >= 3
  86. if ~isempty(varargin{2})
  87. clusterRange = varargin{2};
  88. end
  89. end
  90. if ~numel(clusterRange) == 2
  91. error('Range of clusters argument must contain a lower and upper bound')
  92. end
  93. if clusterRange(1) > clusterRange(2)
  94. error('Cluster range lower bound should be the first entry')
  95. end
  96. if nargin >= 4
  97. if ~isempty(varargin{3})
  98. randseed = varargin{3};
  99. end
  100. end
  101. disp('Loading embedding space')
  102. % Loading the glove twitter word embedding
  103. try
  104. emb = readWordEmbedding('glove.twitter.27B.200d.txt');
  105. catch ME
  106. error(['Failed to load the GloVe embedding file. ', ...
  107. 'Please ensure that the file exists and that the path is correct.\n', ...
  108. 'Specified path: %s\nOriginal error: %s'], 'glove.twitter.27B.200d.txt', ME.message);
  109. end
  110. % Define stopwords list (MATLAB's built-in stopwords list)
  111. stopwords = stopWords('Language','en');
  112. % Adding our own list of stop words
  113. % Perhaps add an option for the user to enter their own stop words here?
  114. stopwords = [stopwords "voice" "noise" "sound" "noises" "sounds" "voices" "answer"];
  115. if nargin >= 5
  116. stopwords = [varargin{5}];
  117. end
  118. % Tokenize the data
  119. documents = tokenizedDocument(data);
  120. % Turn every word into lower case
  121. documents = lower(documents);
  122. % Apply MATLAB's autocorrect function
  123. documents = correctSpelling(documents);
  124. % Remove stop words
  125. documents = removeWords(documents, stopwords);
  126. % Initialize an empty cell to store embeddings
  127. docVectors = cell(numel(documents), 1);
  128. disp('Pre-processing text data')
  129. % Pre-processing of the text data
  130. for i = 1:numel(documents)
  131. if strcmp(data(i), "No answer") % Answers that were left empty were automatically filled with No answer
  132. docVectors{i} = single(zeros(1, emb.Dimension)); % Handle missing data
  133. continue
  134. end
  135. words = string(documents(i));
  136. wordVectors = word2vec(emb, words); % Retrieve word vectors
  137. wordVectors = wordVectors(~any(isnan(wordVectors), 2), :); % Remove missing words
  138. if ~isempty(wordVectors)
  139. % Average the word vectors for each document
  140. % This is a quick way to handle guesses with multiple words
  141. docVectors{i} = mean(wordVectors, 1);
  142. else
  143. docVectors{i} = single(zeros(1, emb.Dimension)); % Handle documents with no match
  144. end
  145. end
  146. % Selecting unique and non-empty guesses words only
  147. docMatrix = cell2mat(docVectors); % Convert to matrix format
  148. % Selecting the non-empty guesses first:
  149. % Turning matrix to logical with 1 when any value in the matrix is not empty
  150. nonEmptyData = docMatrix ~= 0;
  151. % Turning matrix into vector of 1 iff all values in that row is empty (i.e. No Answer)
  152. emptyVec = zeros([1,length(nonEmptyData)]);
  153. for i = 1:length(nonEmptyData)
  154. emptyVec(i) = all(nonEmptyData(i,:) == 0);
  155. end
  156. % Only 1 iff all values in that row is not empty
  157. nonEmptyDataRow = ~emptyVec;
  158. % Selecting non-empty guess
  159. fullMat = docMatrix(nonEmptyDataRow,:);
  160. % Saving the index of the empty guesses
  161. emptyIdx = find(emptyVec);
  162. % Creating a function that will insert these empty guesses back
  163. % This will be used after running the k-means
  164. insert = @(n,x,a) [x(1:a-1), n, x(a:end)];
  165. % Creating the matrix with the unique guesses only
  166. [X, ~, ic] = unique(fullMat , 'rows');
  167. % Performing the first step of the algorithm
  168. disp('Running the first step of the k-means')
  169. sumSim = zeros([length(data) length(data)]);
  170. silhouetteAvg = zeros([numiter,clusterRange(2)]);
  171. for itr = 1:numiter
  172. % Perform K-means clustering
  173. disp(['Iteration Number: ' num2str(itr)])
  174. % Choosing the number of clusters for this iteration
  175. avg = zeros([1,clusterRange(2)]);
  176. % Running the algorithm multiple times to check which number of
  177. % clusters results in the maximum number of clusters
  178. for r = clusterRange(1):clusterRange(2)
  179. rng(itr * randseed) % Set random seed
  180. idx = kmeans(X, r, 'Replicates', 5,'Distance','cosine');
  181. silhouetteScores = (silhouette(gather(X), gather(idx)));
  182. avg(r) = mean(silhouetteScores);
  183. end
  184. % Save the average silhouette score for every run of the k-means
  185. silhouetteAvg(itr,:) = avg;
  186. % Find the run that produced the highest avg silhouette score
  187. numClusters = find(avg == max(avg(clusterRange(1):clusterRange(2))), 1);
  188. disp(['Chosen Number of Clusters: ' num2str(numClusters)])
  189. % Run the k-means again with the chosen number of clusters
  190. rng(itr * randseed)
  191. idxMain = kmeans(X, numClusters, 'Replicates', 5, 'Distance','cosine');
  192. % Insert the non-unique guesses
  193. OGIDX = zeros([1,length(ic)]);
  194. for ii = 1:length(ic)
  195. OGIDX(ii) = idxMain(ic(ii));
  196. end
  197. idxMain = OGIDX;
  198. % Insert the 'No Answer' guesses as zeros
  199. for ii = 1:length(find(emptyVec))
  200. idxMain = insert(0,idxMain,emptyIdx(ii));
  201. end
  202. % Reshape the classification of the guesses into the original format
  203. % (Participant x Guess)
  204. idxMain = idxMain';
  205. idxReshaped = reshape(idxMain,size(data,1),size(data,2));
  206. % Count the number of times a sound had a guess in a certain cluster
  207. sumsOfCluster = zeros([numClusters, size(data,2)]);
  208. for i = 1:numClusters
  209. for ii = 1:size(data,2)
  210. sumsOfCluster(i,ii) = gather(sum(idxReshaped(:,ii) == i));
  211. end
  212. end
  213. % Find the cluster that the guesses of each of the sounds was
  214. % classified to the most and take that as the classification for that
  215. % sounds
  216. [~, idxSum] = max(sumsOfCluster, [], 1);
  217. % Create a sound x sound matrix which contains information about which
  218. % sound was classified with another one. For example: if the value at
  219. % row 2 and column 3 is 1, then sound 2 and sound 3 were classified in
  220. % the same category
  221. similarity = zeros(size(data,2));
  222. for i = 1:length(idxSum)
  223. similarity(i,:) = idxSum == idxSum(i);
  224. end
  225. % Save that matrix for each iteration
  226. sumSim = sumSim + similarity;
  227. end
  228. allSilChoice = zeros([1,numiter]);
  229. % Save the chosen number of clusters for each iteration
  230. for i = 1:length(silhouetteAvg)
  231. allSilChoice(i) = find(silhouetteAvg(i,:) == max(silhouetteAvg(i,clusterRange(1):clusterRange(2))), 1);
  232. end
  233. % Create a similarity matrix by dividing the resulting matrix from last
  234. % step with the number of iterations. This is easier to use with
  235. % classification algorithms
  236. similarityAll = sumSim/numiter;
  237. varargout{1} = similarityAll;
  238. varargout{2} = allSilChoice;
  239. varargout{3} = silhouetteAvg;
  240. end

textAnalysis.m at commit 7955dde, under MIT · at the source

Overview

Authors: Anas Al-Naji1,2, Ricarda I. Schubotz1,2, Anoushiravan Zahedi1,2
  1. Institute of Psychology, University of Muenster,Fliednerstr. 21, 48149 Muenster, Germany
  2. Otto Creutzfeldt-Center for Cognitive and Behavioral Neuroscience, University of Muenster,Fliednerstr. 21, 48149 Muenster, Germany
Institutions: University of Münster (Germany)
Journal: Scientific data, volume 13, issue 1, article 1134
Dates: received 1 December 2025; accepted 29 July 2026; published online 5 August 2026
Type: Data paper · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41597-026-08034-0 · PMID 42557253 · PMCID PMC13443867 · OpenAlex W7172480138
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), methods / tools (subfield)
MeSH: Acoustic Stimulation*, Unsupervised Machine Learning*, Cognitive Neuroscience, Emotions, Humans (* major topic)
Journal subjects: Data Descriptor
Citations: not cited yet (Europe PMC); 34 references in the paper

Abstract

This study presents a standardized naturalistic audio stimulus dataset designed for use in trial-wise cognitive neuroscience, neuroimaging, and behavioral research. The dataset provides short, recognizable auditory stimuli that are normed for emotional valence and startlingness. To create such a dataset, the current study collected 291 audio files from a range of sources and standardized them to a duration of 1.5 s. A final sample of 361 participants rated the audio clips on emotional valence, startlingness, and recognizability, and subsequently freely described the audios by typing what they believed the sound to be. The text responses of the participants were embedded and clustered using an unsupervised machine-learning algorithm to derive a participant-grounded organization of auditory object categories. The results indicate that the audio clips were generally recognizable, while emotional valence and startlingness ratings varied across stimuli.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

ajn3333/Sound-Standardization

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 7955ddee6b30262dafbeb6c596540e4a48ce41cd, 10 August 2026
Languages: MATLAB (3)
Size: 11 files, 3 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
5 files

Code availability

The code for the project can also be found on the following GitHub link: https://github.com/ajn3333/Sound-Standardization.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 3 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The data and the code used for the analysis are available from the Open Science Framework22.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 5 MeSH terms, 30 references.

Cite

This paper

Al-Naji, A., Schubotz, R. I., & Zahedi, A. (2026). A standardized naturalistic audio stimulus dataset with unsupervised labeling. Scientific data, 13(1), 1134. https://doi.org/10.1038/s41597-026-08034-0

BibTeX

@article{alnaji2026standardized,
author = {Al-Naji, Anas and Schubotz, Ricarda I. and Zahedi, Anoushiravan},
title = {{A standardized naturalistic audio stimulus dataset with unsupervised labeling}},
journal = {Scientific data},
year = {2026},
month = aug,
volume = {13},
number = {1},
pages = {1134},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/s41597-026-08034-0},
url = {https://doi.org/10.1038/s41597-026-08034-0},
pmid = {42557253},
pmcid = {PMC13443867}
}

RIS

TY - JOUR
AU - Al-Naji, Anas
AU - Schubotz, Ricarda I.
AU - Zahedi, Anoushiravan
TI - A standardized naturalistic audio stimulus dataset with unsupervised labeling
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/08/05
VL - 13
IS - 1
SP - 1134
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/s41597-026-08034-0
UR - https://doi.org/10.1038/s41597-026-08034-0
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41597-026-08034-0",
"type": "article-journal",
"title": "A standardized naturalistic audio stimulus dataset with unsupervised labeling",
"container-title": "Scientific data",
"author": [
{
"family": "Al-Naji",
"given": "Anas"
},
{
"family": "Schubotz",
"given": "Ricarda I."
},
{
"family": "Zahedi",
"given": "Anoushiravan"
}
],
"container-title-short": "Sci Data",
"volume": "13",
"issue": "1",
"page": "1134",
"DOI": "10.1038/s41597-026-08034-0",
"PMID": "42557253",
"PMCID": "PMC13443867",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41597-026-08034-0",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
5
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41598-026-48067-4
Semantic processing and individual suggestibility modulate motor preparation and perceived distance for looming sounds entering the peripersonal space.
Journal: Scientific reports
In common: 3 references
[2] doi:10.1162/imag.a.1331 [code]
Beyond the canonical HRF: Flexible temporal modeling reveals unconstrained BOLD profiles during naturalistic viewing.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: 3 references
[3] doi:10.3389/fnins.2026.1804069 [code]
Neural and autonomic regulation during brief mindfulness and relaxation interventions in clinical populations: a multimodal MEG study protocol.
Journal: Frontiers in neuroscience
In common: 3 references
[4] doi:10.1186/s12915-026-02630-7 [code]
Phasic modulation of attentional rhythmic sampling according to task demands.
Journal: BMC biology
In common: Statistics and Machine Learning Toolbox, 2 references
[5] doi:10.1007/s00429-026-03111-x [code]
Direction selectivity in naturalistic action observation: distributed representations across the action observation network.
Journal: Brain structure & function
In common: Statistics and Machine Learning Toolbox, 2 references
[6] doi:10.1523/eneuro.0346-25.2026 [code]
Effects of TMS on the Decoding and Electrophysiology of Priority in Working Memory.
Journal: eNeuro
In common: Statistics and Machine Learning Toolbox, 2 references
[7] doi:10.1016/j.ynirp.2026.100363
Narrative coherence shapes functional connectivity in default mode and frontoparietal networks.
Journal: Neuroimage. Reports
In common: 2 references
[8] doi:10.1038/s41597-026-06616-6 [code]
Sustained Attention Task (gradCPT) Dataset using simultaneous EEG-fMRI and DTI.
Journal: Scientific data
In common: methods / tools, 2 references
[9] doi:10.1016/j.neuroimage.2026.122171 [code]
A conserved node degree-based backbone and flexible hub organization of brain connectome during naturalistic movie watching.
Journal: NeuroImage
In common: Statistics and Machine Learning Toolbox, methods / tools, 1 reference
[10] doi:10.1038/s41597-026-07676-4 [code]
A neuroimaging dataset combining movie-watching, eye-tracking, sensorimotor mapping, and cognitive tasks.
Journal: Scientific data
In common: Statistics and Machine Learning Toolbox, methods / tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.