OSCR

SpeechDETECT: an explainable automated speech processing pipeline for early detection of neurological and health changes.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches
  1. [1] § Methodology › Overview of the SpeechDETECT architecture › Component 2.2. cepstral coefficients and spectral features ↔ speechdetect/AcousticFeatureExtractor.py, lines 136–214 · score 0.78 · Spectral Centroid, Cepstral Coefficients, Spectral features, LTAS, Mel, spectrum
  2. [2] § Methodology › Overview of the SpeechDETECT architecture › Component 2.4. loudness and intensity ↔ speechdetect/AcousticFeatureExtractor.py, lines 306–353 · score 0.70 · Sound Pressure, squared amplitude, SPL, RMS, loudness, Root
  3. [3] § Methodology › Overview of the SpeechDETECT architecture › Component 3- computing acoustic and temporal features ↔ speechdetect/acoustic_features/speech_fluency_and_speech_production_dynamics.py, lines 407–421 · score 0.68 · syllabic interval duration, production dynamics, syllable, phoneme, word, fluency
  4. [4] § Methodology › Overview of the SpeechDETECT architecture › Component 2.3. voice quality ↔ speechdetect/acoustic_features/voice_quality.py, lines 244–335 · score 0.65 · Harmonic Ratio, Voice Quality, Noise, NHR, consecutive, sound
  5. [5] § Methodology › Overview of the SpeechDETECT architecture › Component 2.2. cepstral coefficients and spectral features ↔ speechdetect/acoustic_features/cepstral_coefficients_and_spectral_features.py, lines 36–55 · score 0.59 · Cepstral Coefficients, Spectral features, LTAS, spectrum
  6. [6] § Methodology › Overview of the SpeechDETECT architecture › Component 2.8. speech production dynamics ↔ speechdetect/acoustic_features/speech_fluency_and_speech_production_dynamics.py, lines 588–602 · score 0.59 · Relative Sentence Duration, Speech production dynamics, fluency
  7. [7] § Methodology › Overview of the SpeechDETECT architecture › Component 2–voice analysis framework: characterizing vocal traits and temporal dynamics ↔ speechdetect/acoustic_features/speech_fluency_and_speech_production_dynamics.py, lines 24–27 · score 0.58 · speech production dynamics, speech fluency
  8. [8] § Methodology › Overview of the SpeechDETECT architecture ↔ speechdetect/AcousticFeatureExtractor.py, lines 1185–1310 · score 0.57 · fundamental frequency, voice quality, dimensions, rhythmic, fluency, spectral
  9. [9] § Methodology › Overview of the SpeechDETECT architecture › Component 2–voice analysis framework: characterizing vocal traits and temporal dynamics ↔ speechdetect/AcousticFeatureExtractor.py, lines 1419–1507 · score 0.56 · broader categories, speech fluency, map, voice
  10. [10] § Methodology › Overview of the SpeechDETECT architecture › Component 2.5. voice signal complexity ↔ speechdetect/acoustic_features/complexity.py, lines 127–145 · score 0.55 · Multiscale Permutation Entropy, MPE
  11. [11] § Methodology › Overview of the SpeechDETECT architecture › Component 2.5. voice signal complexity ↔ speechdetect/AcousticFeatureExtractor.py, lines 1185–1310 · score 0.54 · Fractal Dimension, Higuchi, HFD, modulate, amplitude, signal
  12. [12] § Methodology › Overview of the SpeechDETECT architecture › Component 2.4. loudness and intensity ↔ speechdetect/acoustic_features/loudness_and_intensity.py, lines 104–134 · score 0.53 · Sound Pressure, SPL, RMS, loudness, Intensity, Amplitude
  13. [13] § Methodology › Overview of the SpeechDETECT architecture › Component 3- computing acoustic and temporal features ↔ speechdetect/AcousticFeatureExtractor.py, lines 136–214 · score 0.51 · Cepstral Coefficients, Spectral Features, spectrogram, ms
  14. [14] § Results › SHAP values visualization (on Pitt corpus) ↔ speechdetect/AcousticFeatureExtractor.py, lines 355–395 · score 0.51 · pause ratio, pause duration, articulatory, syllable, Error, speech

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,507 lines · 69 KB · no license · 7 matches

  1. """
  2. AcousticFeatureExtractor - Main interface for extracting acoustic features from speech recordings
  3. """
  4. import torch
  5. import os
  6. import numpy as np
  7. import inspect
  8. import librosa
  9. import logging
  10. import csv
  11. from typing import List, Dict, Union, Any, Optional
  12. from .acoustic_features import (\
  13. get_pitch, calculate_time_varying_jitter, get_formants_frame_based,
  14. compute_msc, spectral_centriod, ltas, alpha_ratio, log_mel_spectrogram,
  15. mfcc, lpc, lpcc, spectral_envelope, calculate_cpp, hammIndex,
  16. plp_features, harmonicity, calculate_lsp_freqs_for_frames, calculate_frame_wise_zcr,
  17. analyze_audio_shimmer, calculate_frame_level_hnr, amplitude_range,
  18. rms_amplitude, spl_per_frame, peak_amplitude, short_time_energy, intensity,
  19. calculate_hfd_per_frame, calculate_frequency_entropy, calculate_amplitude_entropy,
  20. # Classes needed for process_file_model
  21. PauseBehavior, SpeechBehavior,
  22. # Statistical functions and names
  23. sma, de, function_names,
  24. )
  25. class AcousticFeatureExtractor:
  26. """
  27. A unified interface for extracting various acoustic features from speech recordings.
  28. """
  29. def __init__(self, sampling_rate=16000, log_level=logging.INFO):
  30. """
  31. Initialize the feature extractor
  32. Parameters:
  33. -----------
  34. sampling_rate : int
  35. Sampling rate of the audio files to process
  36. log_level : int
  37. Logging level (default: logging.INFO)
  38. """
  39. self.sampling_rate = sampling_rate
  40. # Configure logging
  41. self.logger = logging.getLogger(__name__)
  42. if not self.logger.handlers:
  43. handler = logging.StreamHandler()
  44. formatter = logging.Formatter('%(asctime)s - %(levelname)s - %(message)s')
  45. handler.setFormatter(formatter)
  46. self.logger.addHandler(handler)
  47. self.logger.setLevel(log_level)
  48. # Available feature types mapped to their extraction methods
  49. self.available_features = {
  50. 'spectral': self.extract_spectral_features,
  51. 'complexity': self.extract_complexity_features,
  52. 'frequency': self.extract_frequency_features,
  53. 'intensity': self.extract_intensity_features,
  54. 'rhythmic': self.extract_rhythmic_features,
  55. 'fluency': self.extract_fluency_features,
  56. 'voice_quality': self.extract_voice_quality_features,
  57. 'all': self.extract_all_features,
  58. 'raw': self.process_file,
  59. 'transcription': self.process_file_model
  60. }
  61. # VAD model and transcription model (to be set later if needed)
  62. self.vad_model = None
  63. self.vad_utils = None
  64. self.transcription_model = None
  65. self.logger.info("AcousticFeatureExtractor initialized with sampling rate %d Hz", sampling_rate)
  66. def set_models(self, vad_model=None, vad_utils=None, transcription_model=None):
  67. """
  68. Set the VAD and transcription models for advanced feature extraction
  69. Parameters:
  70. -----------
  71. vad_model : object
  72. Voice Activity Detection model
  73. vad_utils : object
  74. Utility functions for the VAD model
  75. transcription_model : object
  76. Transcription model for speech-to-text
  77. """
  78. self.vad_model = vad_model
  79. self.vad_utils = vad_utils
  80. self.transcription_model = transcription_model
  81. models_set = []
  82. if vad_model is not None:
  83. models_set.append("VAD model")
  84. if vad_utils is not None:
  85. models_set.append("VAD utils")
  86. if transcription_model is not None:
  87. models_set.append("transcription model")
  88. self.logger.info("Models configured: %s", ", ".join(models_set))
  89. def extract_all_features(self, audio_path):
  90. """
  91. Extract all available acoustic features from an audio file
  92. Parameters:
  93. -----------
  94. audio_path : str
  95. Path to the audio file
  96. Returns:
  97. --------
  98. dict
  99. Dictionary containing all extracted features
  100. """
  101. self.logger.info("Extracting all features from: %s", audio_path)
  102. features = {}
  103. # Extract all available feature sets except 'all' to avoid infinite recursion
  104. for feature_type, feature_method in self.available_features.items():
  105. # Skip 'all' to prevent infinite recursion
  106. if feature_type == 'all':
  107. continue
  108. try:
  109. feature_result = feature_method(audio_path)
  110. features.update(feature_result)
  111. except Exception as e:
  112. self.logger.error("Error extracting %s features: %s", feature_type, str(e))
  113. self.logger.info("Extracted %d features in total from 'all' option", len(features))
  114. return features
  115. def extract_spectral_features(self, audio_path):
  116. """Extract spectral features from audio"""
  117. self.logger.debug("Extracting spectral features from: %s", audio_path)
  118. try:
  119. # Load audio file
  120. data, fs = librosa.load(audio_path, sr=self.sampling_rate)
  121. # Define window parameters
  122. window_length_ms = 50
  123. window_step_ms = 25
  124. features = {}
  125. # Spectral features extraction
  126. try:
  127. # Modulation Spectrum Coefficients
  128. msc = compute_msc(data, fs, nfft=512, window_length_ms=window_length_ms, window_step_ms=window_step_ms, num_msc=13)
  129. features.update(self.process_matrix(msc, 'MSC'))
  130. # Spectral centroids
  131. centroids = spectral_centriod(data, fs, window_length_ms, window_step_ms)
  132. features.update(self.process_row(np.array(centroids), 'CENTRIODS'))
  133. # Long Term Average Spectrum
  134. LTAS, _ = ltas(data, fs, window_length_ms, window_step_ms)
  135. features.update(self.process_row(LTAS, 'LTAS'))
  136. # Alpha ratio
  137. ALPHA_RATIO = alpha_ratio(data, fs, window_length_ms, window_step_ms, (0, 1000), (1000, 5000))
  138. features['ALPHA_RATIO'] = ALPHA_RATIO
  139. # Log Mel Spectrogram
  140. LOG_MEL_SPECTROGRAM = log_mel_spectrogram(data, fs, window_length_ms, window_step_ms, melbands=8, fmin=20, fmax=6500)
  141. features.update(self.process_matrix(LOG_MEL_SPECTROGRAM, 'LOG_MEL_SPECTROGRAM'))
  142. # MFCCs
  143. MFCC = mfcc(data, fs, window_length_ms, window_step_ms, melbands=26, lifter=20)[:15]
  144. features.update(self.process_matrix(MFCC, 'MFCC'))
  145. # Linear Prediction Coefficients
  146. LPC = lpc(data, fs, window_length_ms, window_step_ms)
  147. features.update(self.process_matrix(LPC, 'LPC'))
  148. # Linear Prediction Cepstral Coefficients
  149. LPCC = lpcc(data, fs, window_length_ms, window_step_ms, lpc_length=8)
  150. features.update(self.process_matrix(LPCC, 'LPCC'))
  151. # Spectral Envelope
  152. ENVELOPE = spectral_envelope(data, fs, window_length_ms, window_step_ms)
  153. features.update(self.process_matrix(ENVELOPE, 'ENVELOPE'))
  154. # Cepstral Peak Prominence
  155. CPP = calculate_cpp(data, fs, window_length_ms, window_step_ms)
  156. features.update(self.process_row(CPP, 'CPP'))
  157. # Hammarberg Index
  158. HAMM_INDEX, _, _ = hammIndex(data, fs)
  159. features.update(self.process_row(HAMM_INDEX, 'HAMMARBERG_INDEX'))
  160. # Perceptual Linear Prediction
  161. PLP = plp_features(data, fs, num_filters=26, fmin=20, fmax=8000)
  162. features.update(self.process_matrix(PLP, 'PLP'))
  163. # Line Spectral Pairs
  164. lspFreq = calculate_lsp_freqs_for_frames(data, fs, window_length_ms, window_step_ms, order=8)
  165. features.update(self.process_matrix(lspFreq, "LSP"))
  166. # Zero Crossing Rate
  167. ZCR = calculate_frame_wise_zcr(data, fs, window_length_ms, window_step_ms)
  168. features.update(self.process_row(ZCR, "ZCR"))
  169. self.logger.info("Extracted %d spectral features", len(features))
  170. except Exception as e:
  171. self.logger.error(f"Error processing spectral features: {e}")
  172. return features
  173. except Exception as e:
  174. self.logger.error("Runtime error extracting spectral features: %s", str(e))
  175. return {}
  176. def extract_complexity_features(self, audio_path):
  177. """Extract complexity features from audio"""
  178. self.logger.debug("Extracting complexity features from: %s", audio_path)
  179. try:
  180. # Load audio file
  181. data, fs = librosa.load(audio_path, sr=self.sampling_rate)
  182. # Define window parameters
  183. window_length_ms = 50
  184. window_step_ms = 25
  185. features = {}
  186. # Complexity features extraction
  187. try:
  188. # Higuchi Fractal Dimension
  189. HFD = calculate_hfd_per_frame(data, fs, window_length_ms, window_step_ms, 10)
  190. features.update(self.process_row(HFD, 'HFD'))
  191. # Frequency Entropy
  192. FREQ_ENTROPY = calculate_frequency_entropy(data, fs, window_length_ms, window_step_ms)
  193. features.update(self.process_row(FREQ_ENTROPY, 'FREQ_ENTROPY'))
  194. # Amplitude Entropy
  195. AMP_ENTROPY = calculate_amplitude_entropy(data, fs, window_length_ms, window_step_ms)
  196. features.update(self.process_row(AMP_ENTROPY, 'AMP_ENTROPY'))
  197. # Zero Crossing Rate can also be considered a complexity feature
  198. ZCR = calculate_frame_wise_zcr(data, fs, window_length_ms, window_step_ms)
  199. features.update(self.process_row(ZCR, "ZCR"))
  200. self.logger.info("Extracted %d complexity features", len(features))
  201. except Exception as e:
  202. self.logger.error(f"Error processing complexity features: {e}")
  203. return features
  204. except Exception as e:
  205. self.logger.error("Runtime error extracting complexity features: %s", str(e))
  206. return {}
  207. def extract_frequency_features(self, audio_path):
  208. """Extract frequency features from audio"""
  209. self.logger.debug("Extracting frequency features from: %s", audio_path)
  210. try:
  211. # Load audio file
  212. data, fs = librosa.load(audio_path, sr=self.sampling_rate)
  213. # Define window parameters
  214. window_length_ms = 50
  215. window_step_ms = 25
  216. features = {}
  217. # Frequency parameters extraction
  218. try:
  219. # Fundamental Frequency (F0)
  220. F0 = get_pitch(data, fs, window_length_ms, window_step_ms)
  221. F0_valid = F0[~np.isnan(F0)]
  222. features.update(self.process_row(F0_valid, 'F0'))
  223. # Jitter
  224. jitter = calculate_time_varying_jitter(F0, fs, window_length_ms, window_step_ms, window_length_ms*2, window_step_ms*2)
  225. features.update(self.process_row(np.array(jitter), 'Jitter'))
  226. # Formants (F1, F2, F3)
  227. F_formants = get_formants_frame_based(data, fs, window_length_ms, window_step_ms, [1, 2, 3])
  228. # Handle the case where F_formants is returned as a tuple
  229. if isinstance(F_formants, tuple):
  230. # Extract the formant data from the tuple (first element is usually the formant data)
  231. formant_data = np.array(F_formants[0])
  232. else:
  233. formant_data = F_formants
  234. # Process each formant
  235. for i in range(formant_data.shape[0]):
  236. features.update(self.process_row(formant_data[i, :], f'F{i+1}'))
  237. # Harmonicity can also be considered a frequency feature
  238. HARMONICITY = harmonicity(data, fs)
  239. features['HARMONICITY'] = HARMONICITY
  240. self.logger.info("Extracted %d frequency features", len(features))
  241. except Exception as e:
  242. self.logger.error(f"Error processing frequency features: {e}")
  243. return features
  244. except Exception as e:
  245. self.logger.error("Runtime error extracting frequency features: %s", str(e))
  246. return {}
  247. def extract_intensity_features(self, audio_path):
  248. """Extract intensity and loudness features from audio"""
  249. self.logger.debug("Extracting intensity features from: %s", audio_path)
  250. try:
  251. # Load audio file
  252. data, fs = librosa.load(audio_path, sr=self.sampling_rate)
  253. # Define window parameters
  254. window_length_ms = 50
  255. window_step_ms = 25
  256. features = {}
  257. # Loudness and intensity parameters extraction
  258. try:
  259. # Root Mean Square amplitude
  260. RMS = rms_amplitude(data, fs, window_length_ms, window_step_ms)
  261. features.update(self.process_row(RMS, 'RMS'))
  262. # Sound Pressure Level
  263. SPL = spl_per_frame(data, fs, window_length_ms, window_step_ms)
  264. features.update(self.process_row(SPL, 'SPL'))
  265. # Peak amplitude
  266. PEAK = peak_amplitude(data, fs, window_length_ms, window_step_ms)
  267. features.update(self.process_row(PEAK, 'PEAK'))
  268. # Short-time energy
  269. STE = short_time_energy(data, fs, window_length_ms, window_step_ms)
  270. features.update(self.process_row(STE, 'STE'))
  271. # Intensity
  272. INTENSITY = intensity(data, fs, window_length_ms, window_step_ms)
  273. features.update(self.process_row(INTENSITY, 'INTENSITY'))
  274. # Amplitude range
  275. APQ_range, APQ_std = amplitude_range(data, fs, window_length_ms, window_step_ms)
  276. features.update(self.process_row(APQ_range, 'Amplitude_Range'))
  277. features.update(self.process_row(APQ_std, 'APQ2'))
  278. self.logger.info("Extracted %d intensity features", len(features))
  279. except Exception as e:
  280. self.logger.error(f"Error processing intensity features: {e}")
  281. return features
  282. except Exception as e:
  283. self.logger.error("Runtime error extracting intensity features: %s", str(e))
  284. return {}
  285. def extract_rhythmic_features(self, audio_path):
  286. """Extract rhythmic features from audio"""
  287. self.logger.debug("Extracting rhythmic features from: %s", audio_path)
  288. # Check if required models are available
  289. if None in (self.vad_model, self.vad_utils, self.transcription_model):
  290. self.logger.warning("VAD and transcription models are required for rhythmic features but not set")
  291. return {}
  292. try:
  293. features = {}
  294. # Create PauseBehavior instance for rhythmic analysis
  295. try:
  296. pause_behavior = PauseBehavior(self.vad_model, self.vad_utils, self.transcription_model)
  297. pause_behavior.configure(audio_path)
  298. # List of methods to extract from pause_behavior for rhythmic features
  299. rhythmic_methods = [
  300. 'syllable_rate', 'speech_rate', 'articulation_rate',
  301. 'mean_pause_duration', 'mean_silence_duration', 'mean_speech_duration',
  302. 'speech_to_pause_ratio', 'percentage_silence', 'percentage_voice'
  303. ]
  304. # Extract relevant rhythmic features
  305. for name in rhythmic_methods:
  306. method = getattr(pause_behavior, name, None)
  307. if method and callable(method):
  308. try:
  309. features[name] = method()
  310. except Exception as e:
  311. self.logger.error(f"Error executing PauseBehavior method {name}: {e}")
  312. self.logger.info("Extracted %d rhythmic features", len(features))
  313. except Exception as e:
  314. self.logger.error(f"Error in rhythmic feature extraction: {e}")
  315. return features
  316. except Exception as e:
  317. self.logger.error("Runtime error extracting rhythmic features: %s", str(e))
  318. return {}
  319. def extract_fluency_features(self, audio_path):
  320. """Extract speech fluency features from audio"""
  321. self.logger.debug("Extracting fluency features from: %s", audio_path)
  322. # Check if required models are available
  323. if None in (self.vad_model, self.vad_utils, self.transcription_model):
  324. self.logger.warning("VAD and transcription models are required for fluency features but not set")
  325. return {}
  326. try:
  327. features = {}
  328. # Create SpeechBehavior instance for fluency analysis
  329. try:
  330. # First, set up PauseBehavior to get speech segmentation
  331. pause_behavior = PauseBehavior(self.vad_model, self.vad_utils, self.transcription_model)
  332. pause_behavior.configure(audio_path)
  333. # Then, set up SpeechBehavior using data from PauseBehavior
  334. speech_behavior = SpeechBehavior(self.vad_model, self.vad_utils, self.transcription_model)
  335. # Copy data from pause_behavior to speech_behavior
  336. speech_behavior.data = pause_behavior.data
  337. speech_behavior.silence_ranges = pause_behavior.silence_ranges
  338. speech_behavior.speech_ranges = pause_behavior.speech_ranges
  339. speech_behavior.transcription_result = pause_behavior.transcription_result
  340. speech_behavior.text = pause_behavior.text
  341. # Perform phoneme alignment
  342. speech_behavior.phoneme_alignment(audio_path)
  343. # List of methods to extract from speech_behavior for fluency features
  344. fluency_methods = [
  345. 'phonation_rate', 'phonation_time', 'articulation_time',
  346. 'mean_duration_of_bursts', 'number_of_pauses', 'number_of_filled_pauses',
  347. 'filled_pauses_per_min', 'mean_length_of_runs', 'hesitation_ratio'
  348. ]
  349. # Extract relevant fluency features
  350. for name in fluency_methods:
  351. method = getattr(speech_behavior, name, None)
  352. if method and callable(method):
  353. try:
  354. features[name] = method()
  355. except Exception as e:
  356. self.logger.error(f"Error executing SpeechBehavior method {name}: {e}")
  357. # Process regularity and PVI features
  358. try:
  359. for i, res in enumerate(speech_behavior.regularity_of_segments()):
  360. features[f"regularity_{i}"] = res
  361. for i, res in enumerate(speech_behavior.alternating_regularity()):
  362. features[f"PVI_{i}"] = res
  363. except Exception as e_reg:
  364. self.logger.error(f"Error processing regularity/PVI features: {e_reg}")
  365. # Process relative sentence duration
  366. try:
  367. relative_sentence_duration = speech_behavior.relative_sentence_duration()
  368. features.update(self.process_row(np.array(relative_sentence_duration), "relative_sentence_duration"))
  369. except Exception as e_rel:
  370. self.logger.error(f"Error processing relative sentence duration: {e_rel}")
  371. self.logger.info("Extracted %d fluency features", len(features))
  372. except Exception as e:
  373. self.logger.error(f"Error in fluency feature extraction: {e}")
  374. return features
  375. except Exception as e:
  376. self.logger.error("Runtime error extracting fluency features: %s", str(e))
  377. return {}
  378. def extract_voice_quality_features(self, audio_path):
  379. """Extract voice quality features from audio"""
  380. self.logger.debug("Extracting voice quality features from: %s", audio_path)
  381. try:
  382. # Load audio file
  383. data, fs = librosa.load(audio_path, sr=self.sampling_rate)
  384. # Define window parameters
  385. window_length_ms = 50
  386. window_step_ms = 25
  387. features = {}
  388. # Voice quality features extraction
  389. try:
  390. # Shimmer
  391. SHIMMER = analyze_audio_shimmer(data, fs, window_length_ms, window_step_ms)
  392. features.update(self.process_row(SHIMMER, 'SHIMMER'))
  393. # Harmonics-to-Noise Ratio and Noise-to-Harmonics Ratio
  394. HNR, NHR = calculate_frame_level_hnr(data, fs, window_length_ms, window_step_ms)
  395. features.update(self.process_row(HNR, 'HNR'))
  396. features.update(self.process_row(NHR, 'NHR'))
  397. # Harmonicity
  398. HARMONICITY = harmonicity(data, fs)
  399. features['HARMONICITY'] = HARMONICITY
  400. # Cepstral Peak Prominence (also a voice quality feature)
  401. CPP = calculate_cpp(data, fs, window_length_ms, window_step_ms)
  402. features.update(self.process_row(CPP, 'CPP'))
  403. # Hammarberg Index (spectral tilt related to voice quality)
  404. HAMM_INDEX, _, _ = hammIndex(data, fs)
  405. features.update(self.process_row(HAMM_INDEX, 'HAMMARBERG_INDEX'))
  406. # Alpha Ratio (spectral tilt related to voice quality)
  407. ALPHA_RATIO = alpha_ratio(data, fs, window_length_ms, window_step_ms, (0, 1000), (1000, 5000))
  408. features['ALPHA_RATIO'] = ALPHA_RATIO
  409. # Jitter (also a voice quality feature)
  410. F0 = get_pitch(data, fs, window_length_ms, window_step_ms)
  411. jitter = calculate_time_varying_jitter(F0, fs, window_length_ms, window_step_ms, window_length_ms*2, window_step_ms*2)
  412. features.update(self.process_row(np.array(jitter), 'Jitter'))
  413. self.logger.info("Extracted %d voice quality features", len(features))
  414. except Exception as e:
  415. self.logger.error(f"Error processing voice quality features: {e}")
  416. return features
  417. except Exception as e:
  418. self.logger.error("Runtime error extracting voice quality features: %s", str(e))
  419. return {}
  420. def process_row(self, row, feature_name, index=-1):
  421. """
  422. Process a single row of data with statistical functions
  423. Parameters:
  424. -----------
  425. row : numpy.ndarray
  426. Row of data to process
  427. feature_name : str
  428. Name of the feature
  429. index : int, optional
  430. Index for matrix features, -1 for non-matrix features
  431. Returns:
  432. --------
  433. dict
  434. Dictionary of processed features
  435. """
  436. # Import all statistical functions individually
  437. from .acoustic_features.statistical_functions import (
  438. max, min, span, maxPos, minPos, amean, linregc1, linregc2,
  439. linregerrA, linregerrQ, stddev, skewness, kurtosis,
  440. quartile1, quartile2, quartile3, iqr1_2, iqr2_3, iqr1_3,
  441. percentile1, percentile99, pctlrange0_1, upleveltime75, upleveltime90
  442. )
  443. # Create a dictionary mapping function names to the actual function
  444. stat_funcs = {
  445. "max": max, "min": min, "span": span, "maxPos": maxPos, "minPos": minPos,
  446. "amean": amean, "linregc1": linregc1, "linregc2": linregc2,
  447. "linregerrA": linregerrA, "linregerrQ": linregerrQ, "stddev": stddev,
  448. "skewness": skewness, "kurtosis": kurtosis, "quartile1": quartile1,
  449. "quartile2": quartile2, "quartile3": quartile3, "iqr1_2": iqr1_2,
  450. "iqr2_3": iqr2_3, "iqr1_3": iqr1_3, "percentile1": percentile1,
  451. "percentile99": percentile99, "pctlrange0_1": pctlrange0_1,
  452. "upleveltime75": upleveltime75, "upleveltime90": upleveltime90
  453. }
  454. row_sma = sma(row)
  455. results = {}
  456. # Apply statistical functions to smoothed signal
  457. for func_name in function_names:
  458. # Get the function from our mapping instead of globals()
  459. func = stat_funcs.get(func_name)
  460. if func is None:
  461. self.logger.warning(f"Statistical function {func_name} not found in stat_funcs mapping.")
  462. continue
  463. try:
  464. result = func(row_sma)
  465. except Exception as e:
  466. self.logger.error(f"Error applying {func_name} to {feature_name}: {e}")
  467. result = None
  468. if index >= 0:
  469. name = f"{feature_name}_sma[{index}]_{func_name}"
  470. else:
  471. name = f"{feature_name}_sma_{func_name}"
  472. results[name] = result
  473. # Apply statistical functions to derivative of smoothed signal
  474. row_sma_de = de(row_sma)
  475. for func_name in function_names:
  476. func = stat_funcs.get(func_name)
  477. if func is None:
  478. self.logger.warning(f"Statistical function {func_name} not found in stat_funcs mapping.")
  479. continue
  480. try:
  481. result = func(row_sma_de)
  482. except Exception as e:
  483. self.logger.error(f"Error applying {func_name} to derivative of {feature_name}: {e}")
  484. result = None
  485. if index >= 0:
  486. name = f"{feature_name}_sma_de[{index}]_{func_name}"
  487. else:
  488. name = f"{feature_name}_sma_de_{func_name}"
  489. results[name] = result
  490. return results
  491. def process_matrix(self, matrix, feature_name):
  492. """
  493. Process a matrix of features with statistical functions
  494. Parameters:
  495. -----------
  496. matrix : numpy.ndarray
  497. Matrix to process
  498. feature_name : str
  499. Name of the feature
  500. Returns:
  501. --------
  502. dict
  503. Dictionary of processed features
  504. """
  505. matrix_results = {}
  506. for i, row in enumerate(matrix):
  507. result = self.process_row(row, feature_name, i)
  508. matrix_results.update(result)
  509. return matrix_results
  510. def length(self, res):
  511. """
  512. Count the total number of features in a list of dictionaries
  513. Parameters:
  514. -----------
  515. res : list
  516. List of dictionaries
  517. Returns:
  518. --------
  519. int
  520. Total number of features
  521. """
  522. return sum([len(elm) for elm in res])
  523. def process_file(self, filepath):
  524. """
  525. Process a file and extract all raw acoustic features
  526. Parameters:
  527. -----------
  528. filepath : str
  529. Path to the audio file
  530. Returns:
  531. --------
  532. dict
  533. Dictionary of all extracted features
  534. """
  535. self.logger.info("Processing file for raw feature extraction: %s", filepath)
  536. try:
  537. data, fs = librosa.load(filepath, sr=self.sampling_rate)
  538. except Exception as e:
  539. self.logger.error(f"Failed to load audio file {filepath}: {e}")
  540. return {}
  541. # Define window parameters
  542. window_length_ms = 50
  543. window_step_ms = 25
  544. # --- Feature Calculation with Error Handling ---
  545. features = {}
  546. # Frequency parameters
  547. try:
  548. F0 = get_pitch(data, fs, window_length_ms, window_step_ms)
  549. F0_valid = F0[~np.isnan(F0)]
  550. features.update(self.process_row(F0_valid, 'F0'))
  551. jitter = calculate_time_varying_jitter(F0, fs, window_length_ms, window_step_ms, window_length_ms*2, window_step_ms*2)
  552. features.update(self.process_row(np.array(jitter), 'Jitter'))
  553. F_formants = get_formants_frame_based(data, fs, window_length_ms, window_step_ms, [1, 2, 3])
  554. # Handle the case where F_formants is returned as a tuple
  555. if isinstance(F_formants, tuple):
  556. # Extract the formant data from the tuple (first element is usually the formant data)
  557. formant_data = np.array(F_formants[0])
  558. else:
  559. formant_data = F_formants
  560. # Process each formant
  561. for i in range(formant_data.shape[0]):
  562. features.update(self.process_row(formant_data[i, :], f'F{i+1}'))
  563. self.logger.info("Processed frequency parameters")
  564. except Exception as e:
  565. self.logger.error(f"Error processing frequency parameters: {e}")
  566. # Spectral features
  567. try:
  568. msc = compute_msc(data, fs, nfft=512, window_length_ms=window_length_ms, window_step_ms=window_step_ms, num_msc=13)
  569. features.update(self.process_matrix(msc, 'MSC'))
  570. centroids = spectral_centriod(data, fs, window_length_ms, window_step_ms)
  571. features.update(self.process_row(np.array(centroids), 'CENTRIODS'))
  572. LTAS, _ = ltas(data, fs, window_length_ms, window_step_ms) # Ignore freq array return
  573. features.update(self.process_row(LTAS, 'LTAS'))
  574. ALPHA_RATIO = alpha_ratio(data, fs, window_length_ms, window_step_ms, (0, 1000), (1000, 5000))
  575. # Alpha ratio is often scalar, handle differently if needed, here adding directly
  576. features['ALPHA_RATIO'] = ALPHA_RATIO
  577. LOG_MEL_SPECTROGRAM = log_mel_spectrogram(data, fs, window_length_ms, window_step_ms, melbands=8, fmin=20, fmax=6500)
  578. features.update(self.process_matrix(LOG_MEL_SPECTROGRAM, 'LOG_MEL_SPECTROGRAM'))
  579. MFCC = mfcc(data, fs, window_length_ms, window_step_ms, melbands=26, lifter=20)[:15]
  580. features.update(self.process_matrix(MFCC, 'MFCC'))
  581. LPC = lpc(data, fs, window_length_ms, window_step_ms)
  582. features.update(self.process_matrix(LPC, 'LPC'))
  583. LPCC = lpcc(data, fs, window_length_ms, window_step_ms, lpc_length=8)
  584. features.update(self.process_matrix(LPCC, 'LPCC'))
  585. ENVELOPE = spectral_envelope(data, fs, window_length_ms, window_step_ms)
  586. features.update(self.process_matrix(ENVELOPE, 'ENVELOPE'))
  587. CPP = calculate_cpp(data, fs, window_length_ms, window_step_ms)
  588. features.update(self.process_row(CPP, 'CPP'))
  589. HAMM_INDEX, _, _ = hammIndex(data, fs) # Ignore freq and Pxx return
  590. features.update(self.process_row(HAMM_INDEX, 'HAMMARBERG_INDEX'))
  591. PLP = plp_features(data, fs, num_filters=26, fmin=20, fmax=8000)
  592. features.update(self.process_matrix(PLP, 'PLP'))
  593. HARMONICITY = harmonicity(data, fs)
  594. features['HARMONICITY'] = HARMONICITY # Often scalar
  595. lspFreq = calculate_lsp_freqs_for_frames(data, fs, window_length_ms, window_step_ms, order=8)
  596. features.update(self.process_matrix(lspFreq, "LSP"))
  597. ZCR = calculate_frame_wise_zcr(data, fs, window_length_ms, window_step_ms)
  598. features.update(self.process_row(ZCR, "ZCR"))
  599. self.logger.info("Processed spectral domain features")
  600. except Exception as e:
  601. self.logger.error(f"Error processing spectral features: {e}")
  602. # Voice quality
  603. try:
  604. SHIMMER = analyze_audio_shimmer(data, fs, window_length_ms, window_step_ms)
  605. features.update(self.process_row(SHIMMER, 'SHIMMER'))
  606. HNR, NHR = calculate_frame_level_hnr(data, fs, window_length_ms, window_step_ms)
  607. features.update(self.process_row(HNR, 'HNR'))
  608. features.update(self.process_row(NHR, 'NHR'))
  609. # Amplitude range returns range and std deviation
  610. APQ_range, APQ_std = amplitude_range(data, fs, window_length_ms, window_step_ms)
  611. features.update(self.process_row(APQ_range, 'Amplitude_Range')) # Renamed from APQ_range
  612. features.update(self.process_row(APQ_std, 'APQ2')) # Renamed from APQ_std
  613. self.logger.info("Processed voice quality features")
  614. except Exception as e:
  615. self.logger.error(f"Error processing voice quality features: {e}")
  616. # Loudness and intensity
  617. try:
  618. RMS = rms_amplitude(data, fs, window_length_ms, window_step_ms)
  619. features.update(self.process_row(RMS, 'RMS'))
  620. SPL = spl_per_frame(data, fs, window_length_ms, window_step_ms)
  621. features.update(self.process_row(SPL, 'SPL'))
  622. PEAK = peak_amplitude(data, fs, window_length_ms, window_step_ms)
  623. features.update(self.process_row(PEAK, 'PEAK'))
  624. STE = short_time_energy(data, fs, window_length_ms, window_step_ms)
  625. features.update(self.process_row(STE, 'STE'))
  626. INTENSITY = intensity(data, fs, window_length_ms, window_step_ms)
  627. features.update(self.process_row(INTENSITY, 'INTENSITY'))
  628. self.logger.info("Processed loudness and intensity parameters")
  629. except Exception as e:
  630. self.logger.error(f"Error processing loudness and intensity features: {e}")
  631. # Complexity
  632. try:
  633. HFD = calculate_hfd_per_frame(data, fs, window_length_ms, window_step_ms, 10)
  634. features.update(self.process_row(HFD, 'HFD'))
  635. FREQ_ENTROPY = calculate_frequency_entropy(data, fs, window_length_ms, window_step_ms)
  636. features.update(self.process_row(FREQ_ENTROPY, 'FREQ_ENTROPY'))
  637. AMP_ENTROPY = calculate_amplitude_entropy(data, fs, window_length_ms, window_step_ms)
  638. features.update(self.process_row(AMP_ENTROPY, 'AMP_ENTROPY'))
  639. self.logger.info("Processed complexity features")
  640. except Exception as e:
  641. self.logger.error(f"Error processing complexity features: {e}")
  642. self.logger.info("Extracted a total of %d raw features", len(features))
  643. return features
  644. def process_file_model(self, filepath):
  645. """
  646. Process a file using VAD and transcription models
  647. Parameters:
  648. -----------
  649. filepath : str
  650. Path to the audio file
  651. Returns:
  652. --------
  653. dict
  654. Dictionary of all extracted features
  655. Raises:
  656. -------
  657. ValueError
  658. If models are not set
  659. ImportError
  660. If required dependencies for fluency/rhythmic features are missing
  661. """
  662. if None in (self.vad_model, self.vad_utils, self.transcription_model):
  663. raise ValueError("VAD and transcription models must be set using set_models() before calling this method")
  664. # Raise the specific import error if possible, otherwise a generic one
  665. try:
  666. # Attempting to instantiate will raise the specific error if placeholder is used
  667. PauseBehavior(None, None, None)
  668. SpeechBehavior(None, None, None)
  669. except Exception:
  670. raise ImportError("Required dependencies for fluency/rhythmic features are missing.")
  671. self.logger.info("Processing file with VAD and transcription models: %s", filepath)
  672. p_results = {}
  673. s_results = {}
  674. prob = {}
  675. try:
  676. # Process rhythmic features
  677. pause_behavior = PauseBehavior(self.vad_model, self.vad_utils, self.transcription_model)
  678. pause_behavior.configure(filepath)
  679. voiceProb_signal = self.vad_model.audio_forward(pause_behavior.data, sr=16000)
  680. voiceProb_signal = np.array(voiceProb_signal[0])
  681. prob = self.process_row(voiceProb_signal, "voiceProb")
  682. # List of methods to exclude
  683. excluded_methods_p = ['__init__', 'configure']
  684. # Iterate through all methods of the pause_behavior instance
  685. self.logger.debug("Extracting pause behavior features")
  686. for name, method in inspect.getmembers(pause_behavior, predicate=inspect.ismethod):
  687. if name not in excluded_methods_p:
  688. try:
  689. p_results[name] = method()
  690. except Exception as e_method:
  691. self.logger.error(f"Error executing PauseBehavior method {name}: {e_method}")
  692. # Process speech features
  693. speech_behavior = SpeechBehavior(self.vad_model, self.vad_utils, self.transcription_model)
  694. # Copy data from pause_behavior to speech_behavior
  695. speech_behavior.data = pause_behavior.data
  696. speech_behavior.silence_ranges = pause_behavior.silence_ranges
  697. speech_behavior.speech_ranges = pause_behavior.speech_ranges
  698. speech_behavior.transcription_result = pause_behavior.transcription_result
  699. speech_behavior.text = pause_behavior.text
  700. # Perform phoneme alignment
  701. alignment_error = speech_behavior.phoneme_alignment(filepath)
  702. self.logger.info("Phoneme alignment completed (total segments: %d, hypothetical: %d)",
  703. alignment_error[0], alignment_error[1])
  704. # List of methods to exclude
  705. excluded_methods_s = ['__init__', 'configure', 'phoneme_alignment', 'relative_sentence_duration', 'regularity_of_segments', 'alternating_regularity']
  706. # Iterate through all methods of the speech_behavior instance
  707. self.logger.debug("Extracting speech behavior features")
  708. for name, method in inspect.getmembers(speech_behavior, predicate=inspect.ismethod):
  709. if name not in excluded_methods_s:
  710. try:
  711. s_results[name] = method()
  712. except Exception as e_method:
  713. self.logger.error(f"Error executing SpeechBehavior method {name}: {e_method}")
  714. # Process regularity and PVI features
  715. self.logger.debug("Processing regularity and PVI features")
  716. try:
  717. for i, res in enumerate(speech_behavior.regularity_of_segments()):
  718. s_results[f"regularity_{i}"] = res
  719. for i, res in enumerate(speech_behavior.alternating_regularity()):
  720. s_results[f"PVI_{i}"] = res
  721. except Exception as e_reg:
  722. self.logger.error(f"Error processing regularity/PVI features: {e_reg}")
  723. # Process relative sentence duration
  724. try:
  725. relative_sentence_duration = speech_behavior.relative_sentence_duration()
  726. s_results.update(self.process_row(np.array(relative_sentence_duration), "relative_sentence_duration"))
  727. except Exception as e_rel:
  728. self.logger.error(f"Error processing relative sentence duration: {e_rel}")
  729. except Exception as e_main:
  730. self.logger.critical(f"Major error during model-based processing: {e_main}")
  731. # Depending on severity, you might want to return partial results or an empty dict
  732. # return {}
  733. # Merge all results
  734. final_results = {**p_results, **prob, **s_results}
  735. self.logger.debug("Feature counts - pause: %d, prob: %d, speech: %d", len(p_results), len(prob), len(s_results))
  736. self.logger.info("Extracted a total of %d features with VAD and transcription models", len(final_results))
  737. return final_results
  738. def extract_features(self, audio_paths: Union[str, List[str], 'torch.utils.data.DataLoader'], features_to_calculate: Optional[List[str]] = None, separate_groups: bool = False, output_dir: Optional[str] = None) -> Dict[str, Any]:
  739. """
  740. Extract specified features from a single audio file, a batch of files, or a PyTorch DataLoader
  741. Parameters:
  742. -----------
  743. audio_paths : str or List[str] or torch.utils.data.DataLoader
  744. Path(s) to the audio file(s) or a PyTorch DataLoader that yields:
  745. - File paths as strings
  746. - Tuples/lists where the first element is a file path
  747. - Audio tensors with shape (batch_size, num_samples) or (num_samples,)
  748. features_to_calculate : List[str], optional
  749. List of feature types to extract. If None, extracts all features.
  750. Valid options: 'spectral', 'complexity', 'frequency', 'intensity',
  751. 'rhythmic', 'fluency', 'voice_quality', 'all',
  752. 'raw', 'transcription'
  753. separate_groups : bool, optional
  754. If True, organizes output by feature groups. If False (default), returns a flat dictionary.
  755. output_dir : str, optional
  756. If provided, extracted features for each file will be saved as a CSV file in this directory
  757. with the same name as the audio file.
  758. Returns:
  759. --------
  760. dict
  761. Dictionary containing extracted features
  762. If separate_groups=False (default):
  763. For a single file: {feature_name: feature_value, ...}
  764. For multiple files/DataLoader: {file_path/idx: {feature_name: feature_value, ...}, ...}
  765. If separate_groups=True:
  766. For a single file: {feature_group: {feature_name: feature_value, ...}, ...}
  767. For multiple files/DataLoader: {file_path/idx: {feature_group: {feature_name: feature_value, ...}, ...}, ...}
  768. """
  769. # Default to all features if none specified
  770. if features_to_calculate is None:
  771. features_to_calculate = ['all']
  772. self.logger.info("Extracting features: %s", ", ".join(features_to_calculate))
  773. # Check if output directory exists and create it if needed
  774. if output_dir is not None:
  775. os.makedirs(output_dir, exist_ok=True)
  776. self.logger.info(f"Features will be saved to directory: {output_dir}")
  777. # Check if feature types are valid and available
  778. requested_unavailable = []
  779. for feature_type in features_to_calculate:
  780. if feature_type not in self.available_features:
  781. raise ValueError(f"Unknown feature type: {feature_type}. Available types: {list(self.available_features.keys())}")
  782. if requested_unavailable:
  783. self.logger.warning(
  784. "The following requested feature types are unavailable due to missing dependencies and will be skipped: %s",
  785. ", ".join(requested_unavailable)
  786. )
  787. # Check if input is a PyTorch DataLoader
  788. try:
  789. import torch
  790. from torch.utils.data import DataLoader
  791. is_dataloader = isinstance(audio_paths, DataLoader)
  792. except ImportError:
  793. is_dataloader = False
  794. # Process a PyTorch DataLoader
  795. if is_dataloader:
  796. self.logger.info("Processing PyTorch DataLoader")
  797. batch_results = {}
  798. # Process each batch from the DataLoader
  799. for batch_idx, batch in enumerate(audio_paths):
  800. self.logger.info(f"Processing batch {batch_idx+1}")
  801. # Handle different batch formats:
  802. # 1. Batch of file paths (strings)
  803. # 2. Batch of tuples where first element is file path
  804. # 3. Batch of audio tensors
  805. if isinstance(batch, torch.Tensor):
  806. # Batch is a tensor containing audio samples
  807. if len(batch.shape) == 1: # Single audio sample
  808. self.logger.debug("Processing single audio tensor")
  809. # Convert tensor to numpy array
  810. audio_data = batch.cpu().numpy()
  811. # Process single audio array
  812. file_result = self._process_audio_data(
  813. audio_data,
  814. features_to_calculate,
  815. f"tensor_batch{batch_idx}"
  816. )
  817. key = f"batch{batch_idx}_item0"
  818. if separate_groups and file_result:
  819. self.logger.info("Extracted %d features total for %s", len(file_result), key)
  820. grouped_file_result = self._organize_features_by_group(file_result)
  821. batch_results[key] = grouped_file_result
  822. else:
  823. batch_results[key] = file_result
  824. # Save features to CSV if output_dir is provided
  825. if output_dir is not None and file_result:
  826. csv_path = os.path.join(output_dir, f"{key}.csv")
  827. self._save_features_to_csv(file_result, csv_path)
  828. else: # Batch of audio samples
  829. self.logger.debug(f"Processing batch of {batch.shape[0]} audio tensors")
  830. for i in range(batch.shape[0]):
  831. audio_data = batch[i].cpu().numpy()
  832. file_result = self._process_audio_data(
  833. audio_data,
  834. features_to_calculate,
  835. f"tensor_batch{batch_idx}_item{i}"
  836. )
  837. key = f"batch{batch_idx}_item{i}"
  838. if separate_groups and file_result:
  839. self.logger.info("Extracted %d features total for %s", len(file_result), key)
  840. grouped_file_result = self._organize_features_by_group(file_result)
  841. batch_results[key] = grouped_file_result
  842. else:
  843. batch_results[key] = file_result
  844. # Save features to CSV if output_dir is provided
  845. if output_dir is not None and file_result:
  846. csv_path = os.path.join(output_dir, f"{key}.csv")
  847. self._save_features_to_csv(file_result, csv_path)
  848. elif isinstance(batch, (list, tuple)):
  849. # Batch is a list or tuple - check first element
  850. if len(batch) > 0 and isinstance(batch[0], str):
  851. # Batch contains file paths
  852. for i, item in enumerate(batch):
  853. file_result = self._process_single_file(item, features_to_calculate)
  854. if separate_groups and file_result:
  855. self.logger.info("Extracted %d features total for %s", len(file_result), item)
  856. grouped_file_result = self._organize_features_by_group(file_result)
  857. batch_results[item] = grouped_file_result
  858. else:
  859. batch_results[item] = file_result
  860. # Save features to CSV if output_dir is provided
  861. if output_dir is not None and file_result:
  862. file_name = os.path.splitext(os.path.basename(item))[0]
  863. csv_path = os.path.join(output_dir, f"{file_name}.csv")
  864. self._save_features_to_csv(file_result, csv_path)
  865. elif len(batch) > 0 and isinstance(batch[0], (list, tuple)) and len(batch[0]) > 0 and isinstance(batch[0][0], str):
  866. # Batch contains tuples/lists where first element is a file path
  867. for i, item in enumerate(batch):
  868. file_path = item[0] # Get the file path (first element)
  869. file_result = self._process_single_file(file_path, features_to_calculate)
  870. if separate_groups and file_result:
  871. self.logger.info("Extracted %d features total for %s", len(file_result), file_path)
  872. grouped_file_result = self._organize_features_by_group(file_result)
  873. batch_results[file_path] = grouped_file_result
  874. else:
  875. batch_results[file_path] = file_result
  876. # Save features to CSV if output_dir is provided
  877. if output_dir is not None and file_result:
  878. file_name = os.path.splitext(os.path.basename(file_path))[0]
  879. csv_path = os.path.join(output_dir, f"{file_name}.csv")
  880. self._save_features_to_csv(file_result, csv_path)
  881. else:
  882. self.logger.warning(f"Unsupported batch format. Skipping batch {batch_idx}")
  883. continue
  884. else:
  885. self.logger.warning(f"Unsupported batch item type: {type(batch)}. Skipping batch {batch_idx}")
  886. continue
  887. self.logger.info("DataLoader processing complete. Processed %d items total.", len(batch_results))
  888. return batch_results
  889. # Process a single file
  890. elif isinstance(audio_paths, str):
  891. result = self._process_single_file(audio_paths, features_to_calculate, separate_groups)
  892. # Save features to CSV if output_dir is provided
  893. if output_dir is not None and result:
  894. file_name = os.path.splitext(os.path.basename(audio_paths))[0]
  895. csv_path = os.path.join(output_dir, f"{file_name}.csv")
  896. # For single file, the result might be already grouped
  897. if separate_groups:
  898. # Flatten the grouped features for CSV output
  899. flattened_features = self._flatten_grouped_features(result)
  900. self._save_features_to_csv(flattened_features, csv_path)
  901. else:
  902. self._save_features_to_csv(result, csv_path)
  903. return result
  904. # Process batch of files
  905. else:
  906. self.logger.info("Processing batch of %d files", len(audio_paths))
  907. batch_results = {}
  908. for i, audio_path in enumerate(audio_paths):
  909. self.logger.info("Processing file %d/%d: %s", i+1, len(audio_paths), audio_path)
  910. file_result = self._process_single_file(audio_path, features_to_calculate)
  911. # If separate_groups is True, organize features by group for each file
  912. if separate_groups and file_result:
  913. self.logger.info("Extracted %d features total for %s", len(file_result), audio_path)
  914. grouped_file_result = self._organize_features_by_group(file_result)
  915. batch_results[audio_path] = grouped_file_result
  916. else:
  917. batch_results[audio_path] = file_result
  918. # Save features to CSV if output_dir is provided
  919. if output_dir is not None and file_result:
  920. file_name = os.path.splitext(os.path.basename(audio_path))[0]
  921. csv_path = os.path.join(output_dir, f"{file_name}.csv")
  922. self._save_features_to_csv(file_result, csv_path)
  923. self.logger.info("Batch processing complete")
  924. return batch_results
  925. def _process_single_file(self, audio_path, features_to_calculate, separate_groups=False):
  926. """
  927. Process a single audio file and extract features
  928. Parameters:
  929. -----------
  930. audio_path : str
  931. Path to the audio file
  932. features_to_calculate : List[str]
  933. List of feature types to extract
  934. separate_groups : bool
  935. If True, organizes output by feature groups
  936. Returns:
  937. --------
  938. dict
  939. Dictionary of extracted features
  940. """
  941. self.logger.info("Processing single file: %s", audio_path)
  942. result = {}
  943. for feature_type in features_to_calculate:
  944. # If 'all' is included, extract all available features and skip other types
  945. if feature_type == 'all':
  946. result = self.extract_all_features(audio_path)
  947. break
  948. # Extract each requested feature type
  949. # Add try-except around the call for robustness
  950. try:
  951. feature_extractor = self.available_features[feature_type]
  952. features = feature_extractor(audio_path)
  953. result.update(features)
  954. except Exception as e:
  955. self.logger.error(f"Error during extraction of '{feature_type}' features for {audio_path}: {e}")
  956. # If separate_groups is True, organize features by group
  957. if separate_groups and result:
  958. self.logger.info("Extracted %d features total for %s", len(result), audio_path)
  959. grouped_result = self._organize_features_by_group(result)
  960. self.logger.info("Extracted features organized into %d groups for %s", len(grouped_result), audio_path)
  961. return grouped_result
  962. self.logger.info("Extracted %d features total for %s", len(result), audio_path)
  963. return result
  964. def _process_audio_data(self, audio_data, features_to_calculate, identifier="audio_tensor"):
  965. """
  966. Process audio data directly from a numpy array and extract features
  967. Parameters:
  968. -----------
  969. audio_data : numpy.ndarray
  970. Audio samples as a numpy array
  971. features_to_calculate : List[str]
  972. List of feature types to extract
  973. identifier : str
  974. Identifier for logging purposes
  975. Returns:
  976. --------
  977. dict
  978. Dictionary of extracted features
  979. """
  980. self.logger.info("Processing audio data: %s", identifier)
  981. result = {}
  982. # Create a function that processes audio data for each feature type
  983. def process_audio_with_feature(feature_type):
  984. # Skip unsupported feature types for direct audio data
  985. if feature_type in ['rhythmic', 'fluency', 'transcription']:
  986. self.logger.warning(f"Feature type '{feature_type}' not supported for direct audio data processing")
  987. return {}
  988. if feature_type == 'all':
  989. # For 'all', just call each supported feature type
  990. all_features = {}
  991. for ft in ['spectral', 'complexity', 'frequency', 'intensity', 'voice_quality']:
  992. all_features.update(process_audio_with_feature(ft))
  993. return all_features
  994. # Extract features directly from audio data
  995. try:
  996. # Most feature extraction functions require signal and sampling rate
  997. fs = self.sampling_rate
  998. data = audio_data
  999. window_length_ms = 50
  1000. window_step_ms = 25
  1001. features = {}
  1002. if feature_type == 'spectral':
  1003. # Extract spectral features from data
  1004. try:
  1005. # Modulation Spectrum Coefficients
  1006. msc = compute_msc(data, fs, nfft=512, window_length_ms=window_length_ms, window_step_ms=window_step_ms, num_msc=13)
  1007. features.update(self.process_matrix(msc, 'MSC'))
  1008. # Other spectral features...
  1009. # (Similar to existing code in extract_spectral_features method)
  1010. except Exception as e:
  1011. self.logger.error(f"Error processing spectral features for {identifier}: {e}")
  1012. elif feature_type == 'complexity':
  1013. # Extract complexity features from data
  1014. try:
  1015. # Higuchi Fractal Dimension
  1016. HFD = calculate_hfd_per_frame(data, fs, window_length_ms, window_step_ms, 10)
  1017. features.update(self.process_row(HFD, 'HFD'))
  1018. # Other complexity features...
  1019. # (Similar to existing code in extract_complexity_features method)
  1020. except Exception as e:
  1021. self.logger.error(f"Error processing complexity features for {identifier}: {e}")
  1022. elif feature_type == 'frequency':
  1023. # Extract frequency features from data
  1024. try:
  1025. # Fundamental Frequency (F0)
  1026. F0 = get_pitch(data, fs, window_length_ms, window_step_ms)
  1027. F0_valid = F0[~np.isnan(F0)]
  1028. features.update(self.process_row(F0_valid, 'F0'))
  1029. # Other frequency features...
  1030. # (Similar to existing code in extract_frequency_features method)
  1031. except Exception as e:
  1032. self.logger.error(f"Error processing frequency features for {identifier}: {e}")
  1033. elif feature_type == 'intensity':
  1034. # Extract intensity features from data
  1035. try:
  1036. # Root Mean Square amplitude
  1037. RMS = rms_amplitude(data, fs, window_length_ms, window_step_ms)
  1038. features.update(self.process_row(RMS, 'RMS'))
  1039. # Other intensity features...
  1040. # (Similar to existing code in extract_intensity_features method)
  1041. except Exception as e:
  1042. self.logger.error(f"Error processing intensity features for {identifier}: {e}")
  1043. elif feature_type == 'voice_quality':
  1044. # Extract voice quality features from data
  1045. try:
  1046. # Shimmer
  1047. SHIMMER = analyze_audio_shimmer(data, fs, window_length_ms, window_step_ms)
  1048. features.update(self.process_row(SHIMMER, 'SHIMMER'))
  1049. # Other voice quality features...
  1050. # (Similar to existing code in extract_voice_quality_features method)
  1051. except Exception as e:
  1052. self.logger.error(f"Error processing voice quality features for {identifier}: {e}")
  1053. return features
  1054. except Exception as e:
  1055. self.logger.error(f"Error processing {feature_type} features from audio data: {e}")
  1056. return {}
  1057. # Process each feature type
  1058. for feature_type in features_to_calculate:
  1059. if feature_type == 'all':
  1060. # Handle 'all' feature type
  1061. all_features = process_audio_with_feature('all')
  1062. result.update(all_features)
  1063. break
  1064. else:
  1065. # Handle specific feature type
  1066. features = process_audio_with_feature(feature_type)
  1067. result.update(features)
  1068. self.logger.info("Extracted %d features total for %s", len(result), identifier)
  1069. return result
  1070. def _organize_features_by_group(self, features: Dict[str, Any]) -> Dict[str, Dict[str, Any]]:
  1071. """
  1072. Organize features by their group
  1073. Parameters:
  1074. -----------
  1075. features : Dict[str, Any]
  1076. Dictionary of features
  1077. Returns:
  1078. --------
  1079. Dict[str, Dict[str, Any]]
  1080. Dictionary of features organized by group
  1081. """
  1082. if not features:
  1083. return {}
  1084. # Group features by their category
  1085. feature_groups = self._group_features(list(features.keys()))
  1086. # Create output dictionary organized by groups
  1087. grouped_features = {}
  1088. for group_name, feature_names in feature_groups.items():
  1089. group_dict = {}
  1090. for name in feature_names:
  1091. if name in features:
  1092. group_dict[name] = features[name]
  1093. if group_dict: # Only include groups with features
  1094. grouped_features[group_name] = group_dict
  1095. # Check for any ungrouped features and add them to 'Other'
  1096. all_grouped_names = [name for group in feature_groups.values() for name in group]
  1097. ungrouped_features = {name: value for name, value in features.items()
  1098. if name not in all_grouped_names}
  1099. if ungrouped_features:
  1100. if 'Other' not in grouped_features:
  1101. grouped_features['Other'] = {}
  1102. grouped_features['Other'].update(ungrouped_features)
  1103. return grouped_features
  1104. def _flatten_grouped_features(self, grouped_features: Dict[str, Dict[str, Any]]) -> Dict[str, Any]:
  1105. """
  1106. Flatten a dictionary of grouped features into a single-level dictionary
  1107. Parameters:
  1108. -----------
  1109. grouped_features : Dict[str, Dict[str, Any]]
  1110. Dictionary of features organized by group
  1111. Returns:
  1112. --------
  1113. Dict[str, Any]
  1114. Flattened dictionary of features
  1115. """
  1116. flattened = {}
  1117. for group_name, group_dict in grouped_features.items():
  1118. for feature_name, feature_value in group_dict.items():
  1119. flattened[feature_name] = feature_value
  1120. return flattened
  1121. def _save_features_to_csv(self, features: Dict[str, Any], csv_path: str) -> None:
  1122. """
  1123. Save extracted features to a CSV file
  1124. Parameters:
  1125. -----------
  1126. features : Dict[str, Any]
  1127. Dictionary of extracted features
  1128. csv_path : str
  1129. Path to save the CSV file
  1130. """
  1131. try:
  1132. with open(csv_path, 'w', newline='') as csvfile:
  1133. writer = csv.writer(csvfile)
  1134. writer.writerow(['Feature', 'Value'])
  1135. for feature_name, feature_value in features.items():
  1136. # Handle different types of feature values
  1137. if isinstance(feature_value, (int, float)) or feature_value is None:
  1138. writer.writerow([feature_name, feature_value])
  1139. elif isinstance(feature_value, (list, np.ndarray)):
  1140. # For arrays, convert to string representation
  1141. if isinstance(feature_value, np.ndarray):
  1142. # Convert to Python list for better CSV compatibility
  1143. feature_list = feature_value.tolist()
  1144. else:
  1145. feature_list = feature_value
  1146. # Limit array size in CSV to prevent excessive file sizes
  1147. if len(feature_list) > 100:
  1148. self.logger.warning(f"Feature {feature_name} has {len(feature_list)} elements, truncating to 100 for CSV output")
  1149. feature_list = feature_list[:100]
  1150. # Join array elements with commas for CSV
  1151. writer.writerow([feature_name, str(feature_list)])
  1152. else:
  1153. # For other types, use string representation
  1154. writer.writerow([feature_name, str(feature_value)])
  1155. self.logger.info(f"Saved features to {csv_path}")
  1156. except Exception as e:
  1157. self.logger.error(f"Error saving features to CSV {csv_path}: {e}")
  1158. def _group_features(self, feature_names: List[str]) -> Dict[str, List[str]]:
  1159. """
  1160. Group features by their category based on name patterns
  1161. Parameters:
  1162. -----------
  1163. feature_names : List[str]
  1164. List of feature names to group
  1165. Returns:
  1166. --------
  1167. Dict[str, List[str]]
  1168. Dictionary mapping category names to lists of feature names
  1169. """
  1170. groups = {}
  1171. # Define common feature prefixes and their corresponding groups
  1172. # More specific prefixes first
  1173. prefix_groups = {
  1174. 'LOG_MEL_SPECTROGRAM': 'Spectral/Mel',
  1175. 'MFCC': 'Spectral/Cepstral',
  1176. 'LPCC': 'Spectral/Cepstral',
  1177. 'LPC': 'Spectral/LPC',
  1178. 'LSP': 'Spectral/LSP',
  1179. 'PLP': 'Spectral/PLP',
  1180. 'F0': 'Pitch',
  1181. 'F1': 'Formants',
  1182. 'F2': 'Formants',
  1183. 'F3': 'Formants',
  1184. 'Jitter': 'Voice Quality/Perturbation',
  1185. 'Shimmer': 'Voice Quality/Perturbation',
  1186. 'APQ': 'Voice Quality/Perturbation',
  1187. 'HNR': 'Voice Quality/Harmonicity',
  1188. 'NHR': 'Voice Quality/Harmonicity',
  1189. 'HARMONICITY': 'Voice Quality/Harmonicity',
  1190. 'CPP': 'Voice Quality/CPP',
  1191. 'HAMMARBERG_INDEX': 'Voice Quality/Spectral Tilt',
  1192. 'ALPHA_RATIO': 'Voice Quality/Spectral Tilt',
  1193. 'MSC': 'Spectral/Modulation',
  1194. 'CENTRIODS': 'Spectral/Shape',
  1195. 'LTAS': 'Spectral/Shape',
  1196. 'ENVELOPE': 'Spectral/Shape',
  1197. 'RMS': 'Intensity/Amplitude',
  1198. 'PEAK': 'Intensity/Amplitude',
  1199. 'Amplitude_Range': 'Intensity/Amplitude',
  1200. 'SPL': 'Intensity/SPL',
  1201. 'STE': 'Intensity/Energy',
  1202. 'INTENSITY': 'Intensity/Energy',
  1203. 'HFD': 'Complexity',
  1204. 'FREQ_ENTROPY': 'Complexity',
  1205. 'AMP_ENTROPY': 'Complexity',
  1206. 'ZCR': 'Complexity',
  1207. 'voiceProb': 'Speech Activity',
  1208. 'relative_sentence_duration': 'Speech Fluency/Timing',
  1209. 'regularity': 'Speech Fluency/Regularity',
  1210. 'PVI': 'Speech Fluency/Regularity'
  1211. # Add more specific groups as needed
  1212. }
  1213. # Assign each feature to a group
  1214. for name in feature_names:
  1215. assigned = False
  1216. # Check specific prefixes first
  1217. for prefix, group in prefix_groups.items():
  1218. # Match start or common statistical patterns
  1219. if name.startswith(prefix) or \
  1220. f'_{prefix}_sma' in name or \
  1221. f'_{prefix}_de' in name:
  1222. if group not in groups:
  1223. groups[group] = []
  1224. groups[group].append(name)
  1225. assigned = True
  1226. break
  1227. # Assign to broader category if no specific match
  1228. if not assigned:
  1229. broad_category = 'Other'
  1230. if 'spectral' in name.lower() or 'spec' in name.lower(): broad_category = 'Spectral/Other'
  1231. elif 'voice' in name.lower() or 'hnr' in name.lower() or 'jitter' in name.lower() or 'shimmer' in name.lower(): broad_category = 'Voice Quality/Other'
  1232. elif 'freq' in name.lower() or 'pitch' in name.lower() or 'formant' in name.lower(): broad_category = 'Frequency/Other'
  1233. elif 'intens' in name.lower() or 'loud' in name.lower() or 'amp' in name.lower() or 'spl' in name.lower() or 'peak' in name.lower() or 'rms' in name.lower(): broad_category = 'Intensity/Other'
  1234. elif 'complex' in name.lower() or 'entropy' in name.lower() or 'hfd' in name.lower(): broad_category = 'Complexity/Other'
  1235. elif 'fluency' in name.lower() or 'rhythm' in name.lower() or 'pause' in name.lower() or 'speech' in name.lower(): broad_category = 'Timing/Fluency/Other'
  1236. if broad_category not in groups:
  1237. groups[broad_category] = []
  1238. groups[broad_category].append(name)
  1239. return groups

AcousticFeatureExtractor.py at commit 0707335, no license · at the source

Overview

Authors: Maryam Zolnoori1,2,3,4,5, Elyas Esmaeili6, Mehdi Naserian6, Ali Zolnour6, Sina Rashidi6, Tahoura Morovati6, Hossein Azadmaleki6, Zhihong Zhang3, James M. Noble7, Margaret V. McDonald4
ORCID iDs: Maryam Zolnoori
  1. Columbia University Irving Medical Center,New York, NY 10027 USA
  2. School of Nursing, Columbia University,New York, NY 10027 USA
  3. Data Science Institute, Columbia University,New York, NY 10027 USA
  4. Center for Home Care Policy & Research, VNS Health,New York, NY 10017 USA
  5. Columbia University Irving Medical Center,560 W 168 St, New York, NY 10032 USA
  6. Independent Researcher, New York, USA
  7. Department of Neurology, Taub Institute for Research on Alzheimer’s Disease and the Aging Brain, GH Sergievsky Center, Columbia University,New York, NY 10032 USA
Institutions: Columbia University Irving Medical Center (United States); Columbia University (United States); VNS Health (United States)
Journal: Health information science and systems, volume 14, issue 1, article 77
Dates: received 8 March 2025; accepted 17 June 2026; published online 24 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1007/s13755-026-00468-5 · PMID 42502354 · PMCID PMC13400534 · OpenAlex W7170325869
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing, Spectral & time-frequency, Connectivity, Complexity
Topic: Dementia and Cognitive Impairment Research (Psychiatry and Mental health, Medicine), according to OpenAlex
Funding: National Institute on Aging (K99AG076808)
Citations: not cited yet (Europe PMC); 145 references in the paper

Abstract

Background: Early detection of cognitive impairment remains a critical public health challenge. While biomarkers such as neuroimaging and cerebrospinal fluid analyses offer high sensitivity, their limited accessibility hampers widespread screening, especially in underserved settings. Speech-based markers have emerged as promising, noninvasive indicators of cognitive decline.

Objective: To develop and validate SpeechDETECT, an end-to-end speech-processing pipeline that captures fine-grained acoustic and temporal markers of cognitive impairment and provides interpretable outputs suitable for large-scale screening.

Methods: SpeechDETECT comprises six modules: (1) noise reduction / amplitude normalization; (2) an eight-domain voice-analysis framework (e.g., frequency parameters, speech fluency); (3) 50 ms segment-level feature extraction; (4) feature visualization; (5) dimensionality reduction / selection (Joint Mutual Information Maximization, LassoNet, PCA); and (6) classifier training with SHapley Additive exPlanations (SHAP). Performance was benchmarked against six acoustic toolkits (e.g., GeMAPS) on two English datasets: the DementiaBank Pitt corpus (train = 166, test = 71) with single cookie-theft picture description task and NIA PREPARE Phase 2 corpus (train = 1 064, test = 267) with multiple speech tasks.

Results: A Multi-Layer Perceptron trained on PCA-derived SpeechDETECT features achieved an F1-score = 0.81% and AUC-ROC = 0.80 on the Pitt test set, outperforming the best competing toolkit (AUC = 0.76). On the PREPARE test set—comprising ≤ 30 s recordings from four speech tasks—the same model attained F1 ≈ 0.67% and AUC-ROC = 0.70, demonstrating good generalizability. Cumulative-gains analysis showed that screening the top 40% of ranked participants captured ~ 70% of cognitively-impaired (CI) cases in Pitt and ~ 63% in PREPARE. SHAP revealed speech-fluency metrics (hesitation rate, pause ratio) and high-frequency formant dynamics as the most discriminative features.

Conclusion: SpeechDETECT delivers accurate (AUC up to 0.80) and interpretable detection of early cognitive impairment across both structured and multi-task speech settings. Its fully automated, domain-informed approach enables scalable, speech-based screening and provides a foundation for multimodal systems that combine acoustic markers with clinical or biomarker data to further improve diagnostic precision. The SpeechDETECT toolkit is openly available on GitHub at https://github.com/SpeechCARE/SpeechDETECT-Toolkit for researchers and clinicians. A demo tutorial video showing pipeline usage is available at https://github.com/SpeechCARE/SpeechDETECT-Toolkit/blob/main/SpeechDETECT.mp4.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

SpeechCARE/SpeechDETECT-Toolkit

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 0707335a7e153736c9c06af7b1a1a8b0779ef6d1, 2 August 2026
Languages: Python (13)
Size: 18 files, 13 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (requirements.txt, setup.py)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (9 files), SciPy (7 files), Matplotlib (1 file), PyTorch (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
14 files

Code availability

The codes used for data analysis is publicly available here: https://github.com/SpeechCARE/SpeechDETECT-Toolkit.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 13 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The dataset used for this study is from DementiaBank dataset, the largest publicly available benchmark with audio recordings for the picture-description task.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 1 funder, 93 references.

Cite

This paper

Zolnoori, M., Esmaeili, E., Naserian, M., Zolnour, A., Rashidi, S., Morovati, T., Azadmaleki, H., Zhang, Z., Noble, J. M., & McDonald, M. V. (2026). SpeechDETECT: an explainable automated speech processing pipeline for early detection of neurological and health changes. Health information science and systems, 14(1), 77. https://doi.org/10.1007/s13755-026-00468-5

BibTeX

@article{zolnoori2026speechdetect,
author = {Zolnoori, Maryam and Esmaeili, Elyas and Naserian, Mehdi and Zolnour, Ali and Rashidi, Sina and Morovati, Tahoura and Azadmaleki, Hossein and Zhang, Zhihong and Noble, James M. and McDonald, Margaret V.},
title = {{SpeechDETECT: an explainable automated speech processing pipeline for early detection of neurological and health changes}},
journal = {Health information science and systems},
year = {2026},
month = jul,
volume = {14},
number = {1},
pages = {77},
publisher = {Springer},
issn = {2047-2501},
doi = {10.1007/s13755-026-00468-5},
url = {https://doi.org/10.1007/s13755-026-00468-5},
pmid = {42502354},
pmcid = {PMC13400534}
}

RIS

TY - JOUR
AU - Zolnoori, Maryam
AU - Esmaeili, Elyas
AU - Naserian, Mehdi
AU - Zolnour, Ali
AU - Rashidi, Sina
AU - Morovati, Tahoura
AU - Azadmaleki, Hossein
AU - Zhang, Zhihong
AU - Noble, James M.
AU - McDonald, Margaret V.
TI - SpeechDETECT: an explainable automated speech processing pipeline for early detection of neurological and health changes
T2 - Health information science and systems
J2 - Health Inf Sci Syst
PY - 2026
DA - 2026/07/24
VL - 14
IS - 1
SP - 77
SN - 2047-2501
PB - Springer
DO - 10.1007/s13755-026-00468-5
UR - https://doi.org/10.1007/s13755-026-00468-5
LA - en
ER -

CSL-JSON

{
"id": "10.1007/s13755-026-00468-5",
"type": "article-journal",
"title": "SpeechDETECT: an explainable automated speech processing pipeline for early detection of neurological and health changes",
"container-title": "Health information science and systems",
"author": [
{
"family": "Zolnoori",
"given": "Maryam"
},
{
"family": "Esmaeili",
"given": "Elyas"
},
{
"family": "Naserian",
"given": "Mehdi"
},
{
"family": "Zolnour",
"given": "Ali"
},
{
"family": "Rashidi",
"given": "Sina"
},
{
"family": "Morovati",
"given": "Tahoura"
},
{
"family": "Azadmaleki",
"given": "Hossein"
},
{
"family": "Zhang",
"given": "Zhihong"
},
{
"family": "Noble",
"given": "James M."
},
{
"family": "McDonald",
"given": "Margaret V."
}
],
"container-title-short": "Health Inf Sci Syst",
"volume": "14",
"issue": "1",
"page": "77",
"DOI": "10.1007/s13755-026-00468-5",
"PMID": "42502354",
"PMCID": "PMC13400534",
"ISSN": "2047-2501",
"publisher": "Springer",
"URL": "https://doi.org/10.1007/s13755-026-00468-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
24
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1002/alz.71365 [code]
Benchmarking speech biomarkers of Alzheimer's against cognitive and neural measures.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: SciPy, Matplotlib, NumPy, 2 references
[2] doi:10.1038/s41514-026-00390-w [code]
Mild cognitive impairment cases affect the predictive power of Alzheimer's disease diagnostic models using routine clinical variables.
Journal: npj aging
In common: SciPy, Matplotlib, NumPy, 2 references
[3] doi:10.1002/advs.77003 [code]
SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[4] doi:10.3390/s26103065 [code]
Subject-Wise Depression Screening from Eight-Channel Resting-State EEG Using Asymmetry-Aware Spectral Features and Connectivity Ablation.
Journal: Sensors (Basel, Switzerland)
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[5] doi:10.1371/journal.pgen.1012242 [code]
Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.
Journal: PLoS genetics
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[6] doi:10.1371/journal.pone.0351872 [code]
Decoding visual object recognition from EEG signals.
Journal: PloS one
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[7] doi:10.1038/s41398-026-04101-7 [code]
Using deep learning to identify brain networks mediating cognitive and motor impairments in alcohol use disorder.
Journal: Translational psychiatry
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[8] doi:10.1002/mrm.70431 [code]
SelExNet: A Self-Supervised Physics-Informed Framework for Multi-Channel Joint RF and Gradient Waveform Optimization in 2D Spatially Selective Excitation.
Journal: Magnetic resonance in medicine
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[9] doi:10.3389/fsysb.2026.1873899 [code]
A systems microbiology framework for reproducible multi-dataset omics integration with application to long COVID.
Journal: Frontiers in systems biology
In common: PyTorch, SciPy, Matplotlib, 1 other tool, 1 reference
[10] doi:10.1093/nar/gkag621 [code]
Optimal gene panel selection for targeted spatial transcriptomics experiments.
Journal: Nucleic acids research
In common: PyTorch, Matplotlib, NumPy, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.