OSCR

TweetyBERT: Automated parsing of birdsong through self-supervised machine learning.

Code ↔ Paper

The paper beside its authors' code: matches between them have not been computed for this paper yet.

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Markdown · 431 lines · 25 KB · CC-BY-4.0

  1. # TweetyBERT: Automated Parsing of Birdsong Through Self-Supervised Machine Learning
  2. This repository contains the code and instructions to replicate the results and figures presented in the paper [TweetyBERT: Automated parsing of birdsong through self-supervised machine learning (bioRxiv preprint)](https://www.biorxiv.org/content/10.1101/2025.04.09.648029v1)
  3. ## Tips + Data Locations For Replicating Figures:
  4. - Detailed instructions for generating the figures in the paper are found at the bottom of the readme.md
  5. - /LLB_Embeddings contain npz file embeddings for reconstructing figure 3 TweetyBERT and spectogram UMAPs
  6. - /LLB_Fold_Data contains the npzs used to derive the v-measure scores + extended metrics for figure 4 and 5
  7. - /results contains the data required to reconstruct figure 6 linear probe / finetuning experiments, also contains seasonality analysis and figure 5 extended metrics raw data (folder called proxy metrics)
  8. - /Seasonality Embeddings contains the npz files used to construct the embeddings for seasonality comaprisons (figure 7)
  9. - Path for the TweetyNET song detection json: files/contains_llb.json Path for the seasonality song detection json: files/contains_seasonality.json
  10. - TweetyNET dataset can be found here: https://doi.org/10.5061/dryad.xgxd254f4 (Use the .wav raw audio recordings rather than .npz files, use the provided song detection json (it contains syllable annotations too))
  11. ## 🐦 TweetyBERT Overview
  12. TweetyBERT combines a convolutional front-end with a transformer architecture to learn representations of bird vocalizations. The model can be used for:
  13. - Automated/Unsupervised labeling of songbird syllables
  14. - Comparing embeddings before/after perturbation
  15. - Visualizing song with dimensionality reduction
  16. For questions or collaboration inquiries, please email: georgev [at] Uoregon.edu
  17. ## 🔧 Repository Structure
  18. The repository is organized as follows:
  19. ```
  20. tweety_bert/
  21. ├── readme.md # This file
  22. ├── pretrain.py # Python script for pretraining TweetyBERT
  23. ├── decoding.py # Python script for UMAP generation and decoder training
  24. ├── run_inference.py # Python script for running inference
  25. ├── figure_generation_scripts/ # Scripts for generating paper figures
  26. ├── scripts/ # Helper scripts and utilities for data processing and analysis
  27. ├── shell_scripts/ # Shell scripts for automation (Alternative workflow / deprecated)
  28. ├── src/ # Core model implementation and primary codebase
  29. │ ├── model.py # Defines the TweetyBERT model architecture
  30. │ ├── spectogram_generator.py # Generates spectrograms from WAV files
  31. │ ├── trainer.py # Handles model pretraining loop and metrics
  32. │ ├── decoder.py # Handles decoder training and saving
  33. │ ├── inference.py # Core script for running inference with a trained model
  34. │ ├── linear_probe.py # Implements linear probe model and trainer
  35. │ ├── analysis.py # Contains functions for UMAP plotting and performance metrics
  36. │ ├── data_class.py # Defines Dataset and Dataloader classes
  37. │ └── utils.py # Utility functions (e.g., loading models, configs)
  38. ├── files/ # Stores NPZ files and JSON annotation databases [User must populate or adjust paths]
  39. ├── experiments/ # Stores model checkpoints and training logs [User must populate or adjust paths]
  40. ├── imgs/ # Stores generated images, plots, and visualizations [Output directory]
  41. └── results/ # Stores output data from model computations and analysis [Output directory]
  42. ├── detect_song.py # Python script for running song detection
  43. ```
  44. * **Root Directory:** Contains the main workflow scripts (`pretrain.py`, `decoding.py`, `run_inference.py`) and this README.
  45. * **`figure_generation_scripts/`**: Contains Python scripts specifically designed to reproduce the figures shown in the associated publication. Edit paths within these scripts as needed.
  46. * **`scripts/`**: A collection of utility Python scripts for various tasks like data conversion, splitting, merging, plotting specific metrics, etc.
  47. * **`shell_scripts/`**: Contains the original bash scripts for running workflows (now largely superseded by the root Python scripts). Kept for reference or alternative use cases.
  48. * **`src/`**: Holds the core Python source code for the TweetyBERT model, data handling, training, inference logic, and analysis functions.
  49. * **`files/`**: Intended location for input data like annotation files (`.json`) and embedding files (`.npz`). You will need to place your data here or modify paths in the scripts.
  50. * **`experiments/`**: Default location where trained model checkpoints (`.pth`), configuration files (`config.json`), and training logs (`training_statistics.json`, `train_files.txt`, `test_files.txt`) are saved.
  51. * **`imgs/`**: Default output directory for generated images, such as UMAP plots, spectrogram visualizations, and other figures.
  52. * **`results/`**: Default output directory for non-image results, like performance metrics (`.txt`, `.csv`).
  53. * **`detect_song.py`**: Python script for running song detection.
  54. ## 🚀 Installation & Environment Setup
  55. The following steps assume you have Conda installed and are using a CUDA-capable GPU (e.g., NVIDIA RTX 4090). Adjust as necessary for your system.
  56. ```bash
  57. # 1. Create and activate a new Conda environment
  58. conda create -n tweetybert python=3.11
  59. conda activate tweetybert
  60. # 2. Install core scientific packages (including librosa)
  61. conda install -c conda-forge \
  62. numpy \
  63. matplotlib \
  64. tqdm \
  65. umap-learn \
  66. hdbscan \
  67. scikit-learn \
  68. pandas \
  69. seaborn \
  70. jupyter \
  71. ipykernel \
  72. librosa
  73. # 3. Install additional dependencies via pip
  74. pip install soundfile shutil-extra glasbey pyqtgraph PyQt5 hmmlearn
  75. # 4. (Optional) Install PyTorch if not already installed (adjust CUDA version if needed)
  76. #
  77. # =====================
  78. # ⚠️ WARNING: Be VERY careful to match the correct CUDA version to your system and GPU drivers! ⚠️
  79. # - If you do NOT have an NVIDIA GPU or do not need GPU acceleration, install the CPU-only version:
  80. # pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu
  81. # - If you have a CUDA-capable GPU, install the version matching your CUDA toolkit (see https://pytorch.org/get-started/locally/)
  82. # - Mismatched CUDA versions can cause import errors or silent failures.
  83. # =====================
  84. # Example for CUDA 12.x:
  85. pip install torch torchvision torchaudio
  86. # 5. Clone this repository
  87. git clone https://github.com/georgevenven/tweety_bert.git
  88. cd tweety_bert
  89. ```
  90. **Important Notes:**
  91. * Ensure you have Python 3.11+ and PyTorch >= 2.0 installed.
  92. * The primary workflow now uses the Python scripts (`pretrain.py`, `decoding.py`, `run_inference.py`) located in the root directory.
  93. * The shell scripts in `shell_scripts/` provide an alternative workflow.
  94. ## 💾 Storage Requirements
  95. Depending on the size of your audio dataset, storage requirements can range from **50 GB** to **1 TB** or more. Ensure you have sufficient disk space.
  96. ## ⚡ GPU & Training Times
  97. Pretraining can take 3-4 hours on a single NVIDIA RTX 4090 GPU, and 100s of hours on a CPU, depending on your dataset size and hyperparameters.
  98. ## 🎶 Song Detection & JSON Format
  99. TweetyBERT uses a JSON file to identify segments of bird song within audio recordings. This file can also optionally include syllable labels for validation.
  100. **Example JSON Structure:**
  101. ```json
  102. {
  103. "filename": "bird_XXXX_YYYY_MM_DD_HH_MM_SS.wav",
  104. "song_present": true,
  105. "segments": [
  106. {
  107. "onset_timebin": 100,
  108. "offset_timebin": 500,
  109. "onset_ms": 1234.56,
  110. "offset_ms": 5678.90
  111. }
  112. ],
  113. "spec_parameters": {
  114. "step_size": 119,
  115. "nfft": 1024
  116. },
  117. "syllable_labels": {
  118. "1": [
  119. [1.00, 2.50]
  120. ]
  121. }
  122. }
  123. ```
  124. * **`filename`**: Name of the WAV file.
  125. * **`song_present`**: Boolean indicating if song is detected.
  126. * **(Important):** Even if `song_present` is true, the file might be skipped later if the `segments` list is empty or contains only very short segments.
  127. * **`segments`**: List of detected song segments with onset/offset times (in timebins and milliseconds).
  128. * **`spec_parameters`**: Parameters used for spectrogram generation (e.g., `step_size`, `nfft`).
  129. * **`syllable_labels` (optional)**: Time intervals for each labeled syllable, keyed by label ID.
  130. ### Generating the Song Detection JSON (Recommended)
  131. This repository includes a wrapper script `detect_song.py` to simplify the process of generating the required JSON file using the [TweetyNet Song Detector](https://github.com/georgevenven/tweety_net_song_detector).
  132. 1. **Prerequisite:** Ensure `git` is installed on your system.
  133. 2. **Navigate:** Go to the root directory of the `tweety_bert` repository.
  134. 3. **Run:** Execute the `detect_song.py` script, providing the path to your WAV files.
  135. **Example:**
  136. ```bash
  137. python detect_song.py --input_dir "/path/to/your/wav/files"
  138. ```
  139. * The first time you run it, the script will automatically clone the detector repository into a local folder named `song_detector`.
  140. * It will then process all `.wav` files found in the specified `--input_dir` (including subdirectories).
  141. * The output is a single JSON file (defaulting to `files/song_detection.json`) containing entries for each processed WAV file, indicating detected song segments.
  142. * **Use the path to this generated JSON file** for the `--song_detection_json_path` argument when running `pretrain.py`, `decoding.py`, or `run_inference.py`.
  143. ## 🏋️ Training the Model (Pretraining)
  144. To pretrain TweetyBERT using the Python script:
  145. 1. Navigate to the root directory of the repository (`tweety_bert/`).
  146. 2. Run `pretrain.py` with appropriate arguments. Key arguments include:
  147. * `--input_dir`: Path to the folder containing WAV files.
  148. * `--song_detection_json_path`: Path to the song detection JSON (optional, uses internal detection if not provided).
  149. * `--experiment_name`: Name for your training run (e.g., "MyTweetyBERTModel"). Results saved to `experiments/<experiment_name>`.
  150. * `--test_percentage`: Percentage of data for the test set (default: 20).
  151. * `--batch_size`, `--learning_rate`, `--context`, `--m`, etc.: Model and training hyperparameters. Use `--help` to see all options.
  152. **Example:**
  153. ```bash
  154. python pretrain.py \
  155. --input_dir "/path/to/your/wav/files" \
  156. --song_detection_json_path "/path/to/your/song_detection.json" \
  157. --experiment_name "MyTweetyBERTModel" \
  158. --test_percentage 20 \
  159. --batch_size 42 \
  160. --learning_rate 3e-4 \
  161. --context 1000 \
  162. --m 250 \
  163. --multi_thread # Add this flag to use multi-threading for spec gen
  164. ```
  165. ## 💡 Generating Embeddings & Training a Decoder
  166. After pretraining, use `decoding.py` to generate UMAP embeddings.
  167. 1. Navigate to the root directory (`tweety_bert/`).
  168. 2. Run `decoding.py`. This script has two main modes controlled by `--mode`:
  169. * **`--mode single`**: Processes a single dataset (potentially a random subset).
  170. * **`--mode grouping`**: Processes data split into temporal groups (relies on `scripts/copy_files_from_wavdir_to_multiple_event_dirs.py` for grouping logic).
  171. 3. Key arguments:
  172. * `--mode`: `single` or `grouping`.
  173. * `--bird_name`: Descriptive name for UMAP/decoder outputs (e.g., "my_canary").
  174. * `--model_name`: Name of the pretrained experiment (must match `experiment_name` used in `pretrain.py`).
  175. * `--wav_folder`: Path to the bird's/dataset's WAV files.
  176. * `--song_detection_json_path`: Path to the detection JSON for these files.
  177. * `--num_samples_umap`: Number of samples for UMAP (e.g., "5e5").
  178. * `--num_random_files_spec` (for `--mode single`): Number of random WAVs to use for spectrogram generation.
  179. **Example (`single` mode):**
  180. ```bash
  181. python decoding.py \
  182. --mode single \
  183. --bird_name "my_canary_decoder" \
  184. --model_name "MyTweetyBERTModel" \
  185. --wav_folder "/path/to/this_birds/wav/files" \
  186. --song_detection_json_path "/path/to/this_birds/song_detection.json" \
  187. --num_random_files_spec 100 \
  188. --num_samples_umap 5e5
  189. ```
  190. **Example (`grouping` mode):**
  191. ```bash
  192. python decoding.py \
  193. --mode grouping \
  194. --bird_name "my_canary_grouped_decoder" \
  195. --model_name "MyTweetyBERTModel" \
  196. --wav_folder "/path/to/this_birds/wav/files" \
  197. --song_detection_json_path "/path/to/this_birds/song_detection.json" \
  198. --num_samples_umap 1e5
  199. ```
  200. *(Note: The grouping mode currently relies on the interaction with `scripts/copy_files_from_wavdir_to_multiple_event_dirs.py`, which is interactive).*
  201. <!--
  202. ## 🔍 Inference (Not Part of the TweetyBERT Paper)
  203. Run inference on new WAV files using a trained decoder.
  204. 1. Navigate to the root directory (`tweety_bert/`).
  205. 2. Run `run_inference.py`. Key arguments:
  206. * `--bird_name`: Name used when training the decoder (used to find the saved decoder state, e.g., "my_canary_decoder").
  207. * `--wav_folder`: Directory of new WAV files for inference.
  208. * `--song_detection_json`: Path to the detection JSON for these new files (optional, uses internal detection if not provided).
  209. * `--apply_post_processing`: Apply smoothing (True/False, default True).
  210. * `--window_size`: Smoothing window size (default 200).
  211. * `--visualize`: Generate output plots (True/False, default False).
  212. **Example:**
  213. ```bash
  214. python run_inference.py \
  215. --bird_name "my_canary_decoder" \
  216. --wav_folder "/path/to/new/wav/files" \
  217. --song_detection_json "/path/to/new/song_detection.json" \
  218. --apply_post_processing True \
  219. --visualize True
  220. ```
  221. The output will be a JSON database (`files/<bird_name>_decoded_database.json`) summarizing detected syllables. Visualizations (if enabled) are saved in `imgs/inference_specs_<bird_name>/`. -->
  222. ## 🗄️ NPZ File Format
  223. The `.npz` files used for embeddings and analysis generally contain the following arrays:
  224. | Array Name | Example Shape | Data Type | Description |
  225. | :-------------------- | :------------------- | :---------- | :-------------------------------------------------------------------------- |
  226. | `embedding_outputs` | `(N, 2)` | `float32` | 2D UMAP embedding coordinates for N timebins. |
  227. | `hdbscan_labels` | `(N,)` | `int64` | Cluster labels assigned by HDBSCAN for each timebin (-1 for noise). |
  228. | `ground_truth_labels` | `(N,)` | `int64` | Human-annotated syllable labels for each timebin. |
  229. | `predictions` | `(N, 196)` | `float32` | Raw output (neural activations) from a TweetyBERT layer before UMAP. |
  230. | `s` | `(N, 196)` | `float32` | Spectrogram data for the N timebins (frequency bins = 196). |
  231. | `hdbscan_colors` | `(C_hdbscan, 3)` | `float64` | RGB color values for each HDBSCAN cluster. |
  232. | `ground_truth_colors` | `(C_gt, 3)` | `float64` | RGB color values for each ground truth syllable class. |
  233. | `original_spectogram` | `(N, 196)` | `float32` | Original full spectrogram corresponding to the N timebins. |
  234. | `vocalization` | `(N,)` | `int64` | Binary array indicating if a timebin contains vocalization (1) or not (0). |
  235. | `file_indices` | `(N,)` | `int64` | Index mapping each timebin to its original source file in `file_map`. |
  236. | `dataset_indices` | `(N,)` | `int64` | Index indicating which dataset or group a timebin belongs to (e.g., for seasonality). |
  237. | `file_map` | `()` (scalar object) | `object` | A dictionary mapping integer file indices to actual file path strings. |
  238. *N = Total number of timebins in the NPZ file.*
  239. *C_hdbscan = Number of unique HDBSCAN clusters.*
  240. *C_gt = Number of unique ground truth syllable classes.*
  241. ## 📄 Regenerating Figures from the Paper
  242. The following instructions outline how to regenerate the figures presented in the paper.
  243. **Note:** You will need to adjust file paths within the scripts to point to your local data locations.
  244. **General Setup:**
  245. 1. Ensure your Conda environment (`tweetybert`) is activated.
  246. 2. Navigate to the root of the cloned `tweety_bert` repository.
  247. 3. Organize your data files (`.npz`, `.wav`, `.json`) as referenced by the scripts, or update the paths within the scripts accordingly.
  248. ---
  249. **Figures 1 & 2:**
  250. These are cartoon schematics, and their direct replication from code is not applicable. Figure 2's masked prediction visualizations can be conceptually generated using `figure_generation_scripts/masked_prediction_figure_generator.py`.
  251. <!-- - **To generate similar masked prediction examples:**
  252. 1. Run the `masked_prediction_figure_generator.py` script from the `tweety_bert` root directory. Provide the paths to your trained model directory, the directory containing spectrogram NPZ files for visualization, and the desired output directory using command-line arguments.
  253. ```bash
  254. python figure_generation_scripts/masked_prediction_figure_generator.py \
  255. --model-dir experiments/TweetyBERT_Paper_Yarden_Model \
  256. --data-dir [Path_to_your_data_here]/llb3_specs \
  257. --output-dir imgs/masked_predictions_for_figure_2 \
  258. --num-samples 10 # Optional: specify number of samples
  259. ```
  260. 2. The script will generate visualization images in the specified output directory. -->
  261. ---
  262. **Figure 3: TweetyBERT and Spectrogram UMAP Embeddings**
  263. * **Data:** Prepare NPZ files containing TweetyBERT embeddings and raw spectrogram embeddings, along with ground truth labels. Assume you place them in `files/LLB_Embedding_Paper/`.
  264. * **Figure 3B & 3C (UMAP plots):**
  265. 1. Run from the `tweety_bert` root directory, providing the input NPZ file and output directory as arguments:
  266. ```bash
  267. python figure_generation_scripts/UMAP_plots_from_npz.py \
  268. "placeholder_input_npz_dir" \
  269. "placeholder_output_dir"
  270. ```
  271. *(This script may open an interactive window for cropping if processing a single file.)*
  272. * **Figure 3A (Interactive UMAP region visualization):**
  273. 1. Run from the `tweety_bert` root directory, providing the input NPZ file as an argument:
  274. ```bash
  275. python figure_generation_scripts/visualizing_song_cluster_phase.py \
  276. "placeholder_embedding_npz_path"
  277. ```
  278. This script requires a GUI; use the Lasso tool in the interactive plot to select a UMAP region. Saved images appear in `imgs/selected_regions/`.
  279. *(Optional arguments like `--collage_mode`, `--max_length`, `--used_group_coloring` can be added.)*
  280. ---
  281. **Figure 4: Machine-derived vs. Human-derived Clusters**
  282. * **Data:** Prepare NPZ files from UMAP folds (e.g., in `files/LLB_Fold_Data_Paper/`). Each file should contain embeddings and labels for a data fold.
  283. * **Figure 4A & 4B (V-Measure Calculation and UMAP Plots):**
  284. 1. **Calculate V-Measure scores:**
  285. * Run from the `tweety_bert` root directory, providing the path to your fold data:
  286. ```bash
  287. python scripts/fold_v_measure_calculation.py "placeholder_dir_for_folds"
  288. ```
  289. This will print V-measure scores for each fold[cite: 113].
  290. 2. **Generate UMAP plots** (will plot both syllable and phrase labels for comparison):
  291. * Run from the `tweety_bert` root directory, providing the input NPZ file and output directory:
  292. ```bash
  293. python figure_generation_scripts/UMAP_plots_from_npz.py \
  294. "placeholder_input_npz_dir" \
  295. "placeholder_output_dir"
  296. ```
  297. *(This script may open an interactive window for cropping if processing a single file, you can close it)*.
  298. * **Figure 4C, 4D, 4E (Spectrograms with HDBSCAN and Ground Truth Labels):**
  299. 1. Run from the `tweety_bert` root directory, providing the path to an NPZ file from your fold data:
  300. ```bash
  301. # Example: Generate 10 random segments of default length (1000)
  302. python figure_generation_scripts/visualizing_hdb_scan_labels.py \
  303. --file_path "npz_file_path" # from which npz file you would like to generate examples \
  304. --output_dir "imgs/specs_plus_labels"
  305. ```
  306. This script generates spectrogram fragments (defaulting to 100 random ones if `--start_idx` is not provided). You will need to manually select one that clearly shows interesting phrases and mostly aligned labels, similar to the paper's figure.
  307. `--output_dir`.
  308. ---
  309. **Figure 5: Comparing Human and Automated Labels for Sequence Analysis**
  310. * **Figure 5A, 5B, 5C (Spectrogram with Spurious Insertions):**
  311. Generated similarly to Figure 4C,D,E using `figure_generation_scripts/visualizing_hdb_scan_labels.py`. You'll need to manually find/select a sample NPZ file and segment that exhibits significant spurious insertions by running the script with appropriate arguments.
  312. Example:
  313. ```bash
  314. # add no smoothing parameter
  315. python figure_generation_scripts/visualizing_hdb_scan_labels.py \
  316. --file_path "npz_file_path" # from which npz file you would like to generate examples \
  317. --output_dir "imgs/specs_plus_labels" \
  318. --no_smoothing
  319. ```
  320. * **Figure 5D, 5E, 5F (UMAP Evaluation and Smoothing Window Analysis):**
  321. 1. Run from the `tweety_bert` root directory, providing the path to the UMAP fold data directory and the desired output directory:
  322. ```bash
  323. python figure_generation_scripts/umap_eval.py \
  324. --folder_path "path_to_umap_eval" \
  325. --output_dir "/results/test"
  326. ```
  327. 2. This script will generate:
  328. * `all_windows_summary.txt` in the `output_dir`, containing statistics for each smoothing window and identifying the optimal one.
  329. * `metrics_by_window.png` (used for Fig 5F) in the `output_dir`.
  330. * Plots for Fig 5D and 5E (normalized confusion matrices) will be in the `output_dir/best_window/` subdirectory (e.g., `06_M_norm_fullreorder.png` and `04_M_norm_diag.png` - *note: the exact filenames might vary slightly based on internal plotting choices*).
  331. ---
  332. **Figure 6: Evaluating TweetyBERT Embeddings Using Linear Probes**
  333. * **Data Generation:**
  334. 1. Edit `scripts/linear_probe_automated_analysis.py`:
  335. * Ensure the list `experiment_configs` has the correct `experiment_path` for your pretrained TweetyBERT model (e.g., `"experiments/TweetyBERT_Paper_Yarden_Model"`).
  336. * Update `train_dir` and `test_dir` within `experiment_configs` to point to your linear probe datasets (e.g., for llb3, llb11, llb16). Adjust paths like `"/media/george-vengrovski/Desk SSD/TweetyBERT/linear_probe_dataset/{dataset}_train"`.
  337. * `results_path` is set to `"results"`. If you would like to train the linear probes from scratch, you will have to generate your own linear probe dataset, contact me for more details.
  338. 2. Run from the `tweety_bert` root directory:
  339. ```bash
  340. python scripts/linear_probe_automated_analysis.py
  341. ```
  342. This will create subdirectories in `results/` like `TweetyBERT_linear_probe_llb3/`, etc., containing `results.json`.
  343. * **Plot Generation:**
  344. 1. Open the Jupyter Notebook: `figure_generation_scripts/linear_probe_analysis.ipynb`.
  345. 2. In the first cell, update `base_path` to point to the parent directory of the results generated above (e.g., `base_path = 'results'`).
  346. 3. Run the first cell of the notebook.
  347. ---
  348. **Figure 7: Seasonal Vocal Plasticity in Canaries**
  349. * **Data:** You need NPZ files containing embeddings for different birds across breeding and non-breeding seasons (e.g., `5494_Seasonality_Final.npz`, `5508_Seasonality_Final.npz`). Place these files in a suitable location, for example, `files/seasonality_embeddings/`.
  350. * **Plot Generation:**
  351. 1. Run the `umap_comparison_figure_generation.py` script from the `tweety_bert` root directory. Provide the desired output directory *first*, followed by the paths to one or more NPZ files containing the seasonality embeddings.
  352. ```bash
  353. python figure_generation_scripts/umap_comparison_figure_generation.py \
  354. "imgs/seasonality_analysis_fig7" \
  355. "files/seasonality_embeddings/5494_Seasonality_Final.npz" \
  356. "files/seasonality_embeddings/5508_Seasonality_Final.npz"
  357. ```
  358. 2. The script will generate various plots in the specified output directory (e.g., `imgs/seasonality_analysis_fig7/`), creating subdirectories for each bird ID found in the NPZ filenames (e.g., `bird_5494/`, `bird_5508/`). The specific PNG files used for the paper are titled `bird_{bird_id}_all_before_vs_all_after_overlap.png`.
  359. ---

readme.md, under CC-BY-4.0 · at the source

Overview

Authors: George Vengrovski1,2, Miranda R Hulsey-Vincent1,2, Melissa A Bemrose2, Timothy J Gardner1,2
  1. Institute of Neuroscience and Department of Biology, University of Oregon, Eugene, OR, USA
  2. Phil and Penny Knight Campus for Accelerating Scientific Impact, University of Oregon, Eugene, OR, USA
Institutions: University of Oregon (United States)
Journal: Patterns (New York, N.Y.), volume 7, issue 4, article 101491
Dates: received 17 June 2025; accepted 30 December 2025; published online 3 March 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.patter.2025.101491 · PMID 42005386 · PMCID PMC13083638 · OpenAlex W7133297113
Open access: gold, a free copy (OpenAlex)
Status: code verified
Methods: Smoothing, state filtering, decompositions, Spectral & time-frequency, Connectivity, Statistics, Machine learning
Keywords: self-supervised, transformers, birdsong, deep learning, time series analysis, dimensionality reduction, clustering, song, animal vocalization, communication, bioacoustics
Topic: Animal Vocal Communication and Behavior (Developmental Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NINDS NIH HHS (R01 NS118424); National Institutes of Health (R01NS118424)
Citations: cited by 4 papers (Europe PMC); 86 references in the paper

Abstract

Deep neural networks can be trained to parse animal vocalizations—serving to identify the units of communication and annotating sequences of vocalizations for subsequent statistical analysis. However, current methods rely on human-labeled data for training. The challenge of parsing animal vocalizations in a fully unsupervised manner remains an open problem. Addressing this challenge, we introduce TweetyBERT, a self-supervised transformer neural network developed for the analysis of birdsong. The model is trained to predict masked or hidden fragments of audio but is not exposed to human supervision or labels. Applied to canary song, TweetyBERT autonomously learns the behavioral units of song, such as notes, syllables, and phrases—capturing intricate acoustic and temporal patterns. This approach of developing self-supervised models specifically tailored to animal communication may significantly accelerate the analysis of unlabeled vocal data.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above.

Zenodo 15391040

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 7 files
Software Heritage: not checked
Found in: “Data and code availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
1 file
At the source:

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 0 scripts, each with its path and the digest of its content;
  • no match between paragraphs and code yet;
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data and code availability

The steps to replicate findings, figures, and numerical results are available at Zenodo (https://doi.org/10.5281/zenodo.15391040).85 The TweetyNET dataset used for model training and ground-truth labels is publicly available via Dryad (https://doi.org/10.5061/dryad.xgxd254f4).86

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 11 keywords, 2 funders, 67 references.

Cite

This paper

Vengrovski, G., Hulsey-Vincent, M. R., Bemrose, M. A., & Gardner, T. J. (2026). TweetyBERT: Automated parsing of birdsong through self-supervised machine learning. Patterns (New York, N.Y.), 7(4), 101491. https://doi.org/10.1016/j.patter.2025.101491

BibTeX

@article{vengrovski2026tweetybert,
author = {Vengrovski, George and Hulsey-Vincent, Miranda R and Bemrose, Melissa A and Gardner, Timothy J},
title = {{TweetyBERT: Automated parsing of birdsong through self-supervised machine learning}},
journal = {Patterns (New York, N.Y.)},
year = {2026},
month = mar,
volume = {7},
number = {4},
pages = {101491},
publisher = {Elsevier},
issn = {2666-3899},
doi = {10.1016/j.patter.2025.101491},
url = {https://doi.org/10.1016/j.patter.2025.101491},
pmid = {42005386},
pmcid = {PMC13083638}
}

RIS

TY - JOUR
AU - Vengrovski, George
AU - Hulsey-Vincent, Miranda R
AU - Bemrose, Melissa A
AU - Gardner, Timothy J
TI - TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
T2 - Patterns (New York, N.Y.)
J2 - Patterns (N Y)
PY - 2026
DA - 2026/03/03
VL - 7
IS - 4
SP - 101491
SN - 2666-3899
PB - Elsevier
DO - 10.1016/j.patter.2025.101491
UR - https://doi.org/10.1016/j.patter.2025.101491
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.patter.2025.101491",
"type": "article-journal",
"title": "TweetyBERT: Automated parsing of birdsong through self-supervised machine learning",
"container-title": "Patterns (New York, N.Y.)",
"author": [
{
"family": "Vengrovski",
"given": "George"
},
{
"family": "Hulsey-Vincent",
"given": "Miranda R"
},
{
"family": "Bemrose",
"given": "Melissa A"
},
{
"family": "Gardner",
"given": "Timothy J"
}
],
"container-title-short": "Patterns (N Y)",
"volume": "7",
"issue": "4",
"page": "101491",
"DOI": "10.1016/j.patter.2025.101491",
"PMID": "42005386",
"PMCID": "PMC13083638",
"ISSN": "2666-3899",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.patter.2025.101491",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
3
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1523/eneuro.0023-26.2026 [code]
Real-Time Segmentation and Classification of Birdsong Syllables for Learning Experiments.
Journal: eNeuro
In common: 7 references
[2] doi:10.1038/s41586-026-10528-1 [code]
A critical initialization for biological neural networks.
Journal: Nature
In common: 3 references
[3] doi:10.1371/journal.pone.0333158 [code]
Alpha-synuclein overexpression reduces neural activity within a basal ganglia vocal nucleus in a zebra finch model.
Journal: PloS one
In common: 2 references
[4] doi:10.1093/sleep/zsag022
Mamba-based deep learning approach for sleep staging on a wireless multimodal wearable system without electroencephalography.
Journal: Sleep
In common: 2 references
[5] doi:10.1038/s41467-026-74466-2 [code]
Neuromorphic hierarchical modular reservoirs.
Journal: Nature communications
In common: 2 references
[6] doi:10.1371/journal.pone.0356243 [code]
Functional organization and natural scene responses across mouse visual cortical areas revealed with encoding manifolds.
Journal: PloS one
In common: 2 references
[7] doi:10.1038/s41540-026-00813-0 [code]
Combining in vitro and in silico approaches to model neural tube patterning and isthmic organizer formation.
Journal: NPJ systems biology and applications
In common: 2 references
[8] doi:10.1016/j.xpro.2026.104804 [code]
Protocol for analyzing slow cortical dynamics in mouse neuronal recordings.
Journal: STAR protocols
In common: 2 references
[9] doi:10.3389/fnins.2026.1893619 [code]
Direct laser carbonization of parylene-C toward microelectrodes for &lt;i&gt;in vivo&lt;/i&gt; action potential detection.
Journal: Frontiers in neuroscience
In common: 2 references
[10] doi:10.1016/j.isci.2026.116945
Neurodegeneration in the olfactory system in Niemann Pick type C1 disease.
Journal: iScience
In common: 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.