PumpKin: A machine-learning pipeline for automatically tracking localized kinematics in freely moving C. elegans.
The 7 matches
- [1] § Materials and methods › Design of a custom Butterworth filter for high-frequency noise reduction ↔ library/pumpkin.py, lines 383–448 · score 0.82 · low pass Butterworth, signed magnitude, motion vectors, grinder points, signal, filter
- [2] § Results › Overview of the PumpKin pipeline ↔ library/pumpkin.py, lines 383–448 · score 0.74 · genetic strain, low pass filter, cutoff frequency, expected pumping rate, signal, food
- [3] § Results › Overview of the PumpKin pipeline ↔ library/pumpkin.py, lines 450–501 · score 0.74 · grinder motion signal, cutoff frequency, valid pumps, binning, smoothing, optional
- [4] § Materials and methods › Training the faster R-CNN networks ↔ library/training.py, lines 22–40 · score 0.71 · COCO v1, CNN model, FPN, torchvision, weights, Faster
- [5] § Materials and methods › Motion compensation for reduction of noise from body motion ↔ library/pumpkin.py, lines 187–234 · score 0.64 · affine transformation, matched feature, RANSAC, flow, motion, frame
- [6] § Materials and methods › Use of peristimulus time histogram (PSTH) to estimate a continuous pumping rate from event times ↔ library/pumpkin.py, lines 450–501 · score 0.63 · trough occurs, detected pump, histogram, binned, pumping rate, filtered
- [7] § Materials and methods › Training the faster R-CNN networks ↔ cropROI.ipynb, lines 18–79 · score 0.53 · ROI centered, pharyngeal bulb, cropped, FRCNN, track, videos
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 501 lines · 22 KB · MIT · 5 matches
- ################################################################################
- # pumpkin.py
- # Written by Erin Shappell for Lu Lab
- #
- # This module contains all functions used to generate a continuous pumping rate
- # estimate from an FRCNN-tracked grinder in freely moving C. elegans.
- #
- ################################################################################
- # Imports
- import re
- import os
- import cv2
- import numpy as np
- from scipy.signal import iirfilter, find_peaks, filtfilt
- from scipy.ndimage import gaussian_filter1d
- ################################################################################
- def sample_dark_pts(frame, point, box_size=(15, 15), grid_size=(3, 3)):
- """
- Sample the darkest pixels from a grid of subregions within a box centered at a given point.
- Inputs:
- frame (ndarray): Input video frame (color image as a NumPy array).
- point (tuple): (x, y) coordinates representing the center of the sampling box.
- box_size (tuple): (width, height) of the box around the center within which to sample.
- Default is (15, 15).
- grid_size (tuple): (rows, cols) specifying how the box is divided into subregions.
- Default is (3, 3).
- Output:
- list: A list of (x, y) tuples corresponding to the darkest pixel in each subregion,
- adjusted to the coordinates of the original frame.
- Returns an empty list if the region is invalid.
- """
- cx, cy = map(int, point) # Ensure integer coordinates
- w, h = box_size # Width and height of the bounding box
- # Compute bounding box coordinates
- x1, y1 = max(cx - w // 2, 0), max(cy - h // 2, 0)
- x2, y2 = min(cx + w // 2, frame.shape[1]), min(cy + h // 2, frame.shape[0])
- # Crop the region of interest (ROI)
- roi = frame[y1:y2, x1:x2]
- # Convert to grayscale if the image is color--if invalid, return empty list
- if roi is not None: roi = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY)
- else: return []
- # Grid division
- rows, cols = grid_size
- step_y, step_x = (y2 - y1) // rows, (x2 - x1) // cols # Compute step size
- sampled_points = []
- # Loop over grid cells
- for i in range(rows):
- for j in range(cols):
- # Define cell boundaries
- y_start, y_end = y1 + i * step_y, y1 + (i + 1) * step_y
- x_start, x_end = x1 + j * step_x, x1 + (j + 1) * step_x
- # Extract the sub-region
- cell = roi[(i * step_y):(i + 1) * step_y, (j * step_x):(j + 1) * step_x]
- # Ensure cell has valid size before processing
- if cell.size == 0: continue
- # Find the darkest point in this cell
- min_val, _, min_loc, _ = cv2.minMaxLoc(cell)
- # Adjust coordinates to global frame
- dark_x = x_start + min_loc[0]
- dark_y = y_start + min_loc[1]
- sampled_points.append((dark_x, dark_y))
- return sampled_points
- ################################################################################
- def calc_motion_vecs(dark_points_t1, dark_points_t2, threshold=200):
- """
- Calculate motion vectors by matching dark points between two consecutive frames.
- Inputs:
- dark_points_t1 (list): List of (x, y) coordinates from frame t-1.
- dark_points_t2 (list): List of (x, y) coordinates from frame t.
- threshold (float): Maximum distance allowed to consider two points as a match. Default is 200.
- Output:
- ndarray: Array of motion vectors [(dx, dy), ...] representing the displacement
- of matched points from t-1 to t. Only vectors below the threshold are included.
- """
- # Initialize array of motion vectors
- motion_vectors = []
- for pt1 in dark_points_t1:
- # Calculate distances from pt1 to all points in t2
- distances = np.linalg.norm(np.array(dark_points_t2) - np.array(pt1), axis=1)
- # Find the closest point in t2
- min_distance = np.min(distances)
- if min_distance <= threshold:
- # Get the index of the closest point
- closest_idx = np.argmin(distances)
- pt2 = dark_points_t2[closest_idx]
- # Calculate the motion (difference in coordinates)
- dx = pt2[0] - pt1[0]
- dy = pt2[1] - pt1[1]
- motion_vectors.append((dx, dy))
- return np.array(motion_vectors)
- ################################################################################
- def find_troughs_and_peaks(signal, distance=10, height=[0.1,5]):
- """
- Detect peaks (local maxima) that are followed by troughs (local minima) within a specified distance.
- Inputs:
- signal (ndarray): 1D array of signal values.
- distance (int): Maximum number of frames allowed between a peak and its corresponding trough.
- Default is 10.
- height (list): Two-element list specifying the minimum and maximum height of peaks/troughs.
- Default is [0.1, 5].
- Outputs:
- tuple:
- - troughs (ndarray): Indices of troughs that follow valid peaks within the specified distance.
- - peaks (ndarray): Indices of peaks that are followed by a trough within the specified distance.
- """
- # Find troughs (invert signal to detect minima)
- troughs, _ = find_peaks(-signal, distance=distance/2, height=(height[0], height[1]))
- # Find peaks
- peaks, _ = find_peaks(signal, distance=distance/2, height=(height[0], height[1]))
- # Keep only peaks that come after a trough
- valid_peaks = []
- valid_troughs = []
- for peak in peaks:
- # Find the first trough after the peak
- following_troughs = troughs[troughs > peak]
- # Remove any already matched troughs to ensure exclusivity
- following_troughs = np.array([t for t in following_troughs if t not in valid_troughs])
- if len(following_troughs) > 0 and np.abs(following_troughs[0] - peak) < distance:
- valid_troughs.append(following_troughs[0])
- valid_peaks.append(peak)
- return np.array(valid_troughs), np.array(valid_peaks)
- ################################################################################
- def make_bin_edges(bin_width, dt, total_time):
- """
- Generate timestamp edges to segment a time series into bins of fixed width.
- Inputs:
- bin_width (float): Desired width of each bin in seconds.
- dt (float): Time step between consecutive samples in seconds.
- total_time (float): Total duration of the signal in seconds.
- Outputs:
- ndarray: Array of timestamps representing the edges of each bin, starting at 0 and
- covering the total_time in increments of bin_width.
- """
- # Compute the timesteps (same code as previous)
- ts = np.arange(0, total_time, dt)
- # Num timesteps
- n_timesteps = np.prod(ts.shape)
- # Warn if binsize doesn't divide the timestep evenly
- if (bin_width % dt) >= dt :
- print("Warning: bin_width doesn't evenly divide the timestep")
- print(bin_width % dt)
- # How many timesteps per bin?
- steps_per_bin = np.around(bin_width / dt)
- # Compute the bin edges using np.arange
- # remember that arange is an open interval at the end
- # so add one to the number of timesteps
- bin_edges = np.arange(0, n_timesteps+1, steps_per_bin) * dt
- return bin_edges
- ################################################################################
- def motion_comp(prev_frame, curr_frame, num_points=500, points_to_use=500):
- """
- Estimate the motion transformation matrix between two sequential image frames.
- Contains code adapted from [https://github.com/itberrios/CV_projects/tree/main]
- Inputs:
- prev_frame (ndarray): The first image frame (RGB).
- curr_frame (ndarray): The second sequential image frame (RGB).
- num_points (int): Number of feature points to detect in the first frame.
- points_to_use (int): Number of matched points to use for motion estimation.
- Outputs:
- A (ndarray or None): Estimated affine transformation matrix (2x3) or None if estimation fails.
- prev_points (ndarray): Feature points detected in the previous frame.
- curr_points (ndarray): Corresponding matched points in the current frame.
- """
- # Convert to grayscale
- prev_gray = cv2.cvtColor(prev_frame, cv2.COLOR_RGB2GRAY)
- curr_gray = cv2.cvtColor(curr_frame, cv2.COLOR_RGB2GRAY)
- # Get features for first frame
- features = cv2.goodFeaturesToTrack(prev_gray, num_points, qualityLevel=0.01, minDistance=10)
- # If no feature points exist, exit
- if features is None: return None, [], []
- # Get matching features in next frame with Sparse Optical Flow Estimation
- matched_features, status, _ = cv2.calcOpticalFlowPyrLK(prev_gray, curr_gray, features, None)
- # Reformat previous and current feature points
- prev_points = features[status==1]
- curr_points = matched_features[status==1]
- # Subsample number of points so we don't overfit
- if points_to_use > prev_points.shape[0]: points_to_use = prev_points.shape[0]
- index = np.random.choice(prev_points.shape[0], size=points_to_use, replace=False)
- prev_points_used = prev_points[index]
- curr_points_used = curr_points[index]
- # If no feature points exist, exit
- if len(prev_points_used) == 0 or len(curr_points_used) == 0: return None, [], []
- # Find transformation matrix from frame 1 to frame 2
- A, _ = cv2.estimateAffine2D(prev_points_used, curr_points_used, method=cv2.RANSAC)
- return A, prev_points, curr_points
- ################################################################################
- def get_grinder_motion(prev_frame, curr_frame, prev_com, curr_com):
- """
- Detects matched grinder keypoints and computes motion vectors between two frames.
- Contains code adapted from [https://github.com/itberrios/CV_projects/tree/main]
- Inputs:
- prev_frame (ndarray): Previous video frame.
- curr_frame (ndarray): Current video frame.
- prev_com (tuple): Centroid (x, y) of the grinder in the previous frame.
- curr_com (tuple): Centroid (x, y) of the grinder in the current frame.
- Outputs:
- prev_grinder_pts (ndarray): Transformed grinder points from previous frame.
- curr_grinder_pts (ndarray): Grinder points sampled in current frame.
- motion (ndarray): Motion vectors between matched points.
- magnitude (ndarray): Magnitude of motion vectors.
- angle (ndarray): Angles (radians) of motion vectors.
- """
- ### Get affine transformation for frame alignment
- # Get frame info
- h, w, _ = prev_frame.shape
- # Get affine transformation matrix for motion compensation between frames
- A, prev_pts, curr_pts = motion_comp(prev_frame, curr_frame, num_points=10000, points_to_use=5000)
- ### Transform previous frame's points using affine transformation
- # First, check that A was obtained correctly
- if A is None: return [],[],[],[],[]
- # Get transformed grinder points from the previous frame using A
- A = np.vstack((A, np.zeros((3,)))) # get 3x3 matrix to transform points
- prev_grinder_pts = sample_dark_pts(prev_frame, prev_com)
- curr_grinder_pts = sample_dark_pts(curr_frame, curr_com)
- # Check if points aren't found
- if not prev_grinder_pts or not curr_grinder_pts: return [],[],[],[],[]
- # Compensate the previous frame's points for motion
- comp_pts = np.hstack((prev_grinder_pts, np.ones((len(prev_grinder_pts), 1)))) @ A.T
- comp_pts = comp_pts[:, :2]
- ### Obtain motion information about the grinder points
- motion = calc_motion_vecs(comp_pts, curr_grinder_pts)
- if len(motion) < 2: return [],[],[],[],[]
- magnitude = np.linalg.norm(motion, ord=2, axis=1)
- angle = np.arctan2(motion[:, 0], motion[:, 1])
- return comp_pts, curr_grinder_pts, motion, magnitude, angle
- ################################################################################
- def process_video(video_path=None, grinder_coms=None):
- """
- Processes a video to compute motion vectors between grinder positions frame-by-frame.
- Inputs:
- video_path (str): Path to the input video file.
- grinder_coms (ndarray): Array of grinder center-of-mass positions per frame, shape (n_frames, 2).
- Outputs:
- mags (list of ndarray): List of arrays containing motion vector magnitudes per frame.
- angs (list of ndarray): List of arrays containing motion vector angles (radians) per frame.
- """
- ### Check that a video path and grinder CoMs were both provided
- if video_path is None: raise ValueError("No video path was provided.")
- if grinder_coms is None: raise ValueError("No grinder CoMs were provided.")
- ### Load video
- video_name = os.path.splitext(os.path.basename(video_path))[0]
- vid = cv2.VideoCapture(video_path)
- width = int(vid.get(cv2.CAP_PROP_FRAME_WIDTH))
- height = int(vid.get(cv2.CAP_PROP_FRAME_HEIGHT))
- fps = vid.get(cv2.CAP_PROP_FPS)
- tot_frames = int(vid.get(cv2.CAP_PROP_FRAME_COUNT))
- ### Initialize lists for motion information and labeled frames
- mags, angs = [], []
- saved_pts, frames = [], []
- # Initialize frames and coms
- success, prev_frame = vid.read()
- prev_com = grinder_coms[0]
- i = 1
- while success and i < grinder_coms.shape[0]:
- success, curr_frame = vid.read()
- curr_com = grinder_coms[i] if success else None
- curr_pts = []
- if success:
- # If a grinder was detected, calculate the motion
- if any(prev_com) and any(curr_com):
- # Get the distances between the grinder points in the previous and current frames
- prev_pts, curr_pts, motion, mag, ang = get_grinder_motion(prev_frame, curr_frame, prev_com, curr_com)
- # Save the motion information from the current frame
- mags.append(mag)
- angs.append(ang)
- # Draw detected motion vectors
- if curr_pts and prev_pts.shape == motion.shape and save:
- plot_frame = plot_vecs(prev_frame.copy(), np.hstack([prev_pts,prev_pts+motion]))
- # If a grinder was NOT detected, fill with blanks and continue
- else:
- mags.append([])
- angs.append([])
- # Save previous frame and grinder CoM for next iteration
- prev_frame = curr_frame.copy()
- prev_com = curr_com
- # Update iterator
- i += 1
- # Save the cluster locations from each frame
- saved_pts.append(curr_pts)
- return mags, angs
- ################################################################################
- def get_strain_cond(video_name):
- """
- Parses a video filename to extract the strain and condition identifiers.
- Inputs:
- video_name (str): Full name of the video file (e.g., 'n2_ff_00001.wmv' or 'n2_ff_00001_cropped.wmv').
- Outputs:
- strain (str): Extracted strain identifier from the filename (e.g., 'n2').
- condition (str): Extracted condition identifier from the filename (e.g., 'ff').
- Raises:
- ValueError: If the filename does not match the expected format 'STRAIN_CONDITION_#####.ext'.
- """
- base = os.path.splitext(video_name)[0] # Remove extension, e.g., '.wmv'
- # Match STRAIN_CONDITION_##### optionally followed by '_cropped'
- match = re.match(r'^([^_]+)_([^_]+)_\d+(?:_cropped)?$', base)
- if not match:
- raise ValueError("Input string must be in format 'STRAIN_CONDITION_#####.ext'")
- strain, condition = match.group(1), match.group(2)
- return strain, condition
- ################################################################################
- def get_filtered_motion(strain_name=None, cond_name=None, mags=None, angs=None, fps=20):
- """
- Calculates the signed and filtered motion magnitude of grinder points across frames.
- Inputs:
- strain_name (str): Name of the genetic strain.
- cond_name (str): Name of the experimental condition.
- mags (list of ndarray): List of motion magnitude arrays per frame.
- angs (list of ndarray): List of motion angle arrays per frame (radians).
- fps (float): Sampling frequency of the video frames (default is 20).
- Outputs:
- signed_mags_filt (ndarray): Low-pass filtered signed average motion magnitude signal.
- fc (float): Cutoff frequency used for filtering based on the strain and condition.
- """
- ### Check that a strain name, condition name, and motion information were all provided
- if strain_name is None: raise ValueError("No strain name was provided.")
- if cond_name is None: raise ValueError("No condition name was provided.")
- if mags is None: raise ValueError("No magnitudes were provided.")
- if angs is None: raise ValueError("No angles were provided.")
- ### Determine if any of the grinder points have moved significantly
- # Store magnitudes
- max_mag = 5 # ignore motion vectors exceeding this threshold
- n_frames = len(mags)
- avg_mags = np.zeros(n_frames) # unsigned magnitude of motion
- avg_angs = np.zeros(n_frames) # average angle of motion
- signs = np.zeros(n_frames) # negative/positive motion
- avg_signed_mags = np.zeros(n_frames) # signed magnitude of motion
- # Main loop for calculating signed motion magnitude
- for i in range(n_frames):
- # Save the average angle values (used to determine sign)
- if len(angs[i]) > 0: avg_angs[i] = np.mean(angs[i])
- else: avg_angs[i] = 0
- # Determine the sign based on the angle value
- if avg_angs[i] > 0: signs[i] = 1
- elif avg_angs[i] < 0: signs[i] = -1
- # Save the largest magnitude + average magnitude from each frame
- if len(mags[i]) > 0 and len([x for x in mags[i] if x < max_mag]) > 0:
- avg_mags[i] = np.mean([x for x in mags[i] if x < max_mag])
- else: avg_mags[i] = 0
- # Save the signed magnitude
- avg_signed_mags[i] = avg_mags[i]*signs[i]
- ### Build low-pass filter and apply to signal
- ### NOTE: if using a new strain, you will need to add it as follows:
- ### elif strain_name == 'new_strain' and (cond_name == 'ff' or cond_name == 'sf'):
- ### fc = 2*expected_max_rate_for_new_strain
- # Use a cutoff frequency (fc) equal to 2*max_expected_pumping_rate
- if strain_name == 'n2' and (cond_name == 'ff' or cond_name == 'sf'):
- fc = 10 # N2 on food has a higher expected pumping rate
- elif strain_name == 'eat2' and (cond_name == 'ff' or cond_name == 'sf'):
- fc = 4 # eat-2 on food have lower expected pumping rate
- elif cond_name == 'fs' or cond_name == 'ss':
- fc = 2 # worms off food have signficantly lower expected pumping rate
- Wn = fc / (fps/2) # fps is the sampling frequency
- b, a = iirfilter(N=2, Wn=Wn, btype='low', ftype='butter') # low-pass Butterworth filter
- signed_mags_filt = filtfilt(b, a, avg_signed_mags) # filtered signed magnitude
- return signed_mags_filt, fc
- ################################################################################
- def get_pumping_rate(video_path=None, grinder_motion=None, fps=20, fc=None,
- min_height=0.5, max_height=5, sigma=1, save=False, save_path=None):
- """
- Estimates the continuous and discrete pumping rate of a grinder based on motion signals.
- Inputs:
- video_path (str): Path to the input video file.
- grinder_motion (ndarray): Array of grinder motion magnitudes per frame.
- fps (float): Frame rate of the video.
- fc (float): Cutoff frequency used for filtering based on the strain.
- min_height (float): Minimum height threshold for peak detection.
- max_height (float): Maximum height threshold for peak detection.
- sigma (float): Standard deviation for Gaussian smoothing of discrete pumping rate.
- save (bool): Whether to save the pumping rate and peak times to disk.
- save_path (str): Directory path to save output files (required if save=True).
- Outputs:
- pr_cont (ndarray): Continuous pumping rate signal (Gaussian smoothed).
- pr_disc (ndarray): Discrete pump counts per time bin.
- peaks (ndarray): Frame indices where valid pump peaks occur.
- troughs (ndarray): Frame indices where valid pump troughs occur.
- """
- ### Check that a video path, grinder motion, and grinder CoMs were both provided
- if video_path is None: raise ValueError("No video path was provided.")
- if grinder_motion is None: raise ValueError("No grinder motion signal was provided.")
- ### Detect peaks only if they are followed by a corresponding trough
- dist = int(fps/(fc/2)) # minimum number of frames it takes to complete a pump
- height_range = [min_height, max_height]
- troughs, peaks = find_troughs_and_peaks(grinder_motion, distance=dist, height=height_range)
- ### Bin the detected pumps (i.e., peaks followed by troughs)
- bins = make_bin_edges(fps,1/fps,grinder_motion.shape[0])
- pr_disc, bin_edges = np.histogram(peaks, bins=bins)
- ### Filter the binned counts to obtain continuous estimate of pumping rate
- pr_cont = gaussian_filter1d(pr_disc.astype('float'), sigma, mode='nearest')
- ### OPTIONAL: save motion and pumping rate estimates
- if save:
- if save_path is None: raise ValueError("No save path was provided.")
- video_name = os.path.splitext(os.path.basename(video_path))[0]
- save_rate_path = save_path + video_name.replace('_cropped', '') + '_PumpKin_rate.csv'
- np.savetxt(save_rate_path, pr_cont)
- print('Continuous pumping rate saved to ' + save_rate_path)
- save_times_path = save_path + video_name.replace('_cropped', '') + '_PumpKin_times.csv'
- np.savetxt(save_times_path, peaks/fps) # saving pump times in seconds, not frames
- print('Pumping times saved to ' + save_times_path)
- return pr_cont, pr_disc, peaks, troughs
pumpkin.py at commit 35b4184, under MIT · at the source
Overview
- Interdisciplinary Program in Bioengineering, Georgia Institute of Technology, Atlanta, GeorgiaUnited States of America
- School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GeorgiaUnited States of America
- School of Chemical & Biomolecular Engineering, Georgia Institute of Technology, Atlanta, GeorgiaUnited States of America
- Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GeorgiaUnited States of America
Abstract
One of the many goals of neuroscience is to understand how the brain encodes and transforms sensory information into behavior. These animal behaviors can be studied at the level of multi-limb poses or through the focused analysis of individual body parts. Techniques for tracking animal pose, such as DeepLabCut and SLEAP, enable detailed studies of large-scale multi-limb behaviors but show reduced accuracy when used for single-keypoint tracking, where insufficient spatial context leads to increased drift and instability in tracking (Arent I, Schmidt FP, Botsch M et al. Marker-less motion capture of insect locomotion with deep neural networks pre-trained on synthetic videos. Frontiers in Behavioral Neuroscience. Vol. 15. 2021. Tang G, Han Y, Sun X, et al. Anti-drift pose tracker (ADPT), a transformer-based network for robust animal pose estimation cross-species. eLife. Vol. 13. 2025). More general techniques, such as Faster Region-based Convolutional Neural Network (Faster R-CNN) and You Only Look Once (YOLO), have also been used to track location-based behaviors such as center-of-mass position and velocity. However, behaviors localized to a single body structure, such as the pharyngeal pumping (i.e., feeding) in the microscopic roundworm Caenorhabditis elegans (C. elegans), are particularly sensitive to noise from moving non-target body parts. This limitation cannot be resolved by simply adding more training data, as doing so often leads to overfitting rather than improved robustness, and instead requires additional processing beyond existing object tracking packages. To address these challenges, we present a fast, automated method that reliably measures pumping in freely moving C. elegans by combining a state-of-the-art object detector (Faster R-CNN) with a tunable noise filter in a technique we call PumpKin. To validate its performance, we demonstrate both its speed (average of 0.4 seconds/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.
lu-lab/PumpKin
35b41846face65d5dca3c04ca4ba68f520e9f219, 11 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
16 files
- EZ-FRCNN.ipynb, Jupyter, 157 lines
- annotate.ipynb, Jupyter, 27 lines
- cropROI.ipynb, Jupyter, 79 lines, 1 match
- library/
__init__.py , Python, 10 lines - library/
annotation.py , Python, 257 lines - library/
characterize.py , Python, 165 lines - library/
image_augs.py , Python, 55 lines - library/
inferencing.py , Python, 624 lines - library/
plotting.py , Python, 54 lines - library/
pumpkin.py , Python, 501 lines, 5 matches - library/
training.py , Python, 372 lines, 1 match - library/
utils.py , Python, 464 lines - pumpkin.ipynb, Jupyter, 97 lines
- saveCoMs.ipynb, Jupyter, 84 lines
- LICENSE, License, 21 lines
- README.md, Text, 82 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 14 scripts, each with its path and the digest of its content;
- 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
Data Availability
Data supporting the findings of this study are available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 12 MeSH terms, 3 funders, 37 references.
Cite
This paper
Shappell, E., Buggs, D., Walcott, J., & Lu, H. (2026). PumpKin: A machine-learning pipeline for automatically tracking localized kinematics in freely moving C. elegans. PLoS computational biology, 22(7), e1014489. https://
BibTeX
@article{shappell2026pum
author = {Shappell, Erin and Buggs, Debra and Walcott, Jennah and Lu, Hang},
title = {{PumpKin: A machine-learning pipeline for automatically tracking localized kinematics in freely moving C. elegans}},
journal = {PLoS computational biology},
year = {2026},
month = jul,
volume = {22},
number = {7},
pages = {e1014489},
publisher = {PLOS},
issn = {1553-734X},
doi = {10.1371/
url = {https://
pmid = {42467749},
pmcid = {PMC13399524}
}
RIS
TY - JOUR
AU - Shappell, Erin
AU - Buggs, Debra
AU - Walcott, Jennah
AU - Lu, Hang
TI - PumpKin: A machine-learning pipeline for automatically tracking localized kinematics in freely moving C. elegans
T2 - PLoS computational biology
J2 - PLoS Comput Biol
PY - 2026
DA - 2026/
VL - 22
IS - 7
SP - e1014489
SN - 1553-734X
PB - PLOS
DO - 10.1371/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1371/
"type": "article-journal",
"title": "PumpKin: A machine-learning pipeline for automatically tracking localized kinematics in freely moving C. elegans",
"container-title": "PLoS computational biology",
"author": [
{
"family": "Shappell",
"given": "Erin"
},
{
"family": "Buggs",
"given": "Debra"
},
{
"family": "Walcott",
"given": "Jennah"
},
{
"family": "Lu",
"given": "Hang"
}
],
"container-title-short":
"volume": "22",
"issue": "7",
"page": "e1014489",
"DOI": "10.1371/
"PMID": "42467749",
"PMCID": "PMC13399524",
"ISSN": "1553-734X",
"publisher": "PLOS",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
17
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-72057-9 [code]
- Sex-specific behavioral feedback modulates sensorimotor processing and drives flexible social behavior.Journal: Nature communicationsIn common: OpenCV, scikit-image, PyTorch, 4 other tools, 2 references
- [2] doi:10.1038/s41586-026-10348-3 [code]
- An enteric neuron ionotropic receptor regulates salt stress resistance.Journal: NatureIn common: OpenCV, scikit-image, PyTorch, 4 other tools, C. elegans, 1 reference
- [3] doi:10.1038/s41467-026-72710-3 [code]
- A modular multi-color fluorescence microscope for simultaneous tracking of cellular activity and behavior.Journal: Nature communicationsIn common: OpenCV, scikit-image, pandas, 3 other tools, C. elegans, 2 references
- [4] doi:10.1038/s41593-026-02262-8 [code]
- Cheese3D enables sensitive detection and analysis of whole-face movement in mice.Journal: Nature neuroscienceIn common: OpenCV, scikit-image, pandas, 3 other tools, 3 references
- [5] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: OpenCV, scikit-image, PyTorch, 4 other tools, 2 references
- [6] doi:10.1038/s41593-026-02232-0 [code]
- Entorhinal cortex represents task-relevant remote locations independently of CA1.Journal: Nature neuroscienceIn common: OpenCV, scikit-image, PyTorch, 4 other tools, 2 references
- [7] doi:10.1038/s41467-026-72709-w [code]
- An epifluorescence microscope design for naturalistic behavior and cellular activity in freely moving Caenorhabditis elegans.Journal: Nature communicationsIn common: OpenCV, scikit-image, PyTorch, 4 other tools, C. elegans, 1 reference
- [8] doi:10.1371/journal.pcbi.1013441 [code]
- Large vision model framework for automated C. elegans analysis: From static morphometry to dynamic neural activity.Journal: PLoS computational biologyIn common: OpenCV, scikit-image, PyTorch, 4 other tools, C. elegans, methods / tools
- [9] doi:10.1038/s41467-026-73476-4 [code]
- Developmental molecular signatures define de novo cortico-brainstem circuit for skilled forelimb movement.Journal: Nature communicationsIn common: OpenCV, scikit-image, PyTorch, 3 other tools, 2 references
- [10] doi:10.1126/sciadv.aed3650 [code]
- Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &
lt;i& gt;m/ z& lt;/ i& gt; mapping and exploration. Journal: Science advancesIn common: OpenCV, scikit-image, PyTorch, 4 other tools, other, methods / tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 14 scripts, and 7 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:1785582721c29f72…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
