OSCR

Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches
  1. [1] § Methods › Model training and evaluation ↔ tangle_tracer/nft_datasets.py, lines 33–82 · score 0.87 · Gaussian blurring, vertical flips, affine, brightness, horizontal, hue
  2. [2] § Methods › Point annotation to segmentation masks ↔ tangle_tracer/annot_conversion/ROIAnnotation.py, lines 306–370 · score 0.81 · Otsu thresholded, binary mask, scikit-image, rgb2hed, isolating, DAB
  3. [3] § Methods › Point annotation to segmentation masks ↔ tangle_tracer/annot_conversion/WSIAnnotation.py, lines 436–500 · score 0.80 · Otsu thresholded, binary mask, scikit-image, rgb2hed, isolating, DAB
  4. [4] § Methods › Model training and evaluation ↔ tangle_tracer/modules.py, lines 13–38 · score 0.73 · Tversky loss, pre trained, encoders, UNet, beta, ResNet50
  5. [5] § Methods › Comparing model predictions with AT8 burden and semi-quantitative scores ↔ tangle_tracer/annot_conversion/ROIAnnotation.py, lines 306–370 · score 0.73 · HED channels, Otsu thresholding, scikit-image, RGB, DAB, binarized
  6. [6] § Methods › Comparing model predictions with AT8 burden and semi-quantitative scores ↔ tangle_tracer/annot_conversion/WSIAnnotation.py, lines 436–500 · score 0.73 · HED channels, Otsu thresholding, scikit-image, RGB, DAB, binarized
  7. [7] § Methods › Model training and evaluation ↔ tangle_tracer/train.py, lines 78–121 · score 0.72 · wandb sweep, weight decay, momentum, Tversky, alpha, epoch
  8. [8] § Methods › Point annotation to segmentation masks ↔ tangle_tracer/annot_conversion/ROIAnnotation.py, lines 241–303 · score 0.70 · largest blob, center bias, scikit-image, background, skimage, empty
  9. [9] § Methods › Point annotation to segmentation masks ↔ tangle_tracer/annot_conversion/WSIAnnotation.py, lines 378–433 · score 0.70 · largest blob, center bias, scikit-image, background, skimage, empty
  10. [10] § Methods › Object detection: training an object detection model from bootstrapped bounding boxes ↔ tangle_tracer/object_detection/train_yolo.py, lines 6–45 · score 0.68 · model.tune, hyperparameter search, Ultralytics, optimized, YOLOv8, train
  11. [11] § Methods › Datatype conversion ↔ tangle_tracer/nft_datasets.py, lines 86–151 · score 0.67 · 1–343, temporal at8, s1, scenes, space, shuffling
  12. [12] § Methods › Datatype conversion ↔ tangle_tracer/annot_conversion/czi_to_zarr.py, lines 68–123 · score 0.67 · multiple scenes, zstd, clevel, Blosc, compression, chunk
  13. [13] § Methods › Model training and evaluation ↔ tangle_tracer/train.py, lines 14–74 · score 0.59 · ImageNet normalized, monitored, Tversky, ResNet50, hyperparameter, alpha
  14. [14] § Methods › Model training and evaluation ↔ tangle_tracer/nft_datasets.py, lines 573–656 · score 0.59 · PyTorch, static, stride, inference, empty, rows
  15. [15] § Methods › Metric choices ↔ tangle_tracer/modules.py, lines 71–145 · score 0.53 · AUPRC, AUROC, intersection, recall, aggregate, union
  16. [16] § Methods › Rotated ROI correction ↔ tangle_tracer/annot_conversion/WSIAnnotation.py, lines 266–304 · score 0.53 · corrected ROI, inscribing, rotated, crop, disk, slicing

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 654 lines · 27 KB · MIT · 4 matches

  1. import numpy as np
  2. import zarr
  3. import cv2
  4. from tqdm import tqdm
  5. import os, shutil, sys
  6. from pathlib import Path
  7. from matplotlib import pyplot as plt
  8. from matplotlib.colors import ListedColormap
  9. import matplotlib as mpl
  10. import zarr
  11. from numcodecs import Blosc
  12. import pandas as pd
  13. import skimage
  14. import pickle
  15. from tangle_tracer.annot_conversion.cz_utils import CzParser
  16. from tangle_tracer.annot_conversion.czi_to_zarr import get_section
  17. # from cz_utils import CzParser
  18. # from zarr_utils import get_section #these imports cause pytest to glitch?
  19. #create a WSIAnnotation class
  20. class WSIAnnotation():
  21. """
  22. Goal: Return an info_dict that can be used to represent bounding boxes within the image (WSI or rectangles?)
  23. """
  24. def __init__(self, img_path, annot_path, scale = 1.0, save = True, overwrite = True,
  25. save_path = '/scratch/sghandian/projects/nft/model_input2/',
  26. reannotations = '/srv/home/lianeb/reannotations.csv'):
  27. """
  28. Initialize WSIAnnotation.
  29. """
  30. try:
  31. self.reannotations = pd.read_csv(reannotations)
  32. except FileNotFoundError:
  33. print("Default reannotation file not found. Either pass in None for reannotations parameter or change path.")
  34. self.reannotations = []
  35. self.scale = scale
  36. img_path = Path(img_path)
  37. cz_annot_path = self.__create_annot_path(annot_path)
  38. self.annot_data = CzParser(cz_annot_path).get_regions()
  39. self.wsi_name = img_path.as_posix().split('/')[-1].split('.')[0]
  40. self.wsi_img = zarr.open(img_path, mode = 'r', dtype = 'i4')
  41. #define bbox size
  42. self.bbox_size = int(400 * self.scale)
  43. #define input tile size for model, pad ROI by this much + jitter amount
  44. self.input_tile_size = int(1024 * self.scale)
  45. #how much to vary random sampling coords of positive class
  46. self.jitter_amount = self.input_tile_size // 2
  47. #calculate total padding size
  48. self.padding_size = self.input_tile_size + self.jitter_amount
  49. #need save_path for zarr dump of corrected ROIs
  50. self.save_path = save_path
  51. self.chunk_size = 1000 #used for zarr chunks
  52. self.overwrite = overwrite
  53. try:
  54. self.img_regions = self._get_img_regions(self.annot_data['regions'])
  55. self.bboxes = self._get_img_bboxes(self.img_regions)
  56. self.masks = self._draw_all_masks(biggest = False)
  57. except FileNotFoundError:
  58. print(f'No annotation data found for {self.wsi_name}. \
  59. Returning WSIAnnot with reference to zarr image only.')
  60. if save:
  61. self.save_annot()
  62. def __create_annot_path(self, annot_path):
  63. basename = os.path.basename(annot_path)
  64. if '_s' in basename:
  65. name = "_".join(basename.split('_')[:-1])
  66. annot_path = Path("/".join(annot_path.as_posix().split('/')[:-1]),
  67. name).with_suffix('.cz')
  68. return annot_path
  69. def __get_nft_bboxes(self, rectangle):
  70. """
  71. Creates bbox DIMENSIONS around each NFT in a given ROI. Filters out NFTs
  72. that are not within this ROI.
  73. Input:
  74. rectangle (str) - Index to self.annot_data['regions'] which contains
  75. metadata specific to this ROI within the larger WSI.
  76. Returns:
  77. bboxes (dict) - Maps an NFT (str) to its bbox dimensions (dict).
  78. """
  79. #define nft and rectangle metadata
  80. rect_annots = self.annot_data['regions'][rectangle]
  81. offset = (rect_annots['Left'] * self.scale,
  82. rect_annots['Top'] * self.scale)
  83. angle = -(rect_annots['Rotation']) #use negative version
  84. rect_width, rect_height = rect_annots['Width'] * self.scale, rect_annots['Height'] * self.scale
  85. rect_center = np.int32([rect_width//2, rect_height//2])
  86. #use local coords if provided
  87. if 'nfts' in rect_annots:
  88. nft_annots = rect_annots['nfts']
  89. else:
  90. nft_annots = self.annot_data['nfts']
  91. #define new point annotations
  92. num_nfts = len(nft_annots) #number of original nfts
  93. #define rotation matrix M
  94. M = self.get_rotation_matrix(angle)
  95. def process_annots_into_bboxes(annot_dict, bbox_size, rect_width, rect_height, scale = 1.0):
  96. #bbox border in each direction from center
  97. bboxes = {}
  98. border = int(bbox_size // 2)
  99. for nft, coord in annot_dict.items():
  100. x,y = int(coord[0] * scale), int(coord[1] * scale)
  101. dims = {
  102. 'xmin': int(x - border),
  103. 'ymin': int(y - border),
  104. 'xmax': int(x + border),
  105. 'ymax': int(y + border),
  106. 'width': int(border * 2),
  107. 'height': int(border * 2), #can likely get rid of width and height
  108. }
  109. #check to make sure the bbox is in this rectangle
  110. if all(int(x) > - border for x in dims.values()): #filtered out by >0 condition
  111. if (dims['xmax'] <= rect_width + border) and (dims['ymax'] <= rect_height + border): #could have problems w/ border and edge NFTs
  112. bboxes[nft] = dims
  113. return bboxes
  114. def apply_rotation_to_bboxes(bboxes, bbox_size, rect_width, rect_height, offset, scale = 1.0):
  115. #check each global NFT annot for existence in this ROI (after rotation)
  116. border = int(bbox_size // 2)
  117. for nft, coord in bboxes.items():
  118. x, y = int(coord[0] * self.scale), int(coord[1] * self.scale)
  119. #convert to local coordinates
  120. xy = np.int32([x - offset[0], y - offset[1]])
  121. #shift coords about center of rectangle
  122. xy -= rect_center
  123. #use rotation matrix, order matters!
  124. if 'nfts' in rect_annots:
  125. x, y = np.int32(M @ xy) + np.flip(rect_center)
  126. #may need to add a check based on final orientation
  127. else:
  128. #rect_center to correct shift after rotation, not for local
  129. #shift coordinates back from origin for global coords!
  130. x, y = np.int32(M @ xy) + rect_center
  131. dims = {
  132. 'xmin': int(x - border),
  133. 'ymin': int(y - border),
  134. 'xmax': int(x + border),
  135. 'ymax': int(y + border),
  136. 'width': int(border * 2),
  137. 'height': int(border * 2), #can likely get rid of width and height
  138. }
  139. #check to make sure the bbox is in this rectangle
  140. if all(int(x) > - border for x in dims.values()): #filtered out by >0 condition
  141. if (dims['xmax'] <= rect_width + border) and (dims['ymax'] <= rect_height + border): #could have problems w/ border and edge NFTs
  142. bboxes[nft] = dims
  143. return bboxes
  144. #post-process these boxes by adding the pad_amount to all of them!
  145. def modify_all_dims(bboxes, pad_amount = self.padding_size):
  146. def modify_dims(bbox_dims, pad_amount = pad_amount):
  147. bbox_dims = bbox_dims.copy()
  148. for dim, val in bbox_dims.items():
  149. if dim in ['width', 'height']:
  150. continue
  151. bbox_dims[dim] = val + pad_amount
  152. return bbox_dims
  153. #copy bboxes and apply modify_dims to each bbox
  154. bboxes = bboxes.copy()
  155. for annot, bbox_dim in bboxes.items():
  156. bboxes[annot] = modify_dims(bbox_dim)
  157. return bboxes
  158. #convert raw NFT annotations from WSI-level into bboxes
  159. bboxes = process_annots_into_bboxes(nft_annots, self.bbox_size, rect_width, rect_height, self.scale)
  160. bboxes = apply_rotation_to_bboxes(bboxes, self.bbox_size, rect_width, rect_height, offset, self.scale)
  161. bboxes = modify_all_dims(bboxes, pad_amount = self.padding_size)
  162. if len(self.reannotations) > 0:
  163. df = self.reannotations
  164. reannot_roi = df[(df['ROI'] == int(rectangle.split('_')[-1])) & (df['WSI'] == self.wsi_name)]
  165. nft_reannots = reannot_roi['ROI_coordinates'].apply(eval).to_dict()
  166. new_keys = ['annotation_' + str(k+num_nfts) for k in nft_reannots.keys()]#nft_reannots.keys()]
  167. new_vals = [list(v) for v in nft_reannots.values()]
  168. nft_reannots = dict(zip(new_keys, new_vals))
  169. #no need to apply rotation to these as they were annotated at the ROI-level (pre-rotated)
  170. new_bboxes = process_annots_into_bboxes(nft_reannots, self.bbox_size, rect_width, rect_height, self.scale)
  171. new_bboxes = modify_all_dims(new_bboxes, pad_amount = 0)
  172. bboxes = {**bboxes, **new_bboxes}
  173. return bboxes
  174. def _transform_rect(self, metadata):
  175. #define and scale dimensions
  176. left, top = int(metadata['Left'] * self.scale), int(metadata['Top'] * self.scale)
  177. width, height = int(metadata['Width'] * self.scale), int(metadata['Height'] * self.scale)
  178. border, angle = 0, -metadata['Rotation'] #must use negative rot
  179. #get a path to a roi jpeg if it exists
  180. roi_path = metadata.get('Path', None)
  181. #calculate rotation M for given angle
  182. rotation_M = self.get_rotation_matrix(angle)
  183. #define points w/ respect to origin
  184. pt_A = [0, 0]
  185. pt_B = [0 + width, 0]
  186. pt_C = [0 + width, 0 + height]
  187. pt_D = [0, 0 + height]
  188. pts = np.int32([pt_A, pt_B, pt_C, pt_D])
  189. #shift to center of rectangle
  190. center = np.int32([[width//2, height//2]] * 4)
  191. pts = pts - center
  192. #define shifting params (border + offset dims)
  193. shift = np.float32([[border + left, border + top]] * 4) + center
  194. #create rotation matrix and input points
  195. in_pts = np.float32(pts @ rotation_M) + shift
  196. #define strict output points (height/width of rectangle)
  197. out_pts = np.float32([[0, 0],
  198. [width - 1, 0],
  199. [width - 1, height - 1],
  200. [0, height - 1]])
  201. #before doing transform, feed in a slice that
  202. #is cropped around the minimal region
  203. minCoords = np.flip(np.int32(in_pts.min(axis = 0)))
  204. maxCoords = np.flip(np.int32(in_pts.max(axis = 0)))
  205. tileSize = maxCoords - minCoords
  206. if not roi_path and self.wsi_img:
  207. img_slice = get_section(self.wsi_img, minCoords,
  208. tile_size = (tileSize))
  209. else:
  210. #load directory from jpeg if there's no WSI but there are ROIs
  211. try:
  212. img_slice = mpl.image.imread(roi_path)
  213. except FileNotFoundError:
  214. print('Path to ROI not found: {roi_path}. Fix path to ROI in metadata if no WSI exists.')
  215. #convert in_pts to local coordinates of img_slice
  216. in_pts -= np.flip(minCoords)
  217. #compute the perspective transform M
  218. M = cv2.getPerspectiveTransform(np.float32(in_pts),
  219. np.float32(out_pts))
  220. #warp the slice with the rotation matrix
  221. out = skimage.transform.warp(img_slice, np.linalg.inv(M),
  222. output_shape=(height, width))
  223. return out
  224. def _get_img_regions(self, rectangles):
  225. """
  226. Takes in the ROI annotations and generates a nested dictionary mapping rectangles to the ROI image,
  227. corrected for rotation, as well as the constitutent bbox dimensions for that specific rectangle.
  228. Input:
  229. rectangles (dict): Dictionary mapping rectangle name to global coordinates and rotation.
  230. Returns:
  231. img_slices (dict): Dictionary mapping rectangle name to corrected ROI view in np.array format,
  232. and coordinates of bboxes.
  233. """
  234. img_slices = {}
  235. #can potentially choose the relevant region first,
  236. for rectangle, metadata in tqdm(rectangles.items()):
  237. #uses self.wsi_img and coords to slice minimal inscribed img,
  238. #then rotates and returns the correct view of this cropped img
  239. img_slice = self._transform_rect(metadata)
  240. #pad array by padding amount on each side,
  241. #MIGHT WANT THIS TO BE TILE SIZE !!! 1024px
  242. padded_img = np.pad(img_slice, ((self.padding_size, self.padding_size),
  243. (self.padding_size, self.padding_size), (0, 0)),
  244. mode = 'constant', constant_values = 1)
  245. #convert img_slice to zarr file on disk?
  246. os.makedirs(Path(self.save_path, 'images', self.wsi_name), exist_ok=True)
  247. wsi_path = Path(self.save_path,'images', self.wsi_name, rectangle)
  248. zarr_arr = self.__dump_to_zarr(padded_img, wsi_path)
  249. #uses rectangle coords + angle to generate bounding boxes
  250. bboxes = self.__get_nft_bboxes(rectangle)
  251. #save to dictionary
  252. img_slices[rectangle] = {'img': zarr_arr, 'bboxes': bboxes}
  253. return img_slices
  254. def __dump_to_zarr(self, img: np.array, path: Path):
  255. # compressor = Blosc(cname='lz4', clevel=1, shuffle=1)
  256. compressor = Blosc(cname='zstd',clevel=5, shuffle=Blosc.BITSHUFFLE)
  257. synchronizer = zarr.ProcessSynchronizer(path.with_suffix('.sync'))
  258. store = zarr.DirectoryStore(path)
  259. if self.overwrite:
  260. try:
  261. shutil.rmtree(path)
  262. except FileNotFoundError:
  263. print('Cannot overwrite ROI that does not exist. Creating new zarr array.')
  264. try:
  265. z = zarr.array(img,
  266. store=store,
  267. chunks = (self.chunk_size, self.chunk_size),
  268. write_empty_chunks=False,
  269. compressor = compressor,
  270. synchronizer = synchronizer)
  271. except zarr.errors.ContainsArrayError:
  272. print(f'Skipping {path} since it already exists.')
  273. z = zarr.open(path, mode = 'r')
  274. return z
  275. def _load_bbox(self, rectangle_img, bbox_object):
  276. """
  277. Takes in an annotation for a bounding box and retuns a slice of the rectangle image.
  278. Input:
  279. rectangle_img (np.array): Slice of WSI in np format
  280. bbox_object (dict): Dictionary {annotation_# : xmin, ymin, xmax, ymax, width, height}
  281. scale (float) : Scale at which the image is produced. Will scale bbox coordinates appropriately.
  282. Returns:
  283. box_slice (np.array): Slice of the image contained in the bounding box, padded to 400px if necessary.
  284. """
  285. left = int(bbox_object['xmin'])
  286. top = int(bbox_object['ymin'])
  287. width = int(bbox_object['width'])
  288. height = int(bbox_object['height'])
  289. #add edge case handling via checking coordinates, altering indexing and padding
  290. box_slice = rectangle_img[slice(np.max([0, top]), top + height),
  291. slice(np.max([0, left]), left + width), :].copy()
  292. return box_slice
  293. def _get_img_bboxes(self, img_regions):
  294. """
  295. Takes in the dictionary of rectangles mapped to its image, as well as its constituent bboxes dimensions,
  296. and produces all associated bbox images in a dictionary mapping annotation name to image.
  297. Input:
  298. img (zarr.array): WSI in Zarr format
  299. img_regions (dict): Dictionary for a single WSI mapping each annotated rectangle name to another dictionary,
  300. containing image region and its associated annotated NFTs in bounding box format.
  301. Example:
  302. img_regions['rectangle_0'] = {'img': img_slice, 'bboxes': bboxes}
  303. Returns:
  304. bbox_imgs (dict): Dictionary mapping an NFT bounding box annotation to an image slice.
  305. Example:
  306. bbox_imgs['annotation_0'] = bbox
  307. """
  308. bbox_imgs = {}
  309. for rectangle in img_regions:
  310. for bbox, dims in img_regions[rectangle]['bboxes'].items():
  311. bbox_imgs[bbox] = self._load_bbox(img_regions[rectangle]['img'], dims)
  312. return bbox_imgs
  313. def _create_blobs(self, mask, threshold=1500, max_distance = 125):
  314. """
  315. Takes in a binarized bbox image and generates a mask representing a single NFT within it.
  316. """
  317. labels = skimage.measure.label(mask, connectivity=2, background=0)
  318. out_mask = np.zeros(mask.shape, dtype='uint8')
  319. sizes = {}
  320. # loop over the unique components
  321. for label in np.unique(labels):
  322. # if this is the background label, ignore it
  323. if label == 0:
  324. continue
  325. # otherwise, construct the label mask and count the
  326. # number of pixels
  327. labelMask = np.zeros(mask.shape, dtype="uint8")
  328. labelMask[labels == label] = 255
  329. numPixels = cv2.countNonZero(labelMask)
  330. # if the number of pixels in the component is sufficiently
  331. # large, then add it to our dictionary mapping labels to sizes
  332. if numPixels > threshold*(self.scale**2):
  333. #enforce center bias
  334. blob_center = np.float32([np.average(indices) for indices in np.where(labelMask == 255)])
  335. img_center = np.float32([200*self.scale, 200*self.scale])
  336. if np.linalg.norm(blob_center - img_center) < max_distance*self.scale:
  337. sizes[label] = numPixels
  338. #get the largest blob
  339. try:
  340. maxLabel = max(sizes, key = sizes.get)
  341. except ValueError: #nothing found
  342. #---------------- consider adding the next smallest blob!!! -------------
  343. return out_mask, out_mask
  344. #pop the max label
  345. maxPixel = sizes.pop(maxLabel)
  346. #generate a mask with the largest label
  347. biggest_mask = np.zeros(mask.shape, dtype="uint8")
  348. biggest_mask[labels == maxLabel] = 255
  349. # biggest_mask_size = cv2.countNonZero(biggest_mask)
  350. out_mask = cv2.add(out_mask, biggest_mask)
  351. #get the biggest blobs within X% of max size
  352. percentage_threshold = 0.50
  353. #loop through all sufficiently large blobs and get sizes
  354. for label, pixelSize in sizes.items():
  355. #check if they're large enough to add to an output mask
  356. if (pixelSize / maxPixel) > percentage_threshold:
  357. #create an empty mask, then assign 1s to labeled region
  358. labelMask = np.zeros(mask.shape, dtype="uint8")
  359. labelMask[labels == label] = 255
  360. #add label mask to output mask
  361. out_mask = cv2.add(out_mask, labelMask)
  362. return out_mask, biggest_mask
  363. #try otsu thresholding
  364. def _draw_mask(self, image, show_img = False, overlay = False):
  365. """
  366. Takes in an image and converts it from RGB to HED channels. The image is then normalized and
  367. converted to uint8, then the DAB channel is isolated and is binarized and thresholded via Otsu.
  368. Morphological operations then preprocess the image, then it's passed to _create_blobs() to generate
  369. the segmented mask for the image. Can use show_img and overlay flags to visualize results.
  370. Input:
  371. image (np.array): RGB image representing a zoomed-in NFT image from an ROI.
  372. show_img (bool): Flag to display image in notebook
  373. overlay (bool): Flag to overlay the mask onto the image. RETURNS the overlaid mask!
  374. Retuns:
  375. mask, biggest_mask (tuple): Mask images that represent the most central mask and the biggest masked
  376. region within the bbox.
  377. """
  378. og_image = np.array(skimage.color.rgb2hed(image).copy())
  379. normed_image = None
  380. normed_image = cv2.normalize(og_image, normed_image, alpha=0, beta=255, norm_type=cv2.NORM_MINMAX, dtype=cv2.CV_8U)
  381. #select DAB color channel
  382. try:
  383. DAB = normed_image[:, :, 2]
  384. except TypeError:
  385. print(f"OG image shape {og_image.shape} and image: {og_image}")
  386. return (None, None)
  387. #create binary mask
  388. thresh = cv2.threshold(DAB, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[1]
  389. #use cross as kernel
  390. kernel = cv2.getStructuringElement(cv2.MORPH_CROSS, (3, 3))
  391. # Apply morphological closing, then opening operations
  392. closing = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel)
  393. # for _ in range(10):
  394. # closing = cv2.morphologyEx(closing, cv2.MORPH_CLOSE, kernel)
  395. # opening = cv2.morphologyEx(closing, cv2.MORPH_OPEN, kernel)
  396. #create blobs and their masks
  397. mask, biggest_mask = self._create_blobs(closing, threshold = 1000 * self.scale)
  398. if overlay:
  399. # Choose colormap
  400. cmap = plt.get_cmap('cool')
  401. # Get the colormap colors
  402. my_cmap = cmap(np.arange(cmap.N))
  403. # Set alpha to 0 for non values of 1
  404. my_cmap[:,-1][:255] = 0
  405. # Create new colormap
  406. my_cmap = ListedColormap(my_cmap)
  407. #apply colormap to mask
  408. mask = my_cmap(mask)
  409. #add an extra dimension to image
  410. image = cv2.cvtColor(np.float32(image), cv2.COLOR_RGB2RGBA)
  411. if show_img:
  412. # fig = plt.figure()
  413. plt.imshow(image)
  414. plt.imshow(mask, cmap = 'binary')
  415. return mask, biggest_mask
  416. def _draw_all_masks(self, biggest = False):
  417. """
  418. Iterates through self.bboxes to retrieve 400x400px images for segmentation. Generates mask(s)
  419. based on strategy defined in _draw_mask(). Returns dictionary mapping NFT annotation name to
  420. a single grayscale image.
  421. """
  422. masks = {}
  423. for (nft, bbox) in self.bboxes.items():
  424. mask, biggest_mask = self._draw_mask(bbox, overlay = True)
  425. if biggest:
  426. masks[nft] = biggest_mask
  427. else:
  428. masks[nft] = mask
  429. return masks
  430. #generate a rotation matrix using degrees
  431. def get_rotation_matrix(self, angle):
  432. """
  433. Generate a 2d rotation matrix with an input angle, defined in degrees.
  434. """
  435. angle = np.deg2rad(angle)
  436. R = np.array([[np.cos(angle), -np.sin(angle)],
  437. [np.sin(angle), np.cos(angle)]])
  438. return R
  439. def save_annot(self):
  440. """
  441. Dumps the WSIAnnot object at the default save location for the specific scale at which
  442. the original WSI was loaded in (self.scale).
  443. Can be loaded again with WSIAnnotation.load_annot(WSI_NAME).
  444. """
  445. save_path = Path(self.save_path,'wsiAnnots', f'scale_{self.scale}')
  446. #create a directory if needed
  447. os.makedirs(save_path, exist_ok=True)
  448. pickle_path = Path(save_path, f"{self.wsi_name}.pickle")
  449. #remove pickle file/directory if it does exist
  450. if os.path.isfile(pickle_path) or os.path.islink(pickle_path):
  451. os.remove(pickle_path)
  452. elif os.path.isdir(pickle_path):
  453. shutil.rmtree(pickle_path)
  454. #dump pickle file
  455. with open(Path(save_path, f"{self.wsi_name}.pickle"), "wb") as file_to_store:
  456. pickle.dump(self, file_to_store)
  457. #----------------------------Utility Functions----------------------------#
  458. def plot_bboxes(wsiAnnot, num_imgs = 12, overlay = True):
  459. bbox_imgs = wsiAnnot.bboxes
  460. masks = wsiAnnot.masks
  461. num_imgs = min([num_imgs, len(bbox_imgs.keys())])
  462. fig, ax = plt.subplots(num_imgs // 4, 4, figsize=(16, num_imgs), sharex = True, sharey = True)
  463. ax = ax.flatten()
  464. max_imgs = (num_imgs // 4) * 4
  465. for i, (nft, bbox) in enumerate(bbox_imgs.items()):
  466. if i > max_imgs - 1 or i >= len(bbox_imgs.keys()) - 1:
  467. break
  468. ax[i].imshow(bbox)
  469. if overlay:
  470. ax[i].imshow(masks[nft])
  471. ax[i].set_title(nft)
  472. fig.suptitle(f'NFT Bounding Boxes - {wsiAnnot.wsi_name}')
  473. fig.tight_layout(rect=[0, 0.02, 1, 0.98])
  474. def load_annot(wsiAnnot_name, load_path = '/scratch/sghandian/nft/model_input_v2/wsiAnnots/', scale = 1.0):
  475. out_path = Path(load_path, f'scale_{scale}', f"{wsiAnnot_name}.pickle")
  476. with open(out_path, "rb") as file_to_read:
  477. try:
  478. loaded_object = pickle.load(file_to_read)
  479. except ModuleNotFoundError: #handle unpickling error (pickled with abs import instead of current relative import)
  480. file_to_read.seek(0)
  481. curr_dir = os.path.realpath(os.path.dirname(__file__))
  482. sys.path.insert(0, curr_dir)
  483. loaded_object = pickle.load(file_to_read)
  484. return loaded_object
  485. #----------------------------Utility Functions----------------------------#
  486. def main(args):
  487. SCALE = args.scale
  488. ANNOT_PATH = Path(args.annot_path)
  489. ZARR_PATH = Path(args.img_path, f'scale_{SCALE}/')
  490. SAVE_PATH = Path(args.save_path)
  491. #choose which wsi to create, can use indices for ease but only need to use one of the following
  492. WSI_NAME = args.wsi
  493. idx = args.idx
  494. filenames = os.listdir(ZARR_PATH)
  495. #if wsi_name passed in or idx given, load that one specifically and save it to disk
  496. if WSI_NAME or idx is not None:
  497. if WSI_NAME:
  498. file = WSI_NAME + '.zarr'
  499. else:
  500. file = filenames[idx]
  501. print(f"WSI name: {file}")
  502. imgPath = Path(ZARR_PATH, file)
  503. annotPath = Path(ANNOT_PATH, Path(file).with_suffix('.cz'))
  504. wsiAnnot_single = WSIAnnotation(imgPath, annotPath, SCALE, save = True,
  505. save_path=SAVE_PATH)
  506. return wsiAnnot_single
  507. #create and save each WSIAnnotation sequentially, unlike parallel_annots.py
  508. for file in filenames:
  509. imgPath = Path(ZARR_PATH, file)
  510. annotPath = Path(ANNOT_PATH, Path(file).with_suffix('.cz'))
  511. if not annotPath.isfile():
  512. print("""Skipping {file} as {annotPath} annotation does not exist.
  513. Check whether annot_path arg is correct if this image does have an annotation.""")
  514. continue #skip files that do not have .cz associated
  515. wsiAnnot = WSIAnnotation(imgPath, annotPath, SCALE, save = True,
  516. save_path=SAVE_PATH)
  517. if __name__ == '__main__':
  518. from argparse import ArgumentParser
  519. #add an argparser
  520. parser = ArgumentParser()
  521. parser.add_argument("--img_path", type=str, required=True) #need to pass one of these args in!
  522. parser.add_argument("--annot_path", type=str, required=True) #need to pass one of these args in!
  523. parser.add_argument("--wsi", type=str) #need to pass one of these args in!
  524. parser.add_argument("--idx", type=int) #need to pass one of these args in!
  525. parser.add_argument("--scale", type=float, default = 1.0)
  526. parser.add_argument("--save_path", type=str, default='/scratch/sghandian/projects/nft/')
  527. args = parser.parse_args()
  528. #run on all images
  529. main(args)
  530. """
  531. Example Usage:
  532. python annot_conversion/point_to_mask.py \
  533. --img_path='/scratch/sghandian/nft/raw_data/wsis/zarr/' \
  534. --annot_path='/scratch/sghandian/nft/raw_data/annotations/updated_annots/' \
  535. --idx=0 --scale=0.1 --save_path='../data/'
  536. """

WSIAnnotation.py at commit 8778679, under MIT · at the source

Overview

Authors: Sina Ghandian1,2,3,4, Liane Albarghouthi1,2,3,4, Kiana Nava5, Shivam R. Rai Sharma6,7, Lise Minaud1,2,3,4, Laurel Beckett8, Naomi Saito8, Charles DeCarli9, Robert A. Rissman10, Andrew F. Teich11, Lee-Way Jin5, Brittany N. Dugger5, Michael J. Keiser1,2,3,4
  1. Institute for Neurodegenerative Diseases, University of California, San Francisco,San Francisco, CA 94158 USA
  2. Bakar Computational Health Sciences Institute, University of California,San Francisco, CA 94158 USA
  3. Department of Pharmaceutical Chemistry, University of California, San Francisco,San Francisco, CA 94158 USA
  4. Department of Bioengineering and Therapeutic Sciences, University of California, San Francisco,San Francisco, CA 94158 USA
  5. Department of Pathology and Laboratory Medicine, School of Medicine, University of California, Davis,Sacramento, CA 95817 USA
  6. Department of Computer Science, University of California, Davis,Davis, CA 95616 USA
  7. Robust and Ubiquitous Networking (RUbiNet) Lab, University of California, Davis,Davis, CA 95616 USA
  8. Division of Biostatistics, Department of Public Health Sciences, University of California Davis,Davis, CA USA
  9. Department of Neurology, School of Medicine, Alzheimer’s Disease Research Center, University of California Davis,Sacramento, CA USA
  10. Department of Neurosciences, University of California San Diego,La Jolla, San Diego, CA USA
  11. Department of Neurology, Taub Institute for Research On Alzheimer’s Disease and Aging Brain, Columbia University Medical Center,New York, NY USA
Journal: Scientific reports, volume 16, issue 1, article 27095
Dates: received 28 April 2025; accepted 6 July 2026; published online 28 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41598-026-61605-4 · PMID 42665582 · PMCID PMC13524952 · OpenAlex W4396992496
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: histology / microscopy (modality), human (organism), other condition (population), Alzheimer's / dementia (population), methods / tools (subfield)
Methods: Spectral & time-frequency, Statistics, Machine learning, Connectivity, fMRI & imaging
Keywords: Digital pathology, Neurofibrillary tangles, NFT, Tau, Deep learning, Semantic segmentation, Object detection, Open source, Alzheimer disease, Neuropathology, Alzheimer's disease, Machine learning, Software, Image processing, Immunohistochemistry, Neurodegeneration, Prion diseases, Data processing
MeSH: Alzheimer Disease*, Deep Learning*, Image Processing, Computer-Assisted*, Neurofibrillary Tangles*, Brain, Humans, tau Proteins (* major topic)
Topic: Medical Image Segmentation Techniques (Computer Vision and Pattern Recognition, Computer Science), according to OpenAlex
Funding: National Institute on Aging (P30AG072972, P50AG008702, R01AG062517); Chan Zuckerberg Initiative (CZI) (DAF2018-191905 (https://doi.org/10.37921/550142lkcjzw))
Citations: not cited yet (Europe PMC); 63 references in the paper
Research resources: RRID:AB_223647

Abstract

Accumulation of abnormal tau protein into neurofibrillary tangles (NFTs) is a pathologic hallmark of Alzheimer disease (AD). Accurate detection of NFTs in tissue samples can reveal relationships with clinical, demographic, and genetic features through deep phenotyping. However, expert manual analysis is time-consuming, subject to observer variability, and cannot handle the data amounts generated by modern imaging. We present a scalable, open-source, deep-learning approach to quantify NFT burden in digital whole slide images (WSIs) of post-mortem human brain tissue. To achieve this, we developed a method to generate detailed NFT boundaries directly from single-point-per-NFT annotations. We then trained a semantic segmentation model on 45 annotated 2400 μm by 1200 μm regions of interest (ROIs) selected from 15 unique temporal cortex WSIs of AD cases from three institutions (University of California (UC)-Davis, UC-San Diego, and Columbia University). Segmenting NFTs at the single-pixel level, the model achieved an area under the receiver operating characteristic of 0.832 and an F1 of 0.527 (196-fold over random) on a held-out test set of 664 NFTs from 20 ROIs (7 WSIs). We compared this to deep object detection, which achieved comparable but coarser-grained performance that was 60% faster. The segmentation and object detection models correlated well with expert semi-quantitative scores at the whole-slide level (Spearman’s rho ρ = 0.654 (p = 6.50e-5) and ρ = 0.513 (p = 3.18e-3), respectively). We openly release this multi-institution deep-learning pipeline to provide detailed NFT spatial distribution and morphology analysis capability at a scale otherwise infeasible by manual assessment.

Supplementary Information: The online version contains supplementary material available at https://doi.org/10.1038/s41598-026-61605-4.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

keiserlab/tangle-tracer

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 877867971521646aefef9124a9b2b44ff775a0d0, 11 December 2025
Languages: Python (29)
Size: 54 files, 29 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (REQUIREMENTS.txt, setup.cfg, setup.py), tests
Not found: CITATION.cff, continuous integration, documentation
Tools: pandas (12 files), NumPy (11 files), PyTorch (10 files), PyTorch Lightning (6 files), OpenCV (5 files), Matplotlib (4 files), scikit-image (3 files), Pillow (2 files), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
31 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 29 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The images (rotated and cropped regions of interest and their segmentation masks) used to train and evaluate all machine learning models and data created and used in this project, including results, are available at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD1165. The raw WSIs are available for download in Zarr format (must be uncompressed with imagecodecs’ Jpegxr codec; see repo below for example). The Python scripts, Jupyter notebooks, and additional code information are available at https://github.com/keiserlab/tangle-tracer. The README.md file contains detailed information on how to execute the pipeline and recreate results.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 13 authors, 18 keywords, 7 MeSH terms, 2 funders, 47 references, 1 RRID.

Cite

This paper

Ghandian, S., Albarghouthi, L., Nava, K., Sharma, S. R. R., Minaud, L., Beckett, L., Saito, N., DeCarli, C., Rissman, R. A., Teich, A. F., Jin, L.-W., Dugger, B. N., & Keiser, M. J. (2026). Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations. Scientific reports, 16(1), 27095. https://doi.org/10.1038/s41598-026-61605-4

BibTeX

@article{ghandian2026learning,
author = {Ghandian, Sina and Albarghouthi, Liane and Nava, Kiana and Sharma, Shivam R. Rai and Minaud, Lise and Beckett, Laurel and Saito, Naomi and DeCarli, Charles and Rissman, Robert A. and Teich, Andrew F. and Jin, Lee-Way and Dugger, Brittany N. and Keiser, Michael J.},
title = {{Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations}},
journal = {Scientific reports},
year = {2026},
month = aug,
volume = {16},
number = {1},
pages = {27095},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-61605-4},
url = {https://doi.org/10.1038/s41598-026-61605-4},
pmid = {42665582},
pmcid = {PMC13524952}
}

RIS

TY - JOUR
AU - Ghandian, Sina
AU - Albarghouthi, Liane
AU - Nava, Kiana
AU - Sharma, Shivam R. Rai
AU - Minaud, Lise
AU - Beckett, Laurel
AU - Saito, Naomi
AU - DeCarli, Charles
AU - Rissman, Robert A.
AU - Teich, Andrew F.
AU - Jin, Lee-Way
AU - Dugger, Brittany N.
AU - Keiser, Michael J.
TI - Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/08/28
VL - 16
IS - 1
SP - 27095
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-61605-4
UR - https://doi.org/10.1038/s41598-026-61605-4
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-61605-4",
"type": "article-journal",
"title": "Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations",
"container-title": "Scientific reports",
"author": [
{
"family": "Ghandian",
"given": "Sina"
},
{
"family": "Albarghouthi",
"given": "Liane"
},
{
"family": "Nava",
"given": "Kiana"
},
{
"family": "Sharma",
"given": "Shivam R. Rai"
},
{
"family": "Minaud",
"given": "Lise"
},
{
"family": "Beckett",
"given": "Laurel"
},
{
"family": "Saito",
"given": "Naomi"
},
{
"family": "DeCarli",
"given": "Charles"
},
{
"family": "Rissman",
"given": "Robert A."
},
{
"family": "Teich",
"given": "Andrew F."
},
{
"family": "Jin",
"given": "Lee-Way"
},
{
"family": "Dugger",
"given": "Brittany N."
},
{
"family": "Keiser",
"given": "Michael J."
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "27095",
"DOI": "10.1038/s41598-026-61605-4",
"PMID": "42665582",
"PMCID": "PMC13524952",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-61605-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
28
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/jnen/nlaf152 [code]
Clinical and pathologic correlations of machine learning quantification of Aβ deposits across 3 brain regions of decedents with Alzheimer disease.
Journal: Journal of neuropathology and experimental neurology
In common: OpenCV, scikit-image, Pillow, 5 other tools, Alzheimer's / dementia, 9 references
[2] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: PyTorch Lightning, OpenCV, scikit-image, 6 other tools
[3] doi:10.1371/journal.pcbi.1014263 [code]
MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery.
Journal: PLoS computational biology
In common: PyTorch Lightning, OpenCV, scikit-image, 6 other tools
[4] doi:10.1038/s41598-026-43798-w [code]
A Machine learning pipeline to investigate tissue ingrowth in cerebral aneurysms using preclinical animal models.
Journal: Scientific reports
In common: OpenCV, scikit-image, Pillow, 4 other tools, histology / microscopy, methods / tools, 1 reference
[5] doi:10.1371/journal.pcbi.1013441 [code]
Large vision model framework for automated C. elegans analysis: From static morphometry to dynamic neural activity.
Journal: PLoS computational biology
In common: OpenCV, scikit-image, Pillow, 4 other tools, methods / tools, 2 references
[6] doi:10.1002/alz.71649 [code]
Postmortem brain MRI reveals differential associations of subcortical and limbic volumes with cortical thinning and neurodegenerative pathologies.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: OpenCV, scikit-image, Pillow, 5 other tools, histology / microscopy, Alzheimer's / dementia, other condition
[7] doi:10.1016/j.phro.2026.101056 [code]
Toward uncertainty-aware manual delineation of brain tumours using eye-tracking and image-derived features.
Journal: Physics and imaging in radiation oncology
In common: OpenCV, scikit-image, Pillow, 5 other tools, other condition, 1 reference
[8] doi:10.1038/s41598-026-57519-w [code]
Automated segmentation of neurons and spinal cord structures in immunofluorescence images using SpineDL.
Journal: Scientific reports
In common: OpenCV, scikit-image, Pillow, 5 other tools, histology / microscopy, methods / tools, other condition
[9] doi:10.1364/boe.605322 [code]
Generalized plaque digitization framework for multi-dimensional mesoscopic images.
Journal: Biomedical optics express
In common: OpenCV, scikit-image, Pillow, 5 other tools, 1 reference
[10] doi:10.1371/journal.pcbi.1013499 [code]
VesiclePy: A machine learning vesicle analysis toolbox for volume electron microscopy.
Journal: PLoS computational biology
In common: OpenCV, scikit-image, Pillow, 5 other tools, histology / microscopy, methods / tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.