OSCR

Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § Results › CD58 ↔ conservation-IgSFSize.py, lines 203–273 · score 0.88 · egg laying mammals, placental mammals, Dermoptera, Euarchontoglires, Eutheria, Lagomorpha
  2. [2] § Materials and Methods › Clustering of IgSF ↔ otherScripts/brotherhoodMethod.py, lines 80–91 · score 0.64 · squareform, weighted, correlation, linkage, metric, cut

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 532 lines · 18 KB · no license · 1 match

  1. import os
  2. import sys
  3. import argparse
  4. import statistics
  5. import subprocess
  6. import pickle
  7. import numpy as np
  8. import pandas as pd
  9. import matplotlib.pyplot as plt
  10. from labellines import labelLines
  11. import warnings
  12. from datetime import datetime
  13. homeDir = os.getcwd()
  14. colors = plt.get_cmap('tab10').colors
  15. os.makedirs(f'{homeDir}/results', exist_ok=True)
  16. warnings.filterwarnings("ignore", message="Tried to label line")
  17. now = datetime.now()
  18. timestamp = now.strftime("%m-%d-%Y__%H-%M-%S")
  19. ##### Need to update DOI
  20. descr = '''
  21. This script provides two figures:
  22. Gene Conservation:
  23. Genes input by their GeneIDs, GeneIDs, GeneAccession, or Cluster arguments will be shown in shades of gray.
  24. Genes input by their Group will be colored by their Group.
  25. Change in IgSF Cluster Size:
  26. If an indiviudal gene is input, the entire cluster this gene belongs to will be considered.
  27. If multiple genes are input, all clusters including these genes will be considered.
  28. By default each all clusters are combined.
  29. However, each change in cluster size can be seperated with the seperateClusters argument.
  30. Please limit the inputs or else the figure becomes very busy
  31. To investigate the gene accessions considered for this analysis, you can check our work with the CheckAccessions argument.
  32. We recommend setting the seperateClusters argument to yes if you intend to use the CheckAccessions argument as the output is easier to interpret.
  33. You must provide at least one GeneID, GeneName, GeneAccession, Group, or Cluster.
  34. This script can be run across different taxonomic levels (Eukaryotes, Gnathostomata, Mammalia, and Primates)
  35. Multiple entries can be supplied for each above argument by separating them with commas.
  36. '''
  37. parser = argparse.ArgumentParser(formatter_class=argparse.RawTextHelpFormatter, description = descr)
  38. parser.add_argument("-gi", "--GeneIDs", help = "Please refer to humanIgSFClusters.ods for acceptable GeneIDs.", default = 'NA')
  39. parser.add_argument("-gn", "--GeneNames", help = "Please refer to humanIgSFClusters.ods for acceptable Gene Names.", default = 'NA')
  40. parser.add_argument("-ga", "--GeneAccession", help = "Please refer to humanIgSFClusters.ods for acceptable Gene Accessions", default = 'NA')
  41. parser.add_argument("-gr", "--Groups", help = "Options: Orange, Green, Red, Purple, Brown, Pink. Please refer to Fig xx", default = 'NA') ##### Need to update figure number
  42. parser.add_argument("-cl", "--Clusters", help = "Options: 0-132. Please refer to humanIgSFClusters.ods", default = 'NA')
  43. parser.add_argument("-tl", "--TaxonomicLevel", help = "Options: Eukaryotes, Gnathostomata, Mammalia, Primates. Default = Gnathostomata", default = 'Gnathostomata')
  44. parser.add_argument("-ca", "--CheckAccessions", help = "Create .tsv file showing queried human and other species clustered IgSF accessions with descriptions. Options: yes or no. Default = no", default = 'no')
  45. parser.add_argument("-sc", "--seperateClusters", help = "Seperates clusters in change in IgSF Cluster Size figure. Options: yes or no. Default = no", default = 'no')
  46. df = pd.read_excel(f"{homeDir}/humanIgSFClusters.ods", engine = "odf",sheet_name='IgSF')
  47. geneIDs = list(map(str,df["GeneID"].values.tolist()))
  48. genes = df["Gene"].values.tolist()
  49. accessions = df["Accession"].values.tolist()
  50. groups = df["Group"].values.tolist()
  51. clusters = list(map(str,df["Cluster"].values.tolist()))
  52. df2 = pd.read_excel(f"{homeDir}/humanIgSFClusters.ods", engine = "odf",sheet_name='Clusters')
  53. clusters2 = df2["Cluster"].values.tolist()
  54. size_u = df2["Size (Unique Genes)"].values.tolist()
  55. subfamily = df2["Subfamily"].values.tolist()
  56. members = df2["Members"].values.tolist()
  57. args = parser.parse_args()
  58. arg1 = args.GeneIDs.split(',')
  59. arg2 = args.GeneNames.split(',')
  60. arg3 = args.GeneAccession.split(',')
  61. arg4 = args.Groups.split(',')
  62. arg5 = args.Clusters.split(',')
  63. arg6 = args.TaxonomicLevel
  64. arg7 = args.CheckAccessions
  65. arg8 = args.seperateClusters
  66. #### Checks for valid inputs
  67. if arg1[0] == 'NA' and arg2[0] == 'NA' and arg3[0] == 'NA' and arg4[0] == 'NA' and arg5[0] == 'NA':
  68. sys.exit(f'Please input at least one GeneID, GeneName, GeneAccession, Group, or Cluster.')
  69. for x in arg1:
  70. if x not in geneIDs and x != 'NA':
  71. sys.exit(f'Invalid GeneID: {x}.')
  72. for x in arg2:
  73. if x not in genes and x != 'NA':
  74. sys.exit(f'Invalid GeneName: {x}.')
  75. for x in arg3:
  76. if x not in accessions and x != 'NA':
  77. sys.exit(f'Invalid GeneAccession: {x}.')
  78. for x in arg4:
  79. if x not in set(groups) and x != 'NA':
  80. sys.exit(f'Invalid Group: {x}. Group must be capatilized')
  81. for x in arg5:
  82. if x not in clusters and x != 'NA':
  83. sys.exit(f'Invalid Cluster: {x}.')
  84. if arg6 not in ['Eukaryotes', 'Gnathostomata', 'Mammalia', 'Primates']:
  85. sys.exit(f'Invalid TaxonomicalLevel: {x}.')
  86. if arg7 != 'yes' and arg7 !='no':
  87. sys.exit(f'Please only input either \"yes\" or \"no\" for the --CheckAccessions argument.')
  88. if arg8 != 'yes' and arg8 !='no':
  89. sys.exit(f'Please only input either \"yes\" or \"no\" for the --seperateClusters argument.')
  90. geneIDs = list(map(int,geneIDs))
  91. clusters = list(map(int,clusters))
  92. if arg1[0] != 'NA':
  93. arg1 = list(map(int,arg1))
  94. if arg5[0] != 'NA':
  95. arg5 = list(map(int,arg5))
  96. ### Create file tag describing inputs
  97. ### If there are many inputs, the tag will instead be the date and time
  98. inputts = []
  99. inputts_black = []
  100. inputts_black2 = []
  101. s1 = ''
  102. s2 = ''
  103. s3 = ''
  104. s4 = ''
  105. s5 = ''
  106. hold = []
  107. for a,b,c,d,e in zip(geneIDs,genes,accessions,groups,clusters):
  108. if a in arg1:
  109. inputts.append(b)
  110. inputts_black.append(b)
  111. if a not in hold:
  112. s1 += str(a)+'-'
  113. hold.append(a)
  114. if b in arg2:
  115. inputts.append(b)
  116. inputts_black.append(b)
  117. if b not in hold:
  118. s2 += b+'-'
  119. hold.append(b)
  120. if c in arg3:
  121. inputts.append(b)
  122. inputts_black.append(b)
  123. if c not in hold:
  124. s3 += c+'-'
  125. hold.append(c)
  126. if d in arg4:
  127. inputts.append(b)
  128. if d not in hold:
  129. s4 += d+'-'
  130. hold.append(d)
  131. if e in arg5:
  132. inputts.append(b)
  133. inputts_black2.append(b)
  134. if e not in hold:
  135. s5 += str(e)+'-'
  136. hold.append(e)
  137. inputts_black2 = list(set(inputts_black2))
  138. if inputts_black == []:
  139. inputts_black = inputts_black2
  140. else:
  141. inputts_black = list(set(inputts_black))
  142. inputts = list(set(inputts))
  143. inputTag = arg6+'_'
  144. if s1 != '':
  145. inputTag += (s1[0:-1]+'_')
  146. if s2 != '':
  147. inputTag += (s2[0:-1]+'_')
  148. if s3 != '':
  149. inputTag += (s3[0:-1]+'_')
  150. if s4 != '':
  151. inputTag += (s4[0:-1]+'_')
  152. if s5 != '':
  153. inputTag += (s5[0:-1]+'_')
  154. inputTag = inputTag[0:-1]
  155. if len(inputTag) > 50:
  156. inputTag = timestamp
  157. geneToGroupDict = {}
  158. group_colors = ['Orange', 'Green', 'Red', 'Purple', 'Brown', 'Pink']
  159. for ge,gr in zip(genes,groups):
  160. geneToGroupDict[ge] = group_colors.index(gr)+1
  161. with open(f'{homeDir}/accessionDescriptionDict.pkl', 'rb') as f:
  162. accessionDescriptionDict = pickle.load(f)
  163. with open(f'{homeDir}/accessionToGeneDict_2.pkl', 'rb') as f:
  164. accessionToGeneDict = pickle.load(f)
  165. exclude = ['GCF_000001405.40']
  166. ### exclude assemblies whose accessions cannot be assigned geneIDs (https://ftp.ncbi.nih.gov/gene/DATA/gene2accession.gz)
  167. file = f'{homeDir}/exclude.txt'
  168. fh = open(file)
  169. for f in fh:
  170. exclude.append(f.strip())
  171. fh.close()
  172. assemblies = []
  173. file = f'{homeDir}/assembliesWithIgSF.txt'
  174. fh = open(file)
  175. for f in fh:
  176. ### exclude taxa with only one assembly
  177. if f.strip() not in ['GCF_019279795.1','GCF_000003605.2','GCF_037176945.1','GCA_003344405.1'] and f.strip() not in exclude:
  178. assemblies.append(f.strip())
  179. fh.close()
  180. #################
  181. if arg6 == 'Eukaryotes':
  182. taxa = ['Amoebozoa', 'Viridiplantae', 'Sar', 'Fungi',
  183. 'Porifera', 'Cnidaria', 'Protostomia', 'Tunicata', 'Echinodermata', 'Cephalochordata', 'Cyclostomata',
  184. 'Chondrichthyes','Actinopteri', 'Cladistia', 'Amphibia', 'Aves', 'Lepidosauria', 'Crocodylia','Testudines', 'Mammalia']
  185. if arg6 == 'Gnathostomata':
  186. taxa = ['Chondrichthyes','Actinopteri', 'Cladistia', 'Amphibia', 'Aves', 'Lepidosauria', 'Crocodylia','Testudines', 'Mammalia']
  187. if arg6 == 'Mammalia':
  188. taxa = [
  189. ## Monotremata (egg-laying mammals)
  190. "Monotremata",
  191. ## Marsupials (pouched mammals)
  192. "Didelphimorphia", # opossums
  193. #"Microbiotheria", # monito del monte
  194. "Dasyuromorphia", # quolls, dunnarts, Tasmanian devil
  195. "Diprotodontia", # kangaroos, koalas, wombats, possums
  196. ### Eutheria (placental mammals)
  197. ## Afrotheria
  198. "Afrosoricida", # tenrecs, golden moles
  199. #"Macroscelidea", # elephant shrews
  200. #"Tubulidentata", # aardvark
  201. "Proboscidea", # elephants
  202. #"Sirenia", # manatees, dugongs
  203. ## Xenarthra
  204. #"Cingulata", # armadillos
  205. #"Pilosa", # sloths and anteaters
  206. ## Laurasiatheria
  207. "Eulipotyphla", # shrews, moles, hedgehogs
  208. "Chiroptera", # bats
  209. "Pholidota", # pangolins
  210. "Carnivora", # cats, dogs, seals, etc.
  211. "Perissodactyla", # horses, rhinos, tapirs
  212. "Cetacea", # whales, dolphins, porpoises
  213. "Artiodactyla", # pigs, deer, cows, etc.
  214. ## Euarchontoglires
  215. "Lagomorpha", # rabbits, hares, pikas
  216. "Rodentia", # rodents
  217. #"Scandentia", # treeshrews
  218. "Dermoptera", # colugos
  219. "Primates" # monkeys, apes, humans
  220. ]
  221. if arg6 == 'Primates':
  222. taxa = [
  223. ## Platyrrhines (New World monkeys)
  224. "Cebidae",
  225. ## Catarrhines (Old World monkeys & apes)
  226. "Cercopithecidae", # Old World monkeys
  227. "Hylobatidae", # gibbons
  228. "Hominidae" # great apes (including humans)
  229. ]
  230. newAssemblies = []
  231. taxDict = {}
  232. file = f'{homeDir}/eukaryotaTaxonomy.tsv'
  233. fh = open(file)
  234. for f in fh:
  235. f = f.split('\t')
  236. if f[0] in assemblies:
  237. for t in taxa:
  238. if t+',' in f[1]:
  239. taxDict[f[0]] = t
  240. newAssemblies.append(f[0])
  241. if t == 'Cetacea' and arg6 == 'Mammalia':
  242. break
  243. fh.close()
  244. ########## Create Gene Conservation Figure
  245. print('Creating Gene Conservation Figure')
  246. df3 = pd.read_excel(f"{homeDir}/gene_conservation.ods", engine = "odf",sheet_name=arg6)
  247. geneMatrix = []
  248. for col in taxa:
  249. col_values = (df3[col].astype(float)).tolist()
  250. holdG = df3['Gene'].tolist()
  251. geneMatrix.append(col_values)
  252. # plt.plot([1,2],[-1,-1],colors[0])
  253. plt.figure(figsize=(8, 4))
  254. for c,g in enumerate(holdG):
  255. if g in inputts:
  256. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
  257. grays = ['0','0.35','0.175']
  258. cc = 0
  259. for c,g in enumerate(holdG):
  260. if g in inputts_black:
  261. if arg4[0] == 'NA':
  262. if len(inputts_black) > 10:
  263. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
  264. elif inputts_black == inputts_black2:
  265. if cc == 10:
  266. cc -= 10
  267. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[cc],label=g)
  268. else:
  269. if cc == 3:
  270. cc -= 3
  271. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=grays[cc],label=g)
  272. cc += 1
  273. else:
  274. if cc == 3:
  275. cc -= 3
  276. if len(inputts_black) > 10:
  277. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
  278. else:
  279. plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=grays[cc],label=g)
  280. cc += 1
  281. labelLines(plt.gca().get_lines(), align=True, fontsize=10)
  282. plt.ylim([-0.08,1.08])
  283. plt.ylabel('Average Normalized Bit Score',fontsize=12)
  284. plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
  285. plt.savefig(f'{homeDir}/results/gene_conservation_{inputTag}.png', bbox_inches='tight',dpi=600) #################
  286. # plt.show()
  287. plt.close()
  288. ##################### create change in IgSF size figure
  289. sizeDiff = []
  290. clusterAccessions = []
  291. otherClusterAccessions = []
  292. for a in newAssemblies:
  293. s = int(subprocess.check_output(f"cut -f2 {homeDir}/clusteredIgSF_withHuman/{a}_clusters | sort -g | uniq | wc -l", shell=True))
  294. holdHuman = s*[0]
  295. holdOther = s*[0]
  296. holdClusterAccessions = []
  297. holdOtherClusterAccessions = []
  298. for hh in holdOther:
  299. holdClusterAccessions.append([])
  300. holdOtherClusterAccessions.append([])
  301. holdGenes = []
  302. file = f'{homeDir}/clusteredIgSF_withHuman/{a}_clusters'
  303. fh = open(file)
  304. for f in fh:
  305. f = f.strip()
  306. f = f.split()
  307. if accessionToGeneDict[f[0]] not in holdGenes:
  308. holdGenes.append(accessionToGeneDict[f[0]])
  309. if f[0] in accessions:
  310. holdHuman[int(f[1])] += 1
  311. holdClusterAccessions[int(f[1])].append(f[0])
  312. else:
  313. holdOther[int(f[1])] += 1
  314. holdOtherClusterAccessions[int(f[1])].append(f[0])
  315. fh.close()
  316. clusterAccessions.append(holdClusterAccessions)
  317. otherClusterAccessions.append(holdOtherClusterAccessions)
  318. sizeDiff.append(list(np.array(holdHuman)-np.array(holdOther)))
  319. print('Creating Change in IgSF Cluster Size Figure')
  320. if arg8 == 'no':
  321. queryClusts = []
  322. # queryClustsLabels = []
  323. for i in inputts:
  324. if int(clusters[genes.index(i)]) not in queryClusts:
  325. queryClusts.append(int(clusters[genes.index(i)]))
  326. # queryClustsLabels.append(members[int(clusters[gene.index(i)])])
  327. print(f'Cluster {clusters[genes.index(i)]} | {subfamily[int(clusters[genes.index(i)])]} | {members[int(clusters[genes.index(i)])]}')
  328. queryAccessions = []
  329. for a,b in zip(clusters,accessions):
  330. if int(a) in queryClusts:
  331. queryAccessions.append(b)
  332. ### Adds all clusters in the combined set where at least one human query gene appears
  333. assemblyEquivClusts = []
  334. for ca in clusterAccessions:
  335. equivClusts = set()
  336. for qa in queryAccessions:
  337. for c,a in enumerate(ca):
  338. if qa in a:
  339. equivClusts.add(c)
  340. assemblyEquivClusts.append(list(sorted(equivClusts)))
  341. ### Creates .tsv file showing queried human and other species clustered IgSF accessions with descriptions
  342. if arg7 == 'yes':
  343. newFile = open(f'{homeDir}/results/clusteredIgSF_{inputTag}.tsv','w')
  344. newFile.write('Taxa\tAssembly\tGeneID\tSpecies\tIntendedGene\tAccessionDescription\n')
  345. c = 0
  346. for a1,a2,b in zip(clusterAccessions,otherClusterAccessions,assemblyEquivClusts):
  347. for bb in b:
  348. holdClusters = []
  349. s = ''
  350. for x in a1[bb]:
  351. if clusters[accessions.index(x)] not in holdClusters:
  352. holdClusters.append(clusters[accessions.index(x)])
  353. s += f'Cluster {str(clusters[accessions.index(x)])} | {members[clusters2.index(clusters[accessions.index(x)])]}, '
  354. newFile.write(f'{s[0:-2]}\n')
  355. for x in a1[bb]:
  356. newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tHuman\t{accessionDescriptionDict[x]}\n')
  357. newFile.write('-\t-\t-\t-\t-\n')
  358. for x in a2[bb]:
  359. newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tOther\t{accessionDescriptionDict[x]}\n')
  360. newFile.write('-\t-\t-\t-\t-\n')
  361. newFile.write(f'-\t-\t-\t-\t-\n')
  362. newFile.write(f'-\t-\t-\t-\t-\n')
  363. c += 1
  364. newFile.close()
  365. taxaSizeDiff = []
  366. for t in taxa:
  367. taxaSizeDiff.append([])
  368. c = 0
  369. for aec,sd in zip(assemblyEquivClusts,sizeDiff):
  370. hold = 0
  371. for ii in aec:
  372. hold += sd[ii]
  373. taxaSizeDiff[taxa.index(taxDict[newAssemblies[c]])].append(int(hold))
  374. c += 1
  375. r = []
  376. rstd = []
  377. for tsd in taxaSizeDiff:
  378. r.append(statistics.mean(tsd))
  379. if len(tsd) < 2:
  380. rstd.append(0)
  381. else:
  382. rstd.append(statistics.stdev(tsd))
  383. alpha = 0.8
  384. cap_size = 0.3
  385. plt.figure(figsize=(8, 4))
  386. plt.plot(range(len(taxa)),r,c=colors[0])
  387. for xi, yi, err in zip(range(0,len(taxa)),r,rstd):
  388. if err != 0:
  389. plt.vlines(xi, yi - err, yi + err,colors=colors[0],linestyles='dashed',alpha=alpha,linewidth=1)
  390. plt.hlines([yi - err, yi + err],xi - cap_size, xi + cap_size,colors=colors[0],alpha=alpha,linewidth=1)
  391. plt.axhline(y=0, color='k', linestyle='--')
  392. plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
  393. plt.ylabel('Δ (Human – Avg. Taxa) Cluster Size',fontsize=12)
  394. plt.savefig(f'{homeDir}/results/IgSFSize_{inputTag}.png', bbox_inches='tight',dpi=600)
  395. # plt.show()
  396. plt.close()
  397. else:
  398. queryClusts = []
  399. color_count = 0
  400. plt.figure(figsize=(8, 4))
  401. for i in inputts:
  402. if color_count == 10:
  403. color_count -= 10
  404. if int(clusters[genes.index(i)]) not in queryClusts:
  405. queryClusts.append(int(clusters[genes.index(i)]))
  406. print(f'Cluster {clusters[genes.index(i)]} | {subfamily[int(clusters[genes.index(i)])]} | {members[int(clusters[genes.index(i)])]}')
  407. queryAccessions = []
  408. for a,b in zip(clusters,accessions):
  409. if int(a) == int(clusters[genes.index(i)]):
  410. queryAccessions.append(b)
  411. ### Adds all clusters in the combined set where at least one human query gene appears
  412. assemblyEquivClusts = []
  413. for ca in clusterAccessions:
  414. equivClusts = set()
  415. for qa in queryAccessions:
  416. for c,a in enumerate(ca):
  417. if qa in a:
  418. equivClusts.add(c)
  419. assemblyEquivClusts.append(list(sorted(equivClusts)))
  420. ### Creates .tsv file showing queried human and other species clustered IgSF accessions with descriptions
  421. if arg7 == 'yes':
  422. newFile = open(f'{homeDir}/results/clusteredIgSF_Cluster-{int(clusters[genes.index(i)])}.tsv','w')
  423. newFile.write('Taxa\tAssembly\tGeneID\tSpecies\tIntendedGene\tAccessionDescription\n')
  424. c = 0
  425. for a1,a2,b in zip(clusterAccessions,otherClusterAccessions,assemblyEquivClusts):
  426. for bb in b:
  427. holdClusters = []
  428. s = ''
  429. for x in a1[bb]:
  430. if clusters[accessions.index(x)] not in holdClusters:
  431. holdClusters.append(clusters[accessions.index(x)])
  432. s += f'Cluster {str(clusters[accessions.index(x)])} | {members[clusters2.index(clusters[accessions.index(x)])]}, '
  433. newFile.write(f'{s[0:-2]}\n')
  434. for x in a1[bb]:
  435. newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tHuman\t{accessionDescriptionDict[x]}\n')
  436. newFile.write('-\t-\t-\t-\t-\n')
  437. for x in a2[bb]:
  438. newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tOther\t{accessionDescriptionDict[x]}\n')
  439. newFile.write('-\t-\t-\t-\t-\n')
  440. newFile.write(f'-\t-\t-\t-\t-\n')
  441. newFile.write(f'-\t-\t-\t-\t-\n')
  442. c += 1
  443. newFile.close()
  444. taxaSizeDiff = []
  445. for t in taxa:
  446. taxaSizeDiff.append([])
  447. c = 0
  448. for aec,sd in zip(assemblyEquivClusts,sizeDiff):
  449. hold = 0
  450. for ii in aec:
  451. hold += sd[ii]
  452. taxaSizeDiff[taxa.index(taxDict[newAssemblies[c]])].append(int(hold))
  453. c += 1
  454. r = []
  455. rstd = []
  456. for tsd in taxaSizeDiff:
  457. r.append(statistics.mean(tsd))
  458. if len(tsd) < 2:
  459. rstd.append(0)
  460. else:
  461. rstd.append(statistics.stdev(tsd))
  462. alpha = 0.8
  463. cap_size = 0.3
  464. plt.plot(range(len(taxa)),r,c=colors[color_count],label=subfamily[int(clusters[genes.index(i)])])
  465. for xi, yi, err in zip(range(0,len(taxa)),r,rstd):
  466. plt.vlines(xi, yi - err, yi + err,colors=colors[color_count],linestyles='dashed',alpha=alpha,linewidth=1)
  467. plt.hlines([yi - err, yi + err],xi - cap_size, xi + cap_size,colors=colors[color_count],alpha=alpha,linewidth=1)
  468. color_count += 1
  469. labelLines(plt.gca().get_lines(), align=True, fontsize=10)
  470. plt.axhline(y=0, color='k', linestyle='--')
  471. plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
  472. plt.ylabel('Δ (Human – Avg. Taxa) Cluster Size',fontsize=12)
  473. plt.savefig(f'{homeDir}/results/IgSFSize_{inputTag}.png', bbox_inches='tight',dpi=600)
  474. # plt.show()

conservation-IgSFSize.py at commit af03288, no license · at the source

Overview

  1. Department of Systems & Computational Biology, Albert Einstein College of Medicine, Bronx, NY 10461, USA
  2. Department of Biochemistry, Albert Einstein College of Medicine, Bronx, NY 10461, USA
Institutions: Albert Einstein College of Medicine (United States)
Journal: Genome biology and evolution, volume 18, issue 6, article evag133
Dates: accepted 25 May 2026; published online 2 June 2026; in print June 2026
Type: Research article · Language: English
License: CC BY-NC
Identifiers: DOI 10.1093/gbe/evag133 · PMID 42230317 · PMCID PMC13262535 · OpenAlex W7163206163
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cellular / molecular (subfield)
Methods: Machine learning
Keywords: extracellular IgSF, gnathostomes, evolution of adaptive immune system, evolution of co-stimulatory immune receptors
MeSH: Evolution, Molecular*, Immunoglobulins*, Animals, Conserved Sequence, Humans, Phylogeny (* major topic)
Topic: T-cell and B-cell Immunology (Immunology, Immunology and Microbiology), according to OpenAlex
Funding: NIH (GM136357)
Citations: not cited yet (Europe PMC); 107 references in the paper

Abstract

The human immunoglobulin superfamily (IgSF) encompasses hundreds of proteins involved in cell–cell adhesion, neural connectivity, junctional organization, and immune regulation, with many serving as key checkpoint proteins. To elucidate the evolutionary history of human IgSF members, we systematically analyzed all available eukaryotic reference genomes to determine when each IgSF subfamily first appeared. The human IgSF was partitioned into six major evolutionary timeframes: Metazoa, Vertebrata, Gnathostomata, Tetrapoda, Amniota, and Mammalia. Although proteins were grouped solely by their conservation across eukaryotes, their biological functions clustered naturally, reflecting how new physiological systems create selective pressures that drive the appearance, retention, and diversification of protein architectures suited to those functions. Conservation and functional analyses indicate that human IgSF members arising in tetrapods and amniotes primary regulate the strength and duration of immune responses and form many of the critical components of the immune synapse, while IgSF genes that appear first in mammals have evolved to fine-tune immune activation thresholds to support maternal–fetal tolerance. Case studies are provided to illustrate three key evolutionary themes: (i) highly conserved yet catalytically inactive proteins retain essential regulatory functions, (ii) functional convergence among evolutionarily distinct IgSF families, and (iii) compensatory evolution within adaptive immunity following a lineage-specific loss of an IgSF member. Together, these findings establish an evolutionary framework for organizing the human IgSF by both ancestry and function, providing a foundation for assessing IgSF importance. Notably, this study facilitates the identification of conserved but understudied proteins that emerged alongside the development of the adaptive immune system, highlighting them as promising candidates for future experimental investigation.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

StevenGrudman/IgSF-Conservation

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: af032888ebd9488e70430d119a5f9e2cb71bdd7c, 25 November 2025
Languages: Python (4), Shell (2)
Size: 1,172 files, 6 scripts
Software Heritage: not archived
Found in: the text, “Discussion”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (2 files), NumPy (2 files), pandas (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
7 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 6 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability

The data underlying this article can be found in the online Supplementary material. In addition, all data and a program that can facilitate browsing results are publicly accessible at https://github.com/StevenGrudman/IgSF-Conservation.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 4 keywords, 6 MeSH terms, 1 funder, 105 references.

Cite

This paper

Grudman, S., & Fiser, A. (2026). Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution. Genome biology and evolution, 18(6), evag133. https://doi.org/10.1093/gbe/evag133

BibTeX

@article{grudman2026conservation,
author = {Grudman, Steven and Fiser, Andras},
title = {{Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution}},
journal = {Genome biology and evolution},
year = {2026},
month = jun,
volume = {18},
number = {6},
pages = {evag133},
publisher = {Oxford University Press},
issn = {1759-6653},
doi = {10.1093/gbe/evag133},
url = {https://doi.org/10.1093/gbe/evag133},
pmid = {42230317},
pmcid = {PMC13262535}
}

RIS

TY - JOUR
AU - Grudman, Steven
AU - Fiser, Andras
TI - Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution
T2 - Genome biology and evolution
J2 - Genome Biol Evol
PY - 2026
DA - 2026/06/01
VL - 18
IS - 6
SP - evag133
SN - 1759-6653
PB - Oxford University Press
DO - 10.1093/gbe/evag133
UR - https://doi.org/10.1093/gbe/evag133
LA - en
ER -

CSL-JSON

{
"id": "10.1093/gbe/evag133",
"type": "article-journal",
"title": "Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution",
"container-title": "Genome biology and evolution",
"author": [
{
"family": "Grudman",
"given": "Steven"
},
{
"family": "Fiser",
"given": "Andras"
}
],
"container-title-short": "Genome Biol Evol",
"volume": "18",
"issue": "6",
"page": "evag133",
"DOI": "10.1093/gbe/evag133",
"PMID": "42230317",
"PMCID": "PMC13262535",
"ISSN": "1759-6653",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/gbe/evag133",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.7554/elife.107393 [code]
Chromosome-scale genome assembly of the European common cuttlefish &lt;i&gt;Sepia officinalis&lt;/i&gt;.
Journal: eLife
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 3 references
[2] doi:10.1016/j.xgen.2026.101217 [code]
ProtoCloud: A prototypical self-explaining model for single-cell analysis.
Journal: Cell genomics
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
[3] doi:10.1016/j.molcel.2026.07.006 [code]
DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.
Journal: Molecular cell
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
[4] doi:10.1038/s41593-026-02321-0 [code]
Microbial reactivation of host androgens directs enteric neuronal regulation of gut motility.
Journal: Nature neuroscience
In common: pandas, SciPy, NumPy, cellular / molecular, 2 references
[5] doi:10.1016/j.nbscr.2026.100149 [code]
Molecular correlates of sleep deprivation in the mouse brain identified by meta-analysis of microarray data.
Journal: Neurobiology of sleep and circadian rhythms
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
[6] doi:10.7554/elife.106134 [code]
Identification and classification of ion channels across the tree of life provide functional insights into understudied CALHM channels.
Journal: eLife
In common: pandas, Matplotlib, cellular / molecular, 3 references
[7] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 1 reference
[8] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 1 reference
[9] doi:10.1002/advs.202523984 [code]
INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: pandas, SciPy, Matplotlib, 1 other tool, 1 reference
[10] doi:10.1523/eneuro.0362-25.2026 [code]
Similarities between &lt;i&gt;Ciona&lt;/i&gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
Journal: eNeuro
In common: pandas, SciPy, Matplotlib, 1 other tool, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.