Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution.
The 2 matches
- [1] § Results › CD58 ↔ conservation-IgSFSize.py, lines 203–273 · score 0.88 · egg laying mammals, placental mammals, Dermoptera, Euarchontoglires, Eutheria, Lagomorpha
- [2] § Materials and Methods › Clustering of IgSF ↔ otherScripts/brotherhoodMethod.py, lines 80–91 · score 0.64 · squareform, weighted, correlation, linkage, metric, cut
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 532 lines · 18 KB · no license · 1 match
- import os
- import sys
- import argparse
- import statistics
- import subprocess
- import pickle
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- from labellines import labelLines
- import warnings
- from datetime import datetime
- homeDir = os.getcwd()
- colors = plt.get_cmap('tab10').colors
- os.makedirs(f'{homeDir}/results', exist_ok=True)
- warnings.filterwarnings("ignore", message="Tried to label line")
- now = datetime.now()
- timestamp = now.strftime("%m-%d-%Y__%H-%M-%S")
- ##### Need to update DOI
- descr = '''
- This script provides two figures:
- Gene Conservation:
- Genes input by their GeneIDs, GeneIDs, GeneAccession, or Cluster arguments will be shown in shades of gray.
- Genes input by their Group will be colored by their Group.
- Change in IgSF Cluster Size:
- If an indiviudal gene is input, the entire cluster this gene belongs to will be considered.
- If multiple genes are input, all clusters including these genes will be considered.
- By default each all clusters are combined.
- However, each change in cluster size can be seperated with the seperateClusters argument.
- Please limit the inputs or else the figure becomes very busy
- To investigate the gene accessions considered for this analysis, you can check our work with the CheckAccessions argument.
- We recommend setting the seperateClusters argument to yes if you intend to use the CheckAccessions argument as the output is easier to interpret.
- You must provide at least one GeneID, GeneName, GeneAccession, Group, or Cluster.
- This script can be run across different taxonomic levels (Eukaryotes, Gnathostomata, Mammalia, and Primates)
- Multiple entries can be supplied for each above argument by separating them with commas.
- '''
- parser = argparse.ArgumentParser(formatter_class=argparse.RawTextHelpFormatter, description = descr)
- parser.add_argument("-gi", "--GeneIDs", help = "Please refer to humanIgSFClusters.ods for acceptable GeneIDs.", default = 'NA')
- parser.add_argument("-gn", "--GeneNames", help = "Please refer to humanIgSFClusters.ods for acceptable Gene Names.", default = 'NA')
- parser.add_argument("-ga", "--GeneAccession", help = "Please refer to humanIgSFClusters.ods for acceptable Gene Accessions", default = 'NA')
- parser.add_argument("-gr", "--Groups", help = "Options: Orange, Green, Red, Purple, Brown, Pink. Please refer to Fig xx", default = 'NA') ##### Need to update figure number
- parser.add_argument("-cl", "--Clusters", help = "Options: 0-132. Please refer to humanIgSFClusters.ods", default = 'NA')
- parser.add_argument("-tl", "--TaxonomicLevel", help = "Options: Eukaryotes, Gnathostomata, Mammalia, Primates. Default = Gnathostomata", default = 'Gnathostomata')
- parser.add_argument("-ca", "--CheckAccessions", help = "Create .tsv file showing queried human and other species clustered IgSF accessions with descriptions. Options: yes or no. Default = no", default = 'no')
- parser.add_argument("-sc", "--seperateClusters", help = "Seperates clusters in change in IgSF Cluster Size figure. Options: yes or no. Default = no", default = 'no')
- df = pd.read_excel(f"{homeDir}/humanIgSFClusters.ods", engine = "odf",sheet_name='IgSF')
- geneIDs = list(map(str,df["GeneID"].values.tolist()))
- genes = df["Gene"].values.tolist()
- accessions = df["Accession"].values.tolist()
- groups = df["Group"].values.tolist()
- clusters = list(map(str,df["Cluster"].values.tolist()))
- df2 = pd.read_excel(f"{homeDir}/humanIgSFClusters.ods", engine = "odf",sheet_name='Clusters')
- clusters2 = df2["Cluster"].values.tolist()
- size_u = df2["Size (Unique Genes)"].values.tolist()
- subfamily = df2["Subfamily"].values.tolist()
- members = df2["Members"].values.tolist()
- args = parser.parse_args()
- arg1 = args.GeneIDs.split(',')
- arg2 = args.GeneNames.split(',')
- arg3 = args.GeneAccession.split(',')
- arg4 = args.Groups.split(',')
- arg5 = args.Clusters.split(',')
- arg6 = args.TaxonomicLevel
- arg7 = args.CheckAccessions
- arg8 = args.seperateClusters
- #### Checks for valid inputs
- if arg1[0] == 'NA' and arg2[0] == 'NA' and arg3[0] == 'NA' and arg4[0] == 'NA' and arg5[0] == 'NA':
- sys.exit(f'Please input at least one GeneID, GeneName, GeneAccession, Group, or Cluster.')
- for x in arg1:
- if x not in geneIDs and x != 'NA':
- sys.exit(f'Invalid GeneID: {x}.')
- for x in arg2:
- if x not in genes and x != 'NA':
- sys.exit(f'Invalid GeneName: {x}.')
- for x in arg3:
- if x not in accessions and x != 'NA':
- sys.exit(f'Invalid GeneAccession: {x}.')
- for x in arg4:
- if x not in set(groups) and x != 'NA':
- sys.exit(f'Invalid Group: {x}. Group must be capatilized')
- for x in arg5:
- if x not in clusters and x != 'NA':
- sys.exit(f'Invalid Cluster: {x}.')
- if arg6 not in ['Eukaryotes', 'Gnathostomata', 'Mammalia', 'Primates']:
- sys.exit(f'Invalid TaxonomicalLevel: {x}.')
- if arg7 != 'yes' and arg7 !='no':
- sys.exit(f'Please only input either \"yes\" or \"no\" for the --CheckAccessions argument.')
- if arg8 != 'yes' and arg8 !='no':
- sys.exit(f'Please only input either \"yes\" or \"no\" for the --seperateClusters argument.')
- geneIDs = list(map(int,geneIDs))
- clusters = list(map(int,clusters))
- if arg1[0] != 'NA':
- arg1 = list(map(int,arg1))
- if arg5[0] != 'NA':
- arg5 = list(map(int,arg5))
- ### Create file tag describing inputs
- ### If there are many inputs, the tag will instead be the date and time
- inputts = []
- inputts_black = []
- inputts_black2 = []
- s1 = ''
- s2 = ''
- s3 = ''
- s4 = ''
- s5 = ''
- hold = []
- for a,b,c,d,e in zip(geneIDs,genes,accessions,groups,clusters):
- if a in arg1:
- inputts.append(b)
- inputts_black.append(b)
- if a not in hold:
- s1 += str(a)+'-'
- hold.append(a)
- if b in arg2:
- inputts.append(b)
- inputts_black.append(b)
- if b not in hold:
- s2 += b+'-'
- hold.append(b)
- if c in arg3:
- inputts.append(b)
- inputts_black.append(b)
- if c not in hold:
- s3 += c+'-'
- hold.append(c)
- if d in arg4:
- inputts.append(b)
- if d not in hold:
- s4 += d+'-'
- hold.append(d)
- if e in arg5:
- inputts.append(b)
- inputts_black2.append(b)
- if e not in hold:
- s5 += str(e)+'-'
- hold.append(e)
- inputts_black2 = list(set(inputts_black2))
- if inputts_black == []:
- inputts_black = inputts_black2
- else:
- inputts_black = list(set(inputts_black))
- inputts = list(set(inputts))
- inputTag = arg6+'_'
- if s1 != '':
- inputTag += (s1[0:-1]+'_')
- if s2 != '':
- inputTag += (s2[0:-1]+'_')
- if s3 != '':
- inputTag += (s3[0:-1]+'_')
- if s4 != '':
- inputTag += (s4[0:-1]+'_')
- if s5 != '':
- inputTag += (s5[0:-1]+'_')
- inputTag = inputTag[0:-1]
- if len(inputTag) > 50:
- inputTag = timestamp
- geneToGroupDict = {}
- group_colors = ['Orange', 'Green', 'Red', 'Purple', 'Brown', 'Pink']
- for ge,gr in zip(genes,groups):
- geneToGroupDict[ge] = group_colors.index(gr)+1
- with open(f'{homeDir}/accessionDescriptionDict.pkl', 'rb') as f:
- accessionDescriptionDict = pickle.load(f)
- with open(f'{homeDir}/accessionToGeneDict_2.pkl', 'rb') as f:
- accessionToGeneDict = pickle.load(f)
- exclude = ['GCF_000001405.40']
- ### exclude assemblies whose accessions cannot be assigned geneIDs (https://ftp.ncbi.nih.gov/gene/DATA/gene2accession.gz)
- file = f'{homeDir}/exclude.txt'
- fh = open(file)
- for f in fh:
- exclude.append(f.strip())
- fh.close()
- assemblies = []
- file = f'{homeDir}/assembliesWithIgSF.txt'
- fh = open(file)
- for f in fh:
- ### exclude taxa with only one assembly
- if f.strip() not in ['GCF_019279795.1','GCF_000003605.2','GCF_037176945.1','GCA_003344405.1'] and f.strip() not in exclude:
- assemblies.append(f.strip())
- fh.close()
- #################
- if arg6 == 'Eukaryotes':
- taxa = ['Amoebozoa', 'Viridiplantae', 'Sar', 'Fungi',
- 'Porifera', 'Cnidaria', 'Protostomia', 'Tunicata', 'Echinodermata', 'Cephalochordata', 'Cyclostomata',
- 'Chondrichthyes','Actinopteri', 'Cladistia', 'Amphibia', 'Aves', 'Lepidosauria', 'Crocodylia','Testudines', 'Mammalia']
- if arg6 == 'Gnathostomata':
- taxa = ['Chondrichthyes','Actinopteri', 'Cladistia', 'Amphibia', 'Aves', 'Lepidosauria', 'Crocodylia','Testudines', 'Mammalia']
- if arg6 == 'Mammalia':
- taxa = [
- ## Monotremata (egg-laying mammals)
- "Monotremata",
- ## Marsupials (pouched mammals)
- "Didelphimorphia", # opossums
- #"Microbiotheria", # monito del monte
- "Dasyuromorphia", # quolls, dunnarts, Tasmanian devil
- "Diprotodontia", # kangaroos, koalas, wombats, possums
- ### Eutheria (placental mammals)
- ## Afrotheria
- "Afrosoricida", # tenrecs, golden moles
- #"Macroscelidea", # elephant shrews
- #"Tubulidentata", # aardvark
- "Proboscidea", # elephants
- #"Sirenia", # manatees, dugongs
- ## Xenarthra
- #"Cingulata", # armadillos
- #"Pilosa", # sloths and anteaters
- ## Laurasiatheria
- "Eulipotyphla", # shrews, moles, hedgehogs
- "Chiroptera", # bats
- "Pholidota", # pangolins
- "Carnivora", # cats, dogs, seals, etc.
- "Perissodactyla", # horses, rhinos, tapirs
- "Cetacea", # whales, dolphins, porpoises
- "Artiodactyla", # pigs, deer, cows, etc.
- ## Euarchontoglires
- "Lagomorpha", # rabbits, hares, pikas
- "Rodentia", # rodents
- #"Scandentia", # treeshrews
- "Dermoptera", # colugos
- "Primates" # monkeys, apes, humans
- ]
- if arg6 == 'Primates':
- taxa = [
- ## Platyrrhines (New World monkeys)
- "Cebidae",
- ## Catarrhines (Old World monkeys & apes)
- "Cercopithecidae", # Old World monkeys
- "Hylobatidae", # gibbons
- "Hominidae" # great apes (including humans)
- ]
- newAssemblies = []
- taxDict = {}
- file = f'{homeDir}/eukaryotaTaxonomy.tsv'
- fh = open(file)
- for f in fh:
- f = f.split('\t')
- if f[0] in assemblies:
- for t in taxa:
- if t+',' in f[1]:
- taxDict[f[0]] = t
- newAssemblies.append(f[0])
- if t == 'Cetacea' and arg6 == 'Mammalia':
- break
- fh.close()
- ########## Create Gene Conservation Figure
- print('Creating Gene Conservation Figure')
- df3 = pd.read_excel(f"{homeDir}/gene_conservation.ods", engine = "odf",sheet_name=arg6)
- geneMatrix = []
- for col in taxa:
- col_values = (df3[col].astype(float)).tolist()
- holdG = df3['Gene'].tolist()
- geneMatrix.append(col_values)
- # plt.plot([1,2],[-1,-1],colors[0])
- plt.figure(figsize=(8, 4))
- for c,g in enumerate(holdG):
- if g in inputts:
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
- grays = ['0','0.35','0.175']
- cc = 0
- for c,g in enumerate(holdG):
- if g in inputts_black:
- if arg4[0] == 'NA':
- if len(inputts_black) > 10:
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
- elif inputts_black == inputts_black2:
- if cc == 10:
- cc -= 10
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[cc],label=g)
- else:
- if cc == 3:
- cc -= 3
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=grays[cc],label=g)
- cc += 1
- else:
- if cc == 3:
- cc -= 3
- if len(inputts_black) > 10:
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=colors[geneToGroupDict[g]])
- else:
- plt.plot(taxa,list(np.array(geneMatrix).T)[c],c=grays[cc],label=g)
- cc += 1
- labelLines(plt.gca().get_lines(), align=True, fontsize=10)
- plt.ylim([-0.08,1.08])
- plt.ylabel('Average Normalized Bit Score',fontsize=12)
- plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
- plt.savefig(f'{homeDir}/results/gene_conservation_{inputTag}.png', bbox_inches='tight',dpi=600) #################
- # plt.show()
- plt.close()
- ##################### create change in IgSF size figure
- sizeDiff = []
- clusterAccessions = []
- otherClusterAccessions = []
- for a in newAssemblies:
- s = int(subprocess.check_output(f"cut -f2 {homeDir}/clusteredIgSF_withHuman/{a}_clusters | sort -g | uniq | wc -l", shell=True))
- holdHuman = s*[0]
- holdOther = s*[0]
- holdClusterAccessions = []
- holdOtherClusterAccessions = []
- for hh in holdOther:
- holdClusterAccessions.append([])
- holdOtherClusterAccessions.append([])
- holdGenes = []
- file = f'{homeDir}/clusteredIgSF_withHuman/{a}_clusters'
- fh = open(file)
- for f in fh:
- f = f.strip()
- f = f.split()
- if accessionToGeneDict[f[0]] not in holdGenes:
- holdGenes.append(accessionToGeneDict[f[0]])
- if f[0] in accessions:
- holdHuman[int(f[1])] += 1
- holdClusterAccessions[int(f[1])].append(f[0])
- else:
- holdOther[int(f[1])] += 1
- holdOtherClusterAccessions[int(f[1])].append(f[0])
- fh.close()
- clusterAccessions.append(holdClusterAccessions)
- otherClusterAccessions.append(holdOtherClusterAccessions)
- sizeDiff.append(list(np.array(holdHuman)-np.array(holdOther)))
- print('Creating Change in IgSF Cluster Size Figure')
- if arg8 == 'no':
- queryClusts = []
- # queryClustsLabels = []
- for i in inputts:
- if int(clusters[genes.index(i)]) not in queryClusts:
- queryClusts.append(int(clusters[genes.index(i)]))
- # queryClustsLabels.append(members[int(clusters[gene.index(i)])])
- print(f'Cluster {clusters[genes.index(i)]} | {subfamily[int(clusters[genes.index(i)])]} | {members[int(clusters[genes.index(i)])]}')
- queryAccessions = []
- for a,b in zip(clusters,accessions):
- if int(a) in queryClusts:
- queryAccessions.append(b)
- ### Adds all clusters in the combined set where at least one human query gene appears
- assemblyEquivClusts = []
- for ca in clusterAccessions:
- equivClusts = set()
- for qa in queryAccessions:
- for c,a in enumerate(ca):
- if qa in a:
- equivClusts.add(c)
- assemblyEquivClusts.append(list(sorted(equivClusts)))
- ### Creates .tsv file showing queried human and other species clustered IgSF accessions with descriptions
- if arg7 == 'yes':
- newFile = open(f'{homeDir}/results/clusteredIgSF_{inputTag}.tsv','w')
- newFile.write('Taxa\tAssembly\tGeneID\tSpecies\tIntendedGene\tAccessionDescription\n')
- c = 0
- for a1,a2,b in zip(clusterAccessions,otherClusterAccessions,assemblyEquivClusts):
- for bb in b:
- holdClusters = []
- s = ''
- for x in a1[bb]:
- if clusters[accessions.index(x)] not in holdClusters:
- holdClusters.append(clusters[accessions.index(x)])
- s += f'Cluster {str(clusters[accessions.index(x)])} | {members[clusters2.index(clusters[accessions.index(x)])]}, '
- newFile.write(f'{s[0:-2]}\n')
- for x in a1[bb]:
- newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tHuman\t{accessionDescriptionDict[x]}\n')
- newFile.write('-\t-\t-\t-\t-\n')
- for x in a2[bb]:
- newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tOther\t{accessionDescriptionDict[x]}\n')
- newFile.write('-\t-\t-\t-\t-\n')
- newFile.write(f'-\t-\t-\t-\t-\n')
- newFile.write(f'-\t-\t-\t-\t-\n')
- c += 1
- newFile.close()
- taxaSizeDiff = []
- for t in taxa:
- taxaSizeDiff.append([])
- c = 0
- for aec,sd in zip(assemblyEquivClusts,sizeDiff):
- hold = 0
- for ii in aec:
- hold += sd[ii]
- taxaSizeDiff[taxa.index(taxDict[newAssemblies[c]])].append(int(hold))
- c += 1
- r = []
- rstd = []
- for tsd in taxaSizeDiff:
- r.append(statistics.mean(tsd))
- if len(tsd) < 2:
- rstd.append(0)
- else:
- rstd.append(statistics.stdev(tsd))
- alpha = 0.8
- cap_size = 0.3
- plt.figure(figsize=(8, 4))
- plt.plot(range(len(taxa)),r,c=colors[0])
- for xi, yi, err in zip(range(0,len(taxa)),r,rstd):
- if err != 0:
- plt.vlines(xi, yi - err, yi + err,colors=colors[0],linestyles='dashed',alpha=alpha,linewidth=1)
- plt.hlines([yi - err, yi + err],xi - cap_size, xi + cap_size,colors=colors[0],alpha=alpha,linewidth=1)
- plt.axhline(y=0, color='k', linestyle='--')
- plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
- plt.ylabel('Δ (Human – Avg. Taxa) Cluster Size',fontsize=12)
- plt.savefig(f'{homeDir}/results/IgSFSize_{inputTag}.png', bbox_inches='tight',dpi=600)
- # plt.show()
- plt.close()
- else:
- queryClusts = []
- color_count = 0
- plt.figure(figsize=(8, 4))
- for i in inputts:
- if color_count == 10:
- color_count -= 10
- if int(clusters[genes.index(i)]) not in queryClusts:
- queryClusts.append(int(clusters[genes.index(i)]))
- print(f'Cluster {clusters[genes.index(i)]} | {subfamily[int(clusters[genes.index(i)])]} | {members[int(clusters[genes.index(i)])]}')
- queryAccessions = []
- for a,b in zip(clusters,accessions):
- if int(a) == int(clusters[genes.index(i)]):
- queryAccessions.append(b)
- ### Adds all clusters in the combined set where at least one human query gene appears
- assemblyEquivClusts = []
- for ca in clusterAccessions:
- equivClusts = set()
- for qa in queryAccessions:
- for c,a in enumerate(ca):
- if qa in a:
- equivClusts.add(c)
- assemblyEquivClusts.append(list(sorted(equivClusts)))
- ### Creates .tsv file showing queried human and other species clustered IgSF accessions with descriptions
- if arg7 == 'yes':
- newFile = open(f'{homeDir}/results/clusteredIgSF_Cluster-{int(clusters[genes.index(i)])}.tsv','w')
- newFile.write('Taxa\tAssembly\tGeneID\tSpecies\tIntendedGene\tAccessionDescription\n')
- c = 0
- for a1,a2,b in zip(clusterAccessions,otherClusterAccessions,assemblyEquivClusts):
- for bb in b:
- holdClusters = []
- s = ''
- for x in a1[bb]:
- if clusters[accessions.index(x)] not in holdClusters:
- holdClusters.append(clusters[accessions.index(x)])
- s += f'Cluster {str(clusters[accessions.index(x)])} | {members[clusters2.index(clusters[accessions.index(x)])]}, '
- newFile.write(f'{s[0:-2]}\n')
- for x in a1[bb]:
- newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tHuman\t{accessionDescriptionDict[x]}\n')
- newFile.write('-\t-\t-\t-\t-\n')
- for x in a2[bb]:
- newFile.write(f'{taxDict[newAssemblies[c]]}\t{newAssemblies[c]}\t{accessionToGeneDict[x]}\tOther\t{accessionDescriptionDict[x]}\n')
- newFile.write('-\t-\t-\t-\t-\n')
- newFile.write(f'-\t-\t-\t-\t-\n')
- newFile.write(f'-\t-\t-\t-\t-\n')
- c += 1
- newFile.close()
- taxaSizeDiff = []
- for t in taxa:
- taxaSizeDiff.append([])
- c = 0
- for aec,sd in zip(assemblyEquivClusts,sizeDiff):
- hold = 0
- for ii in aec:
- hold += sd[ii]
- taxaSizeDiff[taxa.index(taxDict[newAssemblies[c]])].append(int(hold))
- c += 1
- r = []
- rstd = []
- for tsd in taxaSizeDiff:
- r.append(statistics.mean(tsd))
- if len(tsd) < 2:
- rstd.append(0)
- else:
- rstd.append(statistics.stdev(tsd))
- alpha = 0.8
- cap_size = 0.3
- plt.plot(range(len(taxa)),r,c=colors[color_count],label=subfamily[int(clusters[genes.index(i)])])
- for xi, yi, err in zip(range(0,len(taxa)),r,rstd):
- plt.vlines(xi, yi - err, yi + err,colors=colors[color_count],linestyles='dashed',alpha=alpha,linewidth=1)
- plt.hlines([yi - err, yi + err],xi - cap_size, xi + cap_size,colors=colors[color_count],alpha=alpha,linewidth=1)
- color_count += 1
- labelLines(plt.gca().get_lines(), align=True, fontsize=10)
- plt.axhline(y=0, color='k', linestyle='--')
- plt.xticks(range(0,len(taxa)),taxa,rotation=90, fontsize=10)
- plt.ylabel('Δ (Human – Avg. Taxa) Cluster Size',fontsize=12)
- plt.savefig(f'{homeDir}/results/IgSFSize_{inputTag}.png', bbox_inches='tight',dpi=600)
- # plt.show()
conservation-IgSFSize.py at commit af03288, no license · at the source
Overview
- Department of Systems & Computational Biology, Albert Einstein College of Medicine, Bronx, NY 10461, USA
- Department of Biochemistry, Albert Einstein College of Medicine, Bronx, NY 10461, USA
Abstract
The human immunoglobulin superfamily (IgSF) encompasses hundreds of proteins involved in cell–cell adhesion, neural connectivity, junctional organization, and immune regulation, with many serving as key checkpoint proteins. To elucidate the evolutionary history of human IgSF members, we systematically analyzed all available eukaryotic reference genomes to determine when each IgSF subfamily first appeared. The human IgSF was partitioned into six major evolutionary timeframes: Metazoa, Vertebrata, Gnathostomata, Tetrapoda, Amniota, and Mammalia. Although proteins were grouped solely by their conservation across eukaryotes, their biological functions clustered naturally, reflecting how new physiological systems create selective pressures that drive the appearance, retention, and diversification of protein architectures suited to those functions. Conservation and functional analyses indicate that human IgSF members arising in tetrapods and amniotes primary regulate the strength and duration of immune responses and form many of the critical components of the immune synapse, while IgSF genes that appear first in mammals have evolved to fine-tune immune activation thresholds to support maternal–fetal tolerance. Case studies are provided to illustrate three key evolutionary themes: (i) highly conserved yet catalytically inactive proteins retain essential regulatory functions, (ii) functional convergence among evolutionarily distinct IgSF families, and (iii) compensatory evolution within adaptive immunity following a lineage-specific loss of an IgSF member. Together, these findings establish an evolutionary framework for organizing the human IgSF by both ancestry and function, providing a foundation for assessing IgSF importance. Notably, this study facilitates the identification of conserved but understudied proteins that emerged alongside the development of the adaptive immune system, highlighting them as promising candidates for future experimental investigation.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.
StevenGrudman/IgSF-Conservation
af032888ebd9488e70430d119a5f9e2cb71bdd7c, 25 November 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
7 files
- conservation-IgSFSize.py
, Python, 532 lines, 1 match - otherScripts/
assignOrthologs1.py , Python, 447 lines - otherScripts/
blastIgSF.py , Python, 24 lines - otherScripts/
brotherhoodMethod.py , Python, 102 lines, 1 match - otherScripts/
submit_orth_human_redo.s , Shell, 14 linesh - otherScripts/
submit_singleJob.sh , Shell, 9 lines - README.md, Text, 27 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 6 scripts, each with its path and the digest of its content;
- 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability
The data underlying this article can be found in the online Supplementary material. In addition, all data and a program that can facilitate browsing results are publicly accessible at https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 4 keywords, 6 MeSH terms, 1 funder, 105 references.
Cite
This paper
Grudman, S., & Fiser, A. (2026). Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution. Genome biology and evolution, 18(6), evag133. https://
BibTeX
@article{grudman2026cons
author = {Grudman, Steven and Fiser, Andras},
title = {{Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution}},
journal = {Genome biology and evolution},
year = {2026},
month = jun,
volume = {18},
number = {6},
pages = {evag133},
publisher = {Oxford University Press},
issn = {1759-6653},
doi = {10.1093/
url = {https://
pmid = {42230317},
pmcid = {PMC13262535}
}
RIS
TY - JOUR
AU - Grudman, Steven
AU - Fiser, Andras
TI - Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution
T2 - Genome biology and evolution
J2 - Genome Biol Evol
PY - 2026
DA - 2026/
VL - 18
IS - 6
SP - evag133
SN - 1759-6653
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "Conservation of Human IgSF Proteins Throughout Eukaryotic Evolution",
"container-title": "Genome biology and evolution",
"author": [
{
"family": "Grudman",
"given": "Steven"
},
{
"family": "Fiser",
"given": "Andras"
}
],
"container-title-short":
"volume": "18",
"issue": "6",
"page": "evag133",
"DOI": "10.1093/
"PMID": "42230317",
"PMCID": "PMC13262535",
"ISSN": "1759-6653",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.7554/elife.107393 [code]
- Chromosome-scale genome assembly of the European common cuttlefish &
lt;i& gt;Sepia officinalis& lt;/ i& gt;. Journal: eLifeIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 3 references - [2] doi:10.1016/j.xgen.2026.101217 [code]
- ProtoCloud: A prototypical self-explaining model for single-cell analysis.Journal: Cell genomicsIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
- [3] doi:10.1016/j.molcel.2026.07.006 [code]
- DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.Journal: Molecular cellIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
- [4] doi:10.1038/s41593-026-02321-0 [code]
- Microbial reactivation of host androgens directs enteric neuronal regulation of gut motility.Journal: Nature neuroscienceIn common: pandas, SciPy, NumPy, cellular / molecular, 2 references
- [5] doi:10.1016/j.nbscr.2026.100149 [code]
- Molecular correlates of sleep deprivation in the mouse brain identified by meta-analysis of microarray data.Journal: Neurobiology of sleep and circadian rhythmsIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 2 references
- [6] doi:10.7554/elife.106134 [code]
- Identification and classification of ion channels across the tree of life provide functional insights into understudied CALHM channels.Journal: eLifeIn common: pandas, Matplotlib, cellular / molecular, 3 references
- [7] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 1 reference
- [8] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 1 reference
- [9] doi:10.1002/advs.202523984 [code]
- INB&
lt;sup& gt;3& lt;/ sup& gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery. Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: pandas, SciPy, Matplotlib, 1 other tool, 1 reference - [10] doi:10.1523/eneuro.0362-25.2026 [code]
- Similarities between &
lt;i& gt;Ciona& lt;/ i& gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/ V Subdivision within a Hindbrain Precursor? Journal: eNeuroIn common: pandas, SciPy, Matplotlib, 1 other tool, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 6 scripts, and 2 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e0eb357bda6d9cdf…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
