Preprocessing on the Go: Practices in Gait-Related Mobile EEG.
The 9 matches
- [1] § Methods › Literature Search ↔ utils/article_fetcher.py, lines 6–39 · score 0.72 · Medical Subject Headings, MeSH, retrieved, query, articles, keywords
- [2] § Results › Overview of Preprocessing Steps ↔ scripts/fig3_stepsnetwork.py, lines 34–43 · score 0.71 · IC decomposition, IC rejection, post ICA, pass filtering, artifact rejection, ERS
- [3] § Results › Overview of Preprocessing Steps ↔ scripts/figB_outcome_stepsnetwork.py, lines 53–68 · score 0.71 · IC decomposition, IC rejection, post ICA, pass filtering, artifact rejection, ERS
- [4] § Methods › Defining Preprocessing Steps ↔ dataframe_plots.ipynb, lines 860–895 · score 0.67 · highpass_filter, Post ICA, Pre ICA Signal, Raw, keywords, Preprocessing
- [5] § Results › Overview of Preprocessing Steps ↔ scripts/fig3_stepsnetwork.py, lines 34–43 · score 0.63 · notch filter, IC rejection, low pass filtering, Artifact rejection, ICA, preprocessing
- [6] § Results › Overview of Preprocessing Steps ↔ dataframe_plots.ipynb, lines 469–486 · score 0.63 · notch filter, IC rejection, low pass filtering, Artifact rejection, ICA, preprocessing
- [7] § Results › Study Cohorts and Gait Tasks ↔ dataframe_plots.ipynb, lines 324–382 · score 0.63 · overground walking, Treadmill walking, healthy adults, patients, paradigm, cohorts
- [8] § Results › Study Cohorts and Gait Tasks ↔ scripts/fig1_cohort_task.py, lines 13–20 · score 0.57 · PD clinical cohorts, PwPD
- [9] § Results › Study Cohorts and Gait Tasks ↔ scripts/fig1_cohort_task.py, lines 13–20 · score 0.53 · PD clinical cohorts, PwPD, HA
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 974 lines · 34 KB · no license · 3 matches
- # %% [markdown]
- # # Review of gait-related mobile EEG preprocessing steps — Data Preparation & Visualization
- #
- # This notebook organizes and visualizes data extracted from the literature via **Elicit**, focusing on experimental and preprocessing parameters relevant to mobile EEG during gait.
- #
- # Our goals here are to:
- # - Structure all study-level data into a clean DataFrame
- # - Explore distributions of cohorts, gait measurement systems, and artifact rejection methods
- # - Generate simple, clear, and reproducible plots for the manuscript and supplementary materials
- # %% [markdown]
- # ## Setup
- #
- # We'll start by importing necessary packages and defining paths.
- # The notebook assumes the CSV file exported from Elicit is available in the `../data/` directory.
- # %%
- import pandas as pd
- import numpy as np
- import ast
- import os
- import itertools
- from pathlib import Path
- from collections import Counter, defaultdict
- import matplotlib.pyplot as plt
- import seaborn as sns
- import networkx as nx
- from matplotlib.patches import Patch, FancyArrowPatch
- from matplotlib.lines import Line2D
- from math import sqrt
- from upsetplot import UpSet, from_indicators
- from utils.config import define_dir, dir_processed
- # %% [markdown]
- # ## Configure the directories
- # %%
- # Get the current directory where the script is executed
- dir_proj = Path.cwd()
- # Define the paths for 'logs' and 'results' directories
- dir_log_results = define_dir(dir_proj, "logs") # Logs directory path
- dir_results = define_dir(dir_proj, "results") # Results directory path
- dir_fulltexts = define_dir(dir_results, "fulltexts") # Full-text articles directory path
- dir_researcharticles = define_dir(dir_results, "researcharticles") # Research articles directory path
- dir_methods = define_dir(dir_results, "methods") # Methods sections directory path
- dir_data = define_dir(dir_proj, "data") # Data directory path
- dir_processed = define_dir(dir_results, "cleanresults") # Processed data directory
- # %% [markdown]
- # ## 1. Load the dataset
- #
- # We’ll start by loading the extracted dataset (`Elicitrevised.csv`), which contains:
- # - Metadata (Title, Citation, Filename)
- # - Experimental details (Cohort, Gait Task, EEG Type)
- # - Preprocessing details and artifact rejection methods
- #
- # Since some fields include multiple values (e.g., multiple gait systems or methods),
- # we’ll clean those for easy downstream use.
- # %%
- data_path = Path("data/20251003_Elicitrevised.csv")
- df = pd.read_csv(data_path, sep=";")
- print(f"Loaded {len(df)} studies")
- df.head(3)
- # %% [markdown]
- # ## Step 2: Organize Data into Structured DataFrames
- #
- # In this section, we will:
- # - Clean and standardize text fields.
- # - Separate parameters with multiple values (e.g., gait systems, artifact rejection methods).
- # - Create easy-to-analyze DataFrames for:
- # - Cohort
- # - Gait Task
- # - EEG Electrode Type
- # - Gait Measurement System
- # - Artifact Rejection Methods
- # - Preprocessing Steps
- # - Outcomes
- #
- # Each DataFrame will have one row per study per method/system, in order to plot them.
- # Some columns (e.g., `Gait_measurement_system`) contain multiple entries separated by `;` or `,`.
- # We will:
- # - Split them into lists.
- # - Explode these lists into separate rows.
- # This will make it possible to count and visualize occurrences of each system or method across studies.
- # %%
- # Standardize column names
- df.columns = (
- df.columns.str.strip()
- .str.replace(" ", "_")
- .str.replace("-", "_")
- )
- # Clean text fields
- text_cols = [
- "Cohort", "Gait_Task", "Dual_layer_cap",
- "Type_of_EEG_electrodes", "Gait_measurement_system",
- "Artifactrej_methods", "step_keywords", "outcome_keywords_script"
- ]
- for col in text_cols:
- if col in df.columns:
- df[col] = (
- df[col]
- .astype(str)
- .str.replace(r"[\r\n]+", " ", regex=True)
- .str.strip()
- )
- # --- Convert invalid values to NaN ---
- df.replace(["", "nan", "None", "NaN"], np.nan, inplace=True)
- # --- Drop duplicates and empty rows ---
- df = df.drop_duplicates().dropna(how="all")
- # --- Utility to clean and explode multi-value columns ---
- def split_and_clean(df, col):
- """Split semicolon/list-like column into multiple rows."""
- if col not in df.columns:
- print(f"Column '{col}' not found.")
- return pd.DataFrame(columns=["Title", "Citation", col])
- temp = df[["Title", "Citation", col]].copy()
- temp[col] = temp[col].astype(str).str.replace(r"[\[\]']", "", regex=True)
- temp[col] = temp[col].str.split(";|,")
- temp = temp.explode(col)
- temp[col] = temp[col].str.strip()
- temp = temp[temp[col].notna() & (temp[col] != "")]
- return temp
- print("Data cleaned and normalized.")
- df.head(3)
- # %%
- # --- Split all relevant columns ---
- df_cohort = split_and_clean(df, "Cohort")
- df_gait_task = split_and_clean(df, "Gait_Task")
- df_gait_system = split_and_clean(df, "Gait_measurement_system")
- df_artifact = split_and_clean(df, "Artifactrej_methods")
- df_eeg_electrodes = split_and_clean(df, "Type_of_EEG_electrodes")
- df_step_keywords = split_and_clean(df, "step_keywords")
- df_outcomes = split_and_clean(df, "outcome_keywords_script")
- # --- Summary ---
- print("Summary of entries:")
- print(f"Cohort entries: {len(df_cohort)}")
- print(f"Gait task entries: {len(df_gait_task)}")
- print(f"Gait systems entries: {len(df_gait_system)}")
- print(f"Artifact rejection entries: {len(df_artifact)}")
- print(f"EEG electrode entries: {len(df_eeg_electrodes)}")
- print(f"Step keyword entries: {len(df_step_keywords)}")
- print(f"Outcome keyword entries: {len(df_outcomes)}")
- # %%
- export_tables = {
- "Cohort": df_cohort,
- "Gait_Task": df_gait_task,
- "Gait_System": df_gait_system,
- "Artifact_Methods": df_artifact,
- "EEG_Electrodes": df_eeg_electrodes,
- "Step_Keywords": df_step_keywords,
- "Outcome_Keywords": df_outcomes
- }
- for name, table in export_tables.items():
- out_path = os.path.join(dir_processed, f"{name}_cleaned.csv")
- table.to_csv(out_path, index=False)
- print(f"Saved: {out_path}")
- # %%
- files = {
- "Cohort": "results/cleanresults/Cohort_cleaned.csv",
- "Gait_Task": "results/cleanresults/Gait_Task_cleaned.csv",
- "Gait_System": "results/cleanresults/Gait_System_cleaned.csv",
- "EEG_Electrodes": "results/cleanresults/EEG_Electrodes_cleaned.csv",
- "Step_Keywords": "results/cleanresults/Step_Keywords_cleaned.csv",
- "Artifact_Methods": "results/cleanresults/Artifact_Methods_cleaned.csv",
- "Outcome_Keywords": "results/cleanresults/Outcome_Keywords_cleaned.csv"
- }
- dfs = {name: pd.read_csv(path, dtype=str).fillna("") for name, path in files.items()}
- # Check sizes
- for name, df in dfs.items():
- print(f"{name}: {df.shape[0]} rows × {df.shape[1]} cols")
- # %%
- display(df_cohort.head(3))
- display(df_artifact.head(3))
- display(df_step_keywords.head(3))
- # %%
- # Check for missing values
- master_titles = set(dfs["Cohort"]["Title"])
- for name, df in dfs.items():
- missing = master_titles - set(df["Title"])
- extra = set(df["Title"]) - master_titles
- print(f"{name}: missing {len(missing)}, extra {len(extra)}")
- # %%
- # Aggregate steps
- steps_grouped = (
- dfs["Step_Keywords"]
- .groupby(["Title", "Citation"])["step_keywords"]
- .apply(lambda x: list(set(x.dropna())))
- .reset_index()
- )
- # Aggregate outcomes
- outcomes_grouped = (
- dfs["Outcome_Keywords"]
- .groupby(["Title", "Citation"])["outcome_keywords_script"]
- .apply(lambda x: list(set(x.dropna())))
- .reset_index()
- )
- # Merge core datasets
- meta = dfs["Cohort"][["Title", "Citation", "Cohort"]].merge(
- dfs["Gait_Task"][["Title", "Gait_Task"]], on="Title", how="left"
- )
- meta = meta.merge(dfs["Gait_System"][["Title", "Gait_measurement_system"]], on="Title", how="left")
- meta = meta.merge(dfs["EEG_Electrodes"][["Title", "Type_of_EEG_electrodes"]], on="Title", how="left")
- meta = meta.merge(steps_grouped, on="Title", how="left")
- meta = meta.merge(outcomes_grouped, on="Title", how="left")
- print(f"✅ Combined dataset: {meta.shape[0]} studies")
- # %% [markdown]
- # ### Creating Clean DataFrames for Each Parameter
- #
- # Now we’ll create structured DataFrames for:
- # - `Cohort`
- # - `Gait_Task`
- # - `Gait_measurement_system` (exploded)
- # - `Type_of_EEG_electrodes` (exploded)
- # - `Artifactrej_methods` (exploded)
- #
- # These will be stored in a dictionary for easy iteration during plotting or export.
- # %%
- available_files = list(dir_processed.glob("*.csv"))
- print("Available processed CSV files:")
- for f in available_files:
- print("-", f.name)
- processed_data = {f.stem: pd.read_csv(f) for f in available_files}
- for name, df in processed_data.items():
- print(f"\n{name.upper()}: {df.shape[0]} rows × {df.shape[1]} cols")
- print(df.head(3))
- # %%
- # Rebuild pipelines per study
- grouped_steps = df_step_keywords.groupby("Title")["step_keywords"].apply(list).reset_index()
- grouped_outcomes = df_outcomes.groupby("Title")["outcome_keywords_script"].apply(list).reset_index()
- merged = pd.merge(grouped_steps, grouped_outcomes, on="Title", how="outer", suffixes=("_steps", "_outcomes"))
- merged["step_keywords"] = merged["step_keywords"].apply(lambda x: x if isinstance(x, list) else [])
- merged["outcome_keywords_script"] = merged["outcome_keywords_script"].apply(lambda x: x if isinstance(x, list) else [])
- print(f"Reconstructed {len(merged)} study pipelines.")
- display(merged.head(3))
- # %% [markdown]
- # ### Exporting Structured Data
- #
- # We’ll now save each structured DataFrame into the project’s `data/processed_data` directory.
- #
- # The base data directory (`dir_data`) is already defined in `utils.config`.
- # This ensures that file paths remain consistent across scripts, notebooks, and collaborators.
- #
- # Each structured dataset will be exported as a CSV file and versioned for reproducibility.
- #
- # %% [markdown]
- # ## Load Processed Data for Visualization
- #
- # Now that the preprocessing and organization are done,
- # we’ll load the processed CSV files (e.g., cohort, gait task, electrode type, etc.)
- # from the `results/cleanresults` folder for visualization.
- # %%
- # --- Data verification cell ---
- import os
- import pandas as pd
- processed_files = [
- "Cohort_cleaned.csv",
- "Gait_Task_cleaned.csv",
- "Gait_System_cleaned.csv",
- "EEG_Electrodes_cleaned.csv",
- "Step_Keywords_cleaned.csv",
- "Artifact_Methods_cleaned.csv",
- "Outcome_Keywords_cleaned.csv"
- ]
- print("✅ Checking processed data structure...\n")
- for f in processed_files:
- path = os.path.join(dir_processed, f)
- if os.path.exists(path):
- df_temp = pd.read_csv(path)
- print(f"{f}: {df_temp.shape[0]} rows × {df_temp.shape[1]} cols")
- print(" Columns:", list(df_temp.columns))
- print(" Sample:\n", df_temp.head(2), "\n")
- else:
- print(f"⚠️ Missing file: {f}")
- # %% [markdown]
- # ## Prepare Aggregated Data for Visualization
- #
- # Next, let’s prepare frequency tables and summary data.
- # For example, how many studies used each gait measurement system,
- # EEG electrode type, and artifact rejection method.
- # %% [markdown]
- # ## Visualize Key Parameters
- #
- # We’ll now create simple, publication-ready bar plots
- # to visualize how frequently each parameter occurs across studies.
- #
- # Libraries: `matplotlib` and `seaborn`.
- # %% [markdown]
- # ## Distribution of Studies by Cohort and Gait Task
- #
- # To understand the experimental landscape,
- # we’ll visualize how different **cohorts** (e.g., healthy adults, patients)
- # are distributed across various **gait tasks** (e.g., treadmill, overground, obstacle walking).
- #
- # This gives a quick view of which populations are most commonly studied
- # under which walking paradigms.
- # %%
- # Load cleaned data
- df_cohort = pd.read_csv(os.path.join(dir_processed, "Cohort_cleaned.csv"))
- df_gait_task = pd.read_csv(os.path.join(dir_processed, "Gait_Task_cleaned.csv"))
- # Merge on Citation to align cohorts and gait tasks per study
- df_cohort_task = pd.merge(df_cohort, df_gait_task, on="Citation", suffixes=("_Cohort", "_Task"))
- # Count occurrences
- pivot = df_cohort_task.pivot_table(index="Cohort", columns="Gait_Task", values="Citation", aggfunc="count", fill_value=0)
- # Plot
- plt.figure(figsize=(10, 6))
- ax = pivot.plot(kind="barh", stacked=True, colormap="tab20", edgecolor="none")
- plt.title("Cohort vs Gait Task Distribution")
- plt.xlabel("Cohort")
- plt.ylabel("Number of Studies")
- # --- Modify legend labels ---
- handles, labels = ax.get_legend_handles_labels()
- label_map = {
- "Overground walking": "Only Overground walking",
- "Treadmill walking": "Only Treadmill walking"
- }
- # Replace only matching labels
- new_labels = [label_map.get(lbl, lbl) for lbl in labels]
- plt.legend(title="Gait Task", bbox_to_anchor=(0.95, 0.95), loc='upper left',
- facecolor="white",
- framealpha=0.9,
- fontsize=9,
- title_fontsize=10)
- plt.tight_layout()
- plt.show()
- # %%
- plt.figure(figsize=(10, 6))
- ax = pivot.plot(kind="barh", stacked=True, colormap="tab20", edgecolor="none")
- plt.title("Cohort vs Gait Task Distribution")
- plt.xlabel("Number of Studies")
- plt.ylabel("Cohort")
- # --- Modify legend labels ---
- handles, labels = ax.get_legend_handles_labels()
- label_map = {
- "Overground walking": "Only Overground walking",
- "Treadmill walking": "Only Treadmill walking"
- }
- # Replace only matching labels
- new_labels = [label_map.get(lbl, lbl) for lbl in labels]
- # --- Place legend inside plot ---
- plt.legend(
- handles,
- new_labels,
- title="Gait Task",
- loc="upper left", # you can try 'upper right' or 'center left' depending on aesthetics
- bbox_to_anchor=(0.95, 0.95),
- frameon=True,
- facecolor="white",
- framealpha=0.9,
- fontsize=10,
- title_fontsize=11
- )
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # ## Heatmap: EEG Electrode Type vs Gait Measurement System
- #
- # We want to visualize how many studies used each combination of EEG electrode type and gait measurement system.
- # The heatmap shows counts for each combination.
- # %%
- plt.figure(figsize=(10, 6))
- heat_data = meta.groupby(["Type_of_EEG_electrodes", "Gait_measurement_system"]).size().unstack(fill_value=0)
- sns.heatmap(heat_data, annot=True, cmap="YlGnBu", fmt="d")
- plt.title("EEG Electrode Types vs Gait Measurement Systems")
- plt.ylabel("Type of EEG Electrodes")
- plt.xlabel("Gait Measurement System")
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # # EEG Preprocessing Network Plot
- # We will visualize the flow of preprocessing steps as a layered network.
- # Each node represents a step, colored by stage, and arrows show transitions between steps across studies.
- #
- # Lets start with importing all required libraries
- # %%
- # Load the two separate datasets
- steps_df = pd.read_csv(os.path.join(dir_processed, "Step_Keywords_cleaned.csv"))
- outcomes_df = pd.read_csv(os.path.join(dir_processed, "Outcome_Keywords_cleaned.csv"))
- # Group keywords by 'Citation' into a semicolon-separated string for each file
- steps_grouped = steps_df.groupby('Citation')['step_keywords'].apply(lambda x: ';'.join(x.dropna())).reset_index()
- outcomes_grouped = outcomes_df.groupby('Citation')['outcome_keywords_script'].apply(lambda x: ';'.join(x.dropna())).reset_index()
- # Rename the outcome column to match the original script's expectation
- outcomes_grouped.rename(columns={'outcome_keywords_script': 'outcome_keywords'}, inplace=True)
- # Merge the two dataframes on 'Citation'
- # An outer merge ensures that articles with only steps or only outcomes are included
- df = pd.merge(steps_grouped, outcomes_grouped, on='Citation', how='outer')
- # Fill any missing values (NaN) with empty strings, as in the original code
- df.fillna("", inplace=True)
- # --- End of Data Loading and Merging Block ---
- # Stage map
- stage_map = {
- "Raw data": ["Raw data"],
- "Pre ICA - Signal Cleaning": [
- "Channel removal",
- "High-pass filter", "Low-pass filter", "Bandpass filter", "Notch filter",
- "Downsample"
- ],
- "Pre ICA - Data Preprocessing": [
- "Artifact rejection", "Bad channel detection","Re-reference", "Epoching"
- ],
- "ICA": ["IC decomposition", "IC rejection"],
- "Post ICA": [
- "Clustering", "Baseline correction",
- "Dipole fitting", "Normalization" ,"Despiking"
- ],
- "Outcome": ["PSD", "ERD/ERS", "ERSP", "CMC"]
- }
- # Build article-wise step tracking (unchanged)
- article_steps = {}
- for idx, row in df.iterrows():
- paper_id = row["Citation"] if "Citation" in df.columns else idx
- steps = row["step_keywords"].split(";") if row["step_keywords"] else []
- outcomes = row["outcome_keywords"].split(";") if row["outcome_keywords"] else []
- # Filter out empty strings that might result from splitting
- steps = [s for s in steps if s]
- outcomes = [o for o in outcomes if o]
- full_sequence = steps + outcomes
- article_steps[paper_id] = full_sequence
- # Calculate transition counts (unchanged, but now works with the merged data)
- transition_counts = Counter()
- for _, row in df.iterrows():
- steps = row["step_keywords"].split(";") if row["step_keywords"] else []
- outcomes = row["outcome_keywords"].split(";") if row["outcome_keywords"] else []
- # Filter out empty strings
- steps = [s for s in steps if s]
- outcomes = [o for o in outcomes if o]
- for i in range(len(steps) - 1):
- transition_counts[(steps[i], steps[i + 1])] += 1
- if steps and outcomes:
- last_step = steps[-1]
- for outcome in outcomes:
- transition_counts[(last_step, outcome)] += 1
- # Node stage mapping (unchanged)
- node_stage = {"raw_data": "Raw data"}
- for stage, keys in stage_map.items():
- for key in keys:
- node_stage[key] = stage
- # Setting layout (unchanged)
- layer_order = [
- "Raw data",
- "Pre ICA - Signal Cleaning",
- "Pre ICA - Data Preprocessing",
- "ICA",
- "Post ICA",
- "Outcome"
- ]
- stage_y = {stage: -i for i, stage in enumerate(layer_order)}
- # Function to get node positions (unchanged)
- def get_node_positions(G, node_stage, stage_map):
- positions = {}
- x_coords = defaultdict(int)
- for node in G.nodes():
- stage = node_stage.get(node, "Raw data")
- y = stage_y.get(stage, -10)
- x = x_coords[y]
- positions[node] = (x, y)
- x_coords[y] += 1
- # Center the nodes in each layer
- for y_val in x_coords:
- num_nodes = x_coords[y_val]
- nodes_at_y = [node for node, pos in positions.items() if pos[1] == y_val]
- for i, node in enumerate(nodes_at_y):
- positions[node] = (i - (num_nodes - 1) / 2.0, y_val)
- return positions
- # Plotting function (unchanged)
- def plot_layered_flowchart(transition_counts, node_stage_map, stage_map, title="Flow of EEG Preprocessing Steps across multiple studies"):
- G = nx.DiGraph()
- for (src, dst), weight in transition_counts.items():
- G.add_edge(src, dst, weight=weight)
- # Add 'raw_data' if it's a source node but not in the graph yet
- if 'raw_data' not in G.nodes() and any(u == 'raw_data' for u, v in transition_counts.keys()):
- G.add_node('raw_data')
- # Color map by node stage
- color_map = {
- "Raw data": "#A9A9A9", # dark gray
- "Pre ICA - Signal Cleaning": "#FF8C42", # deep orange
- "Pre ICA - Data Preprocessing": "#20B2AA", # teal
- "ICA": "#9370DB", # medium purple
- "Post ICA": "#D9534F", # red / crimson
- "Outcome": "#3CB371", # medium sea green
- }
- node_colors = [color_map.get(node_stage_map.get(node, "Raw data"), "gray") for node in G.nodes()]
- node_sizes = [300 + 200 * G.degree(n) for n in G.nodes()]
- pos = get_node_positions(G, node_stage_map, stage_map)
- fig, ax = plt.subplots(figsize=(18, 12))
- nx.draw_networkx_nodes(G, pos, node_size=node_sizes, node_color=node_colors, ax=ax)
- nx.draw_networkx_labels(G, pos, font_size=9, ax=ax)
- # Edge color and width rules
- def get_edge_style(weight):
- if weight >= 46:
- return "black", 4 + weight * 0.15
- elif weight >= 30:
- return "#023e8a", 3.8 + weight * 0.13
- elif weight >= 15:
- return "brown", 3 + weight * 0.12
- elif weight >= 5:
- return "gray", 2 + weight * 0.1
- else:
- return "#cccccc", 1.0
- # Draw arrows with arrowheads using FancyArrowPatch
- for u, v in G.edges():
- weight = G[u][v]['weight']
- color, width = get_edge_style(weight)
- start = pos[u]
- end = pos[v]
- node_radius = 0.3 # Adjust this based on node sizes
- # Offset calculation
- dx, dy = end[0] - start[0], end[1] - start[1]
- dist = sqrt(dx**2 + dy**2)
- if dist == 0: continue
- arrow_start = (start[0] + dx * node_radius / dist, start[1] + dy * node_radius / dist)
- arrow_end = (end[0] - dx * node_radius / dist, end[1] - dy * node_radius / dist)
- arrow = FancyArrowPatch(
- posA=arrow_start,
- posB=arrow_end,
- connectionstyle="arc3,rad=0.2",
- arrowstyle="->",
- mutation_scale=20,
- color=color,
- linewidth=width,
- alpha=0.8
- )
- ax.add_patch(arrow)
- # Legend 1: Node colors (stage types)
- node_legend_elements = [
- Patch(facecolor=color, edgecolor="black", label=stage)
- for stage, color in color_map.items()
- ]
- # Legend 2: Arrow colors (article frequency)
- arrow_legend_elements = [
- Line2D([0], [0], color="#cccccc", lw=2, label="1-4 articles"),
- Line2D([0], [0], color="gray", lw=2, label="5–14 articles"),
- Line2D([0], [0], color="brown", lw=2, label="15–29 articles"),
- Line2D([0], [0], color="#023e8a", lw=2, label="30+ articles")
- ]
- # Place the legends
- first_legend = ax.legend(handles=node_legend_elements, title="Preprocessing Stages", loc="upper left", bbox_to_anchor=(1.02, 1), borderaxespad=0.)
- ax.add_artist(first_legend)
- ax.legend(handles=arrow_legend_elements, title="Transition Frequency", loc="upper left", bbox_to_anchor=(1.10, 0.7), borderaxespad=0.)
- plt.title(title, fontsize=16)
- plt.axis("off")
- plt.tight_layout(rect=[0.1, 0.25, 0.75, 1])
- plt.show()
- # Run plot
- plot_layered_flowchart(transition_counts, node_stage, stage_map)
- # %%
- import pandas as pd
- import networkx as nx
- import matplotlib.pyplot as plt
- from matplotlib.patches import FancyArrowPatch, Patch
- from matplotlib.lines import Line2D
- from collections import Counter, defaultdict
- from math import sqrt
- import os
- # --- Data Loading and Merging Block ---
- # Assumes the CSV files are in the same directory as the script.
- # If not, provide the full path to the files.
- try:
- steps_df = pd.read_csv(os.path.join(dir_processed, "Step_Keywords_cleaned.csv"))
- outcomes_df = pd.read_csv(os.path.join(dir_processed, "Outcome_Keywords_cleaned.csv"))
- except FileNotFoundError:
- print("Error: Make sure 'Step_Keywords_cleaned.csv' and 'Outcome_Keywords_cleaned.csv' are in the correct directory.")
- exit()
- # Group keywords by 'Citation' into a semicolon-separated string for each file
- steps_grouped = steps_df.groupby('Citation')['step_keywords'].apply(lambda x: ';'.join(x.dropna())).reset_index()
- outcomes_grouped = outcomes_df.groupby('Citation')['outcome_keywords_script'].apply(lambda x: ';'.join(x.dropna())).reset_index()
- # Rename the outcome column to match the original script's expectation
- outcomes_grouped.rename(columns={'outcome_keywords_script': 'outcome_keywords'}, inplace=True)
- # Merge the two dataframes on 'Citation'
- df = pd.merge(steps_grouped, outcomes_grouped, on='Citation', how='outer')
- df.fillna("", inplace=True)
- # --- End of Data Loading and Merging Block ---
- # Stage map
- stage_map = {
- "Raw data": ["raw_data"],
- "Pre ICA - Signal Cleaning": [
- "channel_removal", "highpass_filter", "lowpass_filter", "bandpass_filter",
- "notch_filter", "downsample"
- ],
- "Pre ICA - Data Preprocessing": [
- "artifact_rejection", "bad_channel_detection", "re_reference", "epoching"
- ],
- "ICA": ["IC_decomposition", "IC_rejection"],
- "Post ICA": [
- "clustering", "baseline_correction", "dipole_fitting", "normalization",
- "despiking", "CA_rejection"
- ],
- "Outcome": ["PSD", "ERD/ERS", "ERSP", "CMC"]
- }
- # --- Descriptive Statistics Calculation ---
- total_papers = len(df['Citation'].unique())
- step_counts = Counter()
- # Get all unique steps and outcomes from the stage map to ensure all are counted
- all_possible_steps = [step for sublist in stage_map.values() for step in sublist]
- # Calculate how many papers mention each step
- for step in all_possible_steps:
- for _, row in df.iterrows():
- full_keyword_list = (row['step_keywords'] + ';' + row['outcome_keywords']).split(';')
- if step in full_keyword_list:
- step_counts[step] += 1
- print("--- Descriptive Statistics of Preprocessing Steps ---")
- print(f"Total number of unique papers analyzed: {total_papers}\n")
- for step, count in sorted(step_counts.items(), key=lambda item: item[1], reverse=True):
- percentage = (count / total_papers) * 100
- print(f"{step}: {count} out of {total_papers} papers ({percentage:.2f}%)")
- print("-" * 50)
- # Calculate transition counts
- transition_counts = Counter()
- for _, row in df.iterrows():
- steps = [s for s in row["step_keywords"].split(";") if s]
- outcomes = [o for o in row["outcome_keywords"].split(";") if o]
- full_sequence = steps + outcomes
- for i in range(len(full_sequence) - 1):
- transition_counts[(full_sequence[i], full_sequence[i+1])] += 1
- print("\n--- Statistics of Step Transitions ---")
- for (src, dst), count in sorted(transition_counts.items(), key=lambda item: item[1], reverse=True):
- print(f"Transition from '{src}' to '{dst}': {count} times")
- print("-" * 50)
- # --- End of Statistics Calculation ---
- # Node stage mapping
- node_stage = {"raw_data": "Raw data"}
- for stage, keys in stage_map.items():
- for key in keys:
- node_stage[key] = stage
- # Setting layout
- layer_order = [
- "Raw data", "Pre ICA - Signal Cleaning", "Pre ICA - Data Preprocessing",
- "ICA", "Post ICA", "Outcome"
- ]
- stage_y = {stage: -i for i, stage in enumerate(layer_order)}
- def get_node_positions(G, node_stage_map):
- positions = {}
- nodes_by_stage = defaultdict(list)
- for node in G.nodes():
- stage = node_stage_map.get(node, "Raw data")
- nodes_by_stage[stage].append(node)
- for stage in layer_order:
- # Sort nodes alphabetically for consistent layout
- nodes = sorted(nodes_by_stage[stage])
- y = stage_y.get(stage)
- num_nodes = len(nodes)
- for i, node in enumerate(nodes):
- # Center the nodes horizontally
- x = i - (num_nodes - 1) / 2.0
- positions[node] = (x, y)
- return positions
- def plot_layered_flowchart(transition_counts, node_stage_map, title="Flow of EEG Preprocessing Steps across multiple studies"):
- G = nx.DiGraph()
- for (src, dst), weight in transition_counts.items():
- G.add_edge(src, dst, weight=weight)
- if 'raw_data' not in G.nodes() and any(u == 'raw_data' for u, v in transition_counts.keys()):
- G.add_node('raw_data')
- color_map = {
- "Raw data": "#A9A9A9",
- "Pre ICA - Signal Cleaning": "#FF8C42",
- "Pre ICA - Data Preprocessing": "#20B2AA",
- "ICA": "#9370DB",
- "Post ICA": "#D9534F",
- "Outcome": "#3CB371",
- }
- node_colors = [color_map.get(node_stage_map.get(node, "Raw data"), "gray") for node in G.nodes()]
- node_sizes = [400 + 250 * G.degree(n) for n in G.nodes()]
- pos = get_node_positions(G, node_stage_map)
- fig, ax = plt.subplots(figsize=(22, 16))
- nx.draw_networkx_nodes(G, pos, node_size=node_sizes, node_color=node_colors, ax=ax, edgecolors='black')
- # Increased font size for better readability
- nx.draw_networkx_labels(G, pos, font_size=11, font_weight='bold', ax=ax)
- def get_edge_style(weight):
- if weight >= 46: return "black", 4 + weight * 0.15
- elif weight >= 30: return "#023e8a", 3.8 + weight * 0.13
- elif weight >= 15: return "brown", 3 + weight * 0.12
- elif weight >= 5: return "gray", 2 + weight * 0.1
- else: return "#cccccc", 1.5
- for u, v in G.edges():
- weight = G[u][v]['weight']
- color, width = get_edge_style(weight)
- start, end = pos[u], pos[v]
- # A simple approximation for radius based on node_size to offset the arrow
- start_node_index = list(G.nodes()).index(u)
- end_node_index = list(G.nodes()).index(v)
- start_radius = (node_sizes[start_node_index] ** 0.5) / 50.0
- end_radius = (node_sizes[end_node_index] ** 0.5) / 50.0
- dx, dy = end[0] - start[0], end[1] - start[1]
- dist = sqrt(dx**2 + dy**2)
- if dist == 0: continue
- arrow_start = (start[0] + dx * start_radius / dist, start[1] + dy * start_radius / dist)
- arrow_end = (end[0] - dx * end_radius / dist, end[1] - dy * end_radius / dist)
- arrow = FancyArrowPatch(
- posA=arrow_start, posB=arrow_end, connectionstyle="arc3,rad=0.2",
- arrowstyle="-|>,head_length=0.8,head_width=0.4", mutation_scale=25,
- color=color, linewidth=width, alpha=0.9
- )
- ax.add_patch(arrow)
- # Legends
- node_legend_elements = [Patch(facecolor=color, edgecolor="black", label=stage) for stage, color in color_map.items()]
- arrow_legend_elements = [
- Line2D([0], [0], color="#cccccc", lw=2, label="1-4 articles"),
- Line2D([0], [0], color="gray", lw=2, label="5–14 articles"),
- Line2D([0], [0], color="brown", lw=2, label="15–29 articles"),
- Line2D([0], [0], color="#023e8a", lw=2, label="30+ articles")
- ]
- first_legend = ax.legend(handles=node_legend_elements, title="Preprocessing Stages", fontsize=12, title_fontsize=14, loc="upper left", bbox_to_anchor=(1.02, 1))
- ax.add_artist(first_legend)
- ax.legend(handles=arrow_legend_elements, title="Transition Frequency", fontsize=12, title_fontsize=14, loc="upper left", bbox_to_anchor=(1.02, 0.7))
- plt.title(title, fontsize=20, fontweight='bold')
- plt.axis("off")
- plt.tight_layout(rect=[0, 0, 0.85, 1]) # Adjust for legend
- plt.savefig("flowchart_with_stats.png", dpi=300, bbox_inches='tight') # Save the figure
- plt.show()
- # Run plot
- plot_layered_flowchart(transition_counts, node_stage)
- # %%
- # =========================================
- # 6️⃣ Define Stage Mapping and Dynamic Assignment
- # =========================================
- stage_map = {
- "Raw data": ["raw_data"],
- "Pre ICA - Signal Cleaning": [
- "channel_removal",
- "highpass_filter", "lowpass_filter", "bandpass_filter", "notch_filter",
- "downsample"
- ],
- "Pre ICA - Data Preprocessing": [
- "artifact_rejection", "bad_channel_detection","re_reference", "epoching"
- ],
- "ICA": ["IC_decomposition", "IC_rejection"],
- "Post ICA": [
- "clustering", "baseline_correction",
- "dipole_fitting", "normalization" ,"despiking", "CA_rejection"
- ],
- "Outcome": ["PSD", "ERD/ERS", "ERSP", "CMC"]
- }
- def assign_stage(step, idx, steps):
- """Assign preprocessing stage based on position and context."""
- if step in ["bandpass_filter", "lowpass_filter", "highpass_filter", "notch_filter", "artifact_rejection"]:
- if "IC_decomposition" in steps[idx + 1:]:
- if "filter" in step:
- return "Pre ICA - Signal Cleaning"
- else:
- return "Pre ICA - Data Preprocessing"
- else:
- return "Post ICA"
- for stage, keywords in stage_map.items():
- if step in keywords:
- return stage
- return "Other"
- # %%
- steps_all = sorted(set([s for sublist in steps_grouped["step_keywords"] for s in sublist]))
- pivot = pd.DataFrame(False, index=steps_grouped["Citation"], columns=steps_all)
- for _, row in steps_grouped.iterrows():
- for step in row["step_keywords"]:
- pivot.loc[row["Title"], step] = True
- upset_data = from_indicators(pivot.columns, data=pivot)
- plt.figure(figsize=(10, 6))
- UpSet(upset_data, show_counts=True, sort_by='degree').plot()
- plt.suptitle("Overlap of EEG Preprocessing Steps Across Studies", fontsize=14)
- plt.show()
- # %%
- df_artifact = pd.read_csv(os.path.join(dir_processed, "Artifact_Methods_cleaned.csv"))
- df_artifact["Citation"] = df_artifact["Citation"].fillna("Unknown Study")
- methods = df_artifact["Artifactrej_methods"].unique()
- pivot = pd.DataFrame(0, index=df_artifact["Citation"].unique(), columns=methods)
- for title, group in df_artifact.groupby("Citation"):
- for m in group["Artifactrej_methods"]:
- pivot.loc[title, m] = 1
- # Horizontal stacked bars
- plt.figure(figsize=(10, len(pivot) * 0.25 + 3))
- bottoms = np.zeros(len(pivot))
- colors = sns.color_palette("tab20", n_colors=len(methods))
- for i, method in enumerate(pivot.columns):
- plt.barh(pivot.index, pivot[method], left=bottoms, color=colors[i], label=method)
- bottoms += pivot[method].values
- plt.xlabel("Method Presence (1 = used)")
- plt.ylabel("Study")
- plt.title("Artifact Rejection Methods Across Studies")
- plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left', title="Methods")
- plt.tight_layout()
- plt.show()
- # %%
- df_artifact = pd.read_csv(os.path.join(dir_processed, "Artifact_Methods_cleaned.csv"))
- df_artifact["Citation"] = df_artifact["Citation"].fillna("Unknown Study")
- methods = df_artifact["Artifactrej_methods"].unique()
- pivot = pd.DataFrame(0, index=df_artifact["Citation"].unique(), columns=methods)
- for title, group in df_artifact.groupby("Citation"):
- for m in group["Artifactrej_methods"]:
- pivot.loc[title, m] = 1
- # Horizontal stacked bars
- plt.figure(figsize=(10, len(pivot) * 0.25 + 3))
- bottoms = np.zeros(len(pivot))
- colors = sns.color_palette("tab20", n_colors=len(methods))
- for i, method in enumerate(pivot.columns):
- plt.barh(pivot.index, pivot[method], left=bottoms, color=colors[i], label=method)
- bottoms += pivot[method].values
- plt.xlabel("Method Presence (1 = used)")
- plt.ylabel("Study")
- plt.title("Artifact Rejection Methods Across Studies")
- plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left', title="Methods")
- plt.tight_layout()
- plt.show()
- # %%
- meta.to_csv("Combined_Metadata.csv", index=False)
- print("Combined dataset saved as Combined_Metadata.csv")
dataframe_plots.ipynb at commit 7508375, no license · at the source
Overview
- Department of Neurology, University Hospital Schleswig‐Holstein Campus Kiel and Kiel University, Kiel, Germany
- Neuropsychology Lab, Department of Psychology, Carl von Ossietzky University of Oldenburg, Oldenburg, Germany
Abstract
Mobile EEG has become popular in investigating brain dynamics during gait in recent years. Within this development, new preprocessing pipelines have been introduced and refined. The diversity of approaches, however, complicates comparisons across studies. To provide clarity, we reviewed studies that combined mobile EEG with gait measurements to map the preprocessing pipelines used in the field. Our review identified substantial heterogeneity in pipeline steps, their order, combinations, and the level of reporting detail. We visualized this heterogeneity as a map, tracing pathways from raw data to outcomes such as Power spectral density (PSD), Event‐related spectral perturbations (ERSP), Event‐related (de‐) synchronization (ERD/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.
vaishalivinod/LitExtract
7508375060633e3369520e33f2affb2b6e6411e0, 10 April 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
18 files
- dataframe_plots.ipynb, Jupyter, 974 lines
- scripts/
clean_elicitdatacsv.py , Python, 69 lines - scripts/
extractmethods.py , Python, 4 lines - scripts/
fig1_cohort_task.py , Python, 93 lines - scripts/
fig2_eegelec_gait.py , Python, 93 lines - scripts/
fig3_stepsnetwork.py , Python, 285 lines - scripts/
fig4_artifactrej.py , Python, 106 lines - scripts/
figB_outcome_stepsnetwor , Python, 421 linesk.py - scripts/
filter_researcharticles. , Python, 6 linespy - scripts/
retrieve_articles.py , Python, 29 lines - utils/
__init__.py , Python, 1 line - utils/
article_fetcher.py , Python, 229 lines - utils/
cleandata.py , Python, 50 lines - utils/
config.py , Python, 23 lines - utils/
log_search.py , Python, 18 lines - utils/
methodstext.py , Python, 87 lines - utils/
saveas.py , Python, 49 lines - README.md, Text, 49 lines
neurogeriatricskiel/LitExtract
7508375060633e3369520e33f2affb2b6e6411e0, 10 April 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
18 files
- dataframe_plots.ipynb, Jupyter, 974 lines, 3 matches
- scripts/
clean_elicitdatacsv.py , Python, 69 lines - scripts/
extractmethods.py , Python, 4 lines - scripts/
fig1_cohort_task.py , Python, 93 lines, 2 matches - scripts/
fig2_eegelec_gait.py , Python, 93 lines - scripts/
fig3_stepsnetwork.py , Python, 285 lines, 2 matches - scripts/
fig4_artifactrej.py , Python, 106 lines - scripts/
figB_outcome_stepsnetwor , Python, 421 lines, 1 matchk.py - scripts/
filter_researcharticles. , Python, 6 linespy - scripts/
retrieve_articles.py , Python, 29 lines - utils/
__init__.py , Python, 1 line - utils/
article_fetcher.py , Python, 229 lines, 1 match - utils/
cleandata.py , Python, 50 lines - utils/
config.py , Python, 23 lines - utils/
log_search.py , Python, 18 lines - utils/
methodstext.py , Python, 87 lines - utils/
saveas.py , Python, 49 lines - README.md, Text, 49 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 34 scripts, each with its path and the digest of its content;
- 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
The data that support the findings of this study are openly available in LitExtract at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Wiley
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 5 keywords, 5 MeSH terms, 1 funder, 74 references.
Cite
This paper
Vinod, V., Papin, L. J., Romijnders, R., Maetzler, W., & Welzel, J. (2026). Preprocessing on the Go: Practices in Gait-Related Mobile EEG. Psychophysiology, 63(6), e70352. https://
BibTeX
@article{vinod2026prepro
author = {Vinod, Vaishali and Papin, Lara Johanna and Romijnders, Robbin and Maetzler, Walter and Welzel, Julius},
title = {{Preprocessing on the Go: Practices in Gait-Related Mobile EEG}},
journal = {Psychophysiology},
year = {2026},
month = jun,
volume = {63},
number = {6},
pages = {e70352},
publisher = {Wiley},
issn = {0048-5772},
doi = {10.1111/
url = {https://
pmid = {42347764},
pmcid = {PMC13296838}
}
RIS
TY - JOUR
AU - Vinod, Vaishali
AU - Papin, Lara Johanna
AU - Romijnders, Robbin
AU - Maetzler, Walter
AU - Welzel, Julius
TI - Preprocessing on the Go: Practices in Gait-Related Mobile EEG
T2 - Psychophysiology
J2 - Psychophysiology
PY - 2026
DA - 2026/
VL - 63
IS - 6
SP - e70352
SN - 0048-5772
PB - Wiley
DO - 10.1111/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1111/
"type": "article-journal",
"title": "Preprocessing on the Go: Practices in Gait-Related Mobile EEG",
"container-title": "Psychophysiology",
"author": [
{
"family": "Vinod",
"given": "Vaishali"
},
{
"family": "Papin",
"given": "Lara Johanna"
},
{
"family": "Romijnders",
"given": "Robbin"
},
{
"family": "Maetzler",
"given": "Walter"
},
{
"family": "Welzel",
"given": "Julius"
}
],
"container-title-short":
"volume": "63",
"issue": "6",
"page": "e70352",
"DOI": "10.1111/
"PMID": "42347764",
"PMCID": "PMC13296838",
"ISSN": "0048-5772",
"publisher": "Wiley",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1111/ejn.70428
- Modulation of Corticomuscular Coupling With Split-Belt Locomotor Adaptation in Healthy Young Males.Journal: The European journal of neuroscienceIn common: EEG, 12 references
- [2] doi:10.1007/s00221-026-07342-6 [code]
- A portable solution for simultaneous human movement and mobile EEG acquisition: readiness potential for basketball free-throw shooting.Journal: Experimental brain researchIn common: EEG, 10 references
- [3] doi:10.1371/journal.pone.0348957 [code]
- Steps against the burden of Parkinson's disease (StepuP): Protocol of a randomized controlled trial elucidating the biomechanical and neurophysiological mechanisms of a speed dependent treadmill training intervention.Journal: PloS oneIn common: Matplotlib, NumPy, EEG, 5 references, author Vaishali Vinod
- [4] doi:10.3390/s26082440 [code]
- Suppressing Non-Stationary Motion Artefacts in Mobile EEG Using Generalized Eigenvalue Decomposition.Journal: Sensors (Basel, Switzerland)In common: methods / tools, EEG, 8 references
- [5] doi:10.1016/j.isci.2026.116647 [code]
- Enhancing mobile brain and body imaging: Open-source solutions for real-world research applications.Journal: iScienceIn common: methods / tools, 6 references
- [6] doi:10.1111/psyp.70265 [code]
- Neurocognitive Dynamics of Translating Information From a Spatial Map Into Action.Journal: PsychophysiologyIn common: EEG, 5 references
- [7] doi:10.1111/psyp.70297 [code]
- Heartbeat-Evoked Responses in M/
EEG: A Systematic Review of Methods With Suggestions for Analysis and Reporting. Journal: PsychophysiologyIn common: pandas, NumPy, methods / tools, EEG, 3 references - [8] doi:10.1002/hbm.70469 [code]
- VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.Journal: Human brain mappingIn common: NetworkX, seaborn, pandas, 2 other tools, methods / tools, 1 reference
- [9] doi:10.1002/mds.70348 [code]
- Electroencephalography-B
ased Clustering Reveals Robust Neurophysiological Subtypes in Parkinson's Disease. Journal: Movement disorders : official journal of the Movement Disorder SocietyIn common: seaborn, pandas, Matplotlib, 1 other tool, EEG, 2 references - [10] doi:10.1038/s41598-026-56070-y [code]
- SSDLabeler: realistic semi-synthetic data generation for multi-label artifact classification in EEG.Journal: Scientific reportsIn common: seaborn, pandas, Matplotlib, 1 other tool, EEG, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 34 scripts, and 9 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:2e0ce0c843250ec3…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
