Labs
OSCR is young, and much of it is an experiment in how research code can be found, read and cited. This page gathers what is experimental, how to build on the registry's data, and what is being built — marked as such.
Experiments you can use today
- The Code ↔ Paper matches
- Method
lexical-v1: paragraphs and code units reduced to technical terms, a pair kept when several rare terms agree. A first method, measured by hand on 25 papers; its next steps are a better reading of the paper's structure, a real parser for the code, and a local model. The reader shows every match with its evidence and score. - The paper from two sources
- When Europe PMC does not answer, the reader reads PubMed Central's copy, whose paragraphs may be numbered otherwise, and places each match by its section and evidence terms.
- The classification
- Rules over a paper's title, keywords, MeSH terms, subjects, abstract and journal, with their reasons; a local model will settle the ambiguous cases once compared with hand labels (how).
- A registry within a fixed number of files
- Whatever the catalogue's size, the site keeps a bounded number of files: authors, journals, institutions, tools and datasets are rendered in the browser from a fixed number of shards, and older papers by the Worker from their records.
The open datasets
OpenScientificCodeRegistry/Database, on Hugging Face: the authors' scripts under verified licences, deduplicated by the SHA-256 of their text, as Parquet blocks (blocks/NNNNN.parquet: sha256, language, size, lines, content; 64 rows a row group; zstd) with one JSON manifest per repository (manifests/<xx>/<repository>.json). Accessing the data gives the layout in full, and how to read it with Python or DuckDB.
Reading one script with range requests
The blocks are cut for the browser: a reader fetches a block's footer once, then only the row group that holds the file it wants — about 78 KB for one script — by HTTP range requests, which Hugging Face answers with the headers a page on another site needs. With hyparquet and its zstd decompressor:
// A script's text from the open dataset, in a browser: the block's footer once, then the one
// row group (64 rows) that holds the file, by HTTP range requests. npm: hyparquet, hyparquet-compressors.
import { asyncBufferFromUrl, parquetReadObjects } from "hyparquet";
import { compressors } from "hyparquet-compressors";
const base = "https://huggingface.co/datasets/OpenScientificCodeRegistry/Database/resolve/main";
// 1. The repository's manifest: its commit, its licence, and where each file's text is.
const repo = "github.com/owner/name";
const shard = [...new Uint8Array(await crypto.subtle.digest("SHA-1", new TextEncoder().encode(repo)))]
.map((b) => b.toString(16).padStart(2, "0")).join("").slice(0, 2);
const manifest = await (await fetch(`${base}/manifests/${shard}/${repo.replace(/[^A-Za-z0-9._-]+/g, "__")}.json`)).json();
const file = manifest.files.find((f) => f.path === "analysis/run.py");
// 2. That row of that block.
const block = await asyncBufferFromUrl({ url: `${base}/blocks/${String(file.block).padStart(5, "0")}.parquet` });
const [row] = await parquetReadObjects({ file: block, compressors, rowStart: file.row, rowEnd: file.row + 1 });
console.log(manifest.license, row.sha256 === file.sha256, row.content);Check the file's sha256 against the manifest's, and keep the licence the manifest names. A block, once published, never changes: its rows can be cached for good.
The tracing maps' format
A map is one JSON file, tracing-map.json, format tracing-map/0.1, deposited on Zenodo under CC0-1.0 once an author validated it (what a map is). Its shape, shortened:
{
"format": "tracing-map/0.1",
"paper": { "doi": "10.1234/abcd", "title": "…", "journal": "…", "published": "2026-09-21",
"authors": ["…"] },
"code": [{
"repo": "github.com/owner/name", "url": "https://github.com/owner/name",
"state": "alive", "license": "MIT", "commit": "7a034b0a…", "commit_date": "2026-09-01T10:00:00+02:00",
"type": "", "software_heritage_archived": true, "level": "inventoried",
"found_by": "text:availability", "section": "Code availability",
"files": [{ "path": "analysis/run.py", "language": "Python", "digest": "c26da9b3…" }]
}],
"alignments": [{
"paragraph": 42, "section": "Methods › EEG preprocessing", "repo": "github.com/owner/name",
"path": "analysis/run.py", "start_line": 10, "end_line": 38, "symbol": "preprocess",
"score": 3.1, "evidence": ["notch 50", "ica", "l_freq"], "method": "lexical-v1"
}],
"proposed": { "by": "oscr", "on": "2026-09-29" },
"validated": { "by": "Family, Given", "orcid": "0000-0000-0000-0000", "on": "2026-10-02", "proof": "orcid" }
}code[].levelfound(the paper cites it),alive(it answers),inventoried(files listed, commit recorded).code[].files[].digest- the SHA-256 of the file's text at the recorded commit.
alignments[].paragraph- the paragraph's position among the
<p>of the<body>of the paper's open-access XML at Europe PMC, counted from 0 in document order;evidenceholds short terms only, never a sentence.
validated appears once an author validated the map. The digest a validation carries is the SHA-256 of the map without proposed and validated, serialized as compact JSON with sorted keys.
The site's own files
Every record the pages show is a public, static JSON file, rebuilt every night. They serve the pages first; their layout may change.
/lookup/NN.json- The DOI lookup: 256 shards;
NNis the first 2 hex characters of the SHA-1 of the lower-cased DOI. Each maps a DOI to[status, day read], and the page's identifier when it has one. /records/<type>/NN.json- The authors (1,024 shards), institutions (512), tools (256), journals (128) and datasets (128): each shard holds
entitiesby key and therowsof their papers. A key's shard is the leading bits of the SHA-1 of the key (an ORCID iD in capitals; any other key in lower case), as many as the number of shards needs, in hexadecimal. /records/paper/NN.json- The records of the papers whose pages are rendered on demand (256 shards, by the SHA-1 of the page's identifier).
/badge.svg,/sitemap.xml- The badge; the sitemap and its shards, listing every page.
The search's API
The search page asks GET /api/search, which anyone may call — each call counts against the site's daily free quota, so call it sparingly:
https://openscicode.org/api/search?q=eeg&modality=eeg&sort=newest&page=1&size=20 https://openscicode.org/api/search?q=tool:mne&format=csv (the first 500 results, CSV or JSON)
Parameters: q (the query language), the filters (status, year, modality, organism, population, subfield, tool, language, journal, data, host, code_license, type, license, matches, oa, repeated for OR), from, to, sort (relevance, newest, oldest, cited), page, size (1–50), format. The answer is JSON: the results, the total (exact up to 500), the facets' counts, and what the search cost in database rows. When the quota is spent it answers 503 with the code quota.
The badge
One static image for every paper, linked to the paper's page from its code's README: how, and its rules.
In development
Not available on this site. What follows is being built on branches of the project's repository that are not merged or deployed. It may change, or not ship.
- A research layer over GitHub
- Researchers' repositories would stay in their own GitHub accounts, operated through a GitHub App with their consent, one authorization at a time, and read in OSCR's own viewer: the code with its tracing maps' lines in colour, editing in the browser, pull requests and releases linked to the paper's versions, and research issues — an error in the code, a mismatch between code and paper, a failed reproduction — attached to a paper and its map. GitHub would be the storage, and the last resort for what the registry cannot show.
- The
oscrcommand line for researchers - A program for the terminal, standard-library Python only, that links a repository to its paper, checks it as the registry does, traces its lines to the paragraphs of the Methods, produces a citation, and works with the GitHub repository — never running the code it reads. Not published yet.
- A public API with tokens
- Versioned read and write routes under
/api/v1/, by personal tokens only — shown once, stored as a digest, each with its rights and an expiry — with limits per token and a reference built from the routes themselves. - Checks that never run the code
- On a repository's pull requests and at any commit: its licence, an environment file saying how the code runs again, its link to the paper's DOI, a usable CITATION.cff, the files its tracing maps point to, file sizes, a README that says how to run it — read as text, nothing built, installed or executed; a change fails only when it breaks what a paper relies on.
Feedback
OSCR has no email address. About a paper — a wrong match, a missed link, a misjudged repository — use that paper's own forms, signed in: a correction and its note reach the operator (how). About the experiments themselves, an idea or a bug, open an issue on the project's GitHub.
