OSCR

Labs

OSCR is young, and much of it is an experiment in how research code can be found, read and cited. This page gathers what is experimental, how to build on the registry's data, and what is being built — marked as such.

Experiments you can use today

The Code ↔ Paper matches
Method lexical-v1: paragraphs and code units reduced to technical terms, a pair kept when several rare terms agree. A first method, measured by hand on 25 papers; its next steps are a better reading of the paper's structure, a real parser for the code, and a local model. The reader shows every match with its evidence and score.
The paper from two sources
When Europe PMC does not answer, the reader reads PubMed Central's copy, whose paragraphs may be numbered otherwise, and places each match by its section and evidence terms.
The classification
Rules over a paper's title, keywords, MeSH terms, subjects, abstract and journal, with their reasons; a local model will settle the ambiguous cases once compared with hand labels (how).
A registry within a fixed number of files
Whatever the catalogue's size, the site keeps a bounded number of files: authors, journals, institutions, tools and datasets are rendered in the browser from a fixed number of shards, and older papers by the Worker from their records.

The open datasets

OpenScientificCodeRegistry/Database, on Hugging Face: the authors' scripts under verified licences, deduplicated by the SHA-256 of their text, as Parquet blocks (blocks/NNNNN.parquet: sha256, language, size, lines, content; 64 rows a row group; zstd) with one JSON manifest per repository (manifests/<xx>/<repository>.json). Accessing the data gives the layout in full, and how to read it with Python or DuckDB.

Reading one script with range requests

The blocks are cut for the browser: a reader fetches a block's footer once, then only the row group that holds the file it wants — about 78 KB for one script — by HTTP range requests, which Hugging Face answers with the headers a page on another site needs. With hyparquet and its zstd decompressor:

// A script's text from the open dataset, in a browser: the block's footer once, then the one
// row group (64 rows) that holds the file, by HTTP range requests. npm: hyparquet, hyparquet-compressors.
import { asyncBufferFromUrl, parquetReadObjects } from "hyparquet";
import { compressors } from "hyparquet-compressors";

const base = "https://huggingface.co/datasets/OpenScientificCodeRegistry/Database/resolve/main";
// 1. The repository's manifest: its commit, its licence, and where each file's text is.
const repo = "github.com/owner/name";
const shard = [...new Uint8Array(await crypto.subtle.digest("SHA-1", new TextEncoder().encode(repo)))]
  .map((b) => b.toString(16).padStart(2, "0")).join("").slice(0, 2);
const manifest = await (await fetch(`${base}/manifests/${shard}/${repo.replace(/[^A-Za-z0-9._-]+/g, "__")}.json`)).json();
const file = manifest.files.find((f) => f.path === "analysis/run.py");
// 2. That row of that block.
const block = await asyncBufferFromUrl({ url: `${base}/blocks/${String(file.block).padStart(5, "0")}.parquet` });
const [row] = await parquetReadObjects({ file: block, compressors, rowStart: file.row, rowEnd: file.row + 1 });
console.log(manifest.license, row.sha256 === file.sha256, row.content);

Check the file's sha256 against the manifest's, and keep the licence the manifest names. A block, once published, never changes: its rows can be cached for good.

The tracing maps' format

A map is one JSON file, tracing-map.json, format tracing-map/0.1, deposited on Zenodo under CC0-1.0 once an author validated it (what a map is). Its shape, shortened:

{
  "format": "tracing-map/0.1",
  "paper": { "doi": "10.1234/abcd", "title": "…", "journal": "…", "published": "2026-09-21",
             "authors": ["…"] },
  "code": [{
    "repo": "github.com/owner/name", "url": "https://github.com/owner/name",
    "state": "alive", "license": "MIT", "commit": "7a034b0a…", "commit_date": "2026-09-01T10:00:00+02:00",
    "type": "", "software_heritage_archived": true, "level": "inventoried",
    "found_by": "text:availability", "section": "Code availability",
    "files": [{ "path": "analysis/run.py", "language": "Python", "digest": "c26da9b3…" }]
  }],
  "alignments": [{
    "paragraph": 42, "section": "Methods › EEG preprocessing", "repo": "github.com/owner/name",
    "path": "analysis/run.py", "start_line": 10, "end_line": 38, "symbol": "preprocess",
    "score": 3.1, "evidence": ["notch 50", "ica", "l_freq"], "method": "lexical-v1"
  }],
  "proposed": { "by": "oscr", "on": "2026-09-29" },
  "validated": { "by": "Family, Given", "orcid": "0000-0000-0000-0000", "on": "2026-10-02", "proof": "orcid" }
}
code[].level
found (the paper cites it), alive (it answers), inventoried (files listed, commit recorded).
code[].files[].digest
the SHA-256 of the file's text at the recorded commit.
alignments[].paragraph
the paragraph's position among the <p> of the <body> of the paper's open-access XML at Europe PMC, counted from 0 in document order; evidence holds short terms only, never a sentence.

validated appears once an author validated the map. The digest a validation carries is the SHA-256 of the map without proposed and validated, serialized as compact JSON with sorted keys.

The site's own files

Every record the pages show is a public, static JSON file, rebuilt every night. They serve the pages first; their layout may change.

/lookup/NN.json
The DOI lookup: 256 shards; NN is the first 2 hex characters of the SHA-1 of the lower-cased DOI. Each maps a DOI to [status, day read], and the page's identifier when it has one.
/records/<type>/NN.json
The authors (1,024 shards), institutions (512), tools (256), journals (128) and datasets (128): each shard holds entities by key and the rows of their papers. A key's shard is the leading bits of the SHA-1 of the key (an ORCID iD in capitals; any other key in lower case), as many as the number of shards needs, in hexadecimal.
/records/paper/NN.json
The records of the papers whose pages are rendered on demand (256 shards, by the SHA-1 of the page's identifier).
/badge.svg, /sitemap.xml
The badge; the sitemap and its shards, listing every page.

The search's API

The search page asks GET /api/search, which anyone may call — each call counts against the site's daily free quota, so call it sparingly:

https://openscicode.org/api/search?q=eeg&modality=eeg&sort=newest&page=1&size=20
https://openscicode.org/api/search?q=tool:mne&format=csv        (the first 500 results, CSV or JSON)

Parameters: q (the query language), the filters (status, year, modality, organism, population, subfield, tool, language, journal, data, host, code_license, type, license, matches, oa, repeated for OR), from, to, sort (relevance, newest, oldest, cited), page, size (1–50), format. The answer is JSON: the results, the total (exact up to 500), the facets' counts, and what the search cost in database rows. When the quota is spent it answers 503 with the code quota.

The badge

One static image for every paper, linked to the paper's page from its code's README: how, and its rules.

In development

Not available on this site. What follows is being built on branches of the project's repository that are not merged or deployed. It may change, or not ship.

A research layer over GitHub
Researchers' repositories would stay in their own GitHub accounts, operated through a GitHub App with their consent, one authorization at a time, and read in OSCR's own viewer: the code with its tracing maps' lines in colour, editing in the browser, pull requests and releases linked to the paper's versions, and research issues — an error in the code, a mismatch between code and paper, a failed reproduction — attached to a paper and its map. GitHub would be the storage, and the last resort for what the registry cannot show.
The oscr command line for researchers
A program for the terminal, standard-library Python only, that links a repository to its paper, checks it as the registry does, traces its lines to the paragraphs of the Methods, produces a citation, and works with the GitHub repository — never running the code it reads. Not published yet.
A public API with tokens
Versioned read and write routes under /api/v1/, by personal tokens only — shown once, stored as a digest, each with its rights and an expiry — with limits per token and a reference built from the routes themselves.
Checks that never run the code
On a repository's pull requests and at any commit: its licence, an environment file saying how the code runs again, its link to the paper's DOI, a usable CITATION.cff, the files its tracing maps point to, file sizes, a README that says how to run it — read as text, nothing built, installed or executed; a change fails only when it breaks what a paper relies on.

Feedback

OSCR has no email address. About a paper — a wrong match, a missed link, a misjudged repository — use that paper's own forms, signed in: a correction and its note reach the operator (how). About the experiments themselves, an idea or a bug, open an issue on the project's GitHub.