OSCR

About

OSCR

OSCR (Open Scientific Code Registry) is a registry of the code that the authors of open-access neuroscience papers publish with their work. It finds that code from the paper's own text, checks it where it lives, records exactly what was found and under which licence, and shows it beside the paper, paragraph by paragraph. It is free, it asks nobody for an email address, and it never claims more than it has verified: the source always prevails.

The mission

A paper's results come from code as much as from its methods section. OSCR exists so that anyone who reads a neuroscience paper can find the code its authors used, know whether that code is really there, read it next to the text it implements, and cite both. It does that at a scale no reader can: every open-access neuroscience paper it can read, every day.

Three commitments follow from that aim, and every other choice on this page serves them:

  • Find the code the authors themselves point to, not code that merely resembles theirs.
  • Verify it at the source, and say in words what was verified, when, and what was not.
  • Respect the authors, their licences and the readers: copy only what a licence allows, collect no reader's email address, and never republish a paper's text.

The problem

Code behind papers is hard to find, hard to verify and hard to trust. The reasons are practical, and each one shaped the registry:

It is hard to find.
A paper says where its code is in many places: a data and code availability statement, the methods, a table of resources, a reference to a software record, a supplementary archive, a footnote. Some of those places are not in any search index. A search engine for papers sees a paper's metadata; the link to the code is often only in its full text.
It is hard to verify.
A link can be dead, point to an empty repository, to a lab's home page, to a tool the authors merely used, or to data rather than code. "Available on request" is a promise, not code. Only a visit to each link, and a look at what it holds, tells these cases apart.
It is hard to trust.
A repository changes after publication. Without the commit that was there, its licence and the list of its files, nobody can say later what the paper's code was. And without the authors' own word, a link found by a machine stays a proposal.
It is hard to read.
Even when the code is there, the reader has to guess which function implements which paragraph of the methods.

How it works, step by step

A harvester runs without pause on the operator's own computer. Each paper goes through the steps below, each written to its private database before the next, so that an interruption loses nothing. Every night at 04:17 (the computer's local time) the public part is exported and the site is rebuilt.

  1. Harvest

    OSCR asks Europe PMC for the open-access papers whose full text Europe PMC holds and whose title or abstract matches a broad neuroscience query (EEG, MEG, fMRI, intracranial recordings, calcium imaging, electrophysiology, neuroimaging, brain stimulation, neurons, cortex, the hippocampus, the brain, and more). It takes the papers published since its last look every hour, and, in between, goes back through the stock month by month. For each paper it reads the full text (the publisher's JATS XML, as Europe PMC gives it), its DataCite metadata, and Crossref's when there is no text. It then enriches the record: Europe PMC's own record (MeSH terms, grants, corrections), OpenAlex (institutions with their ROR identifiers, the paper's topic, its open-access status, a preprint), and the Retraction Watch data (retractions, corrections, expressions of concern).

  2. Classification in scope

    The query is broad on purpose: "neural" also matches artificial neural networks, "cortical" cortical bone. So each paper is classified, first by rules over its title, keywords, MeSH terms, subjects, abstract and journal, each rule saying which term fired. A paper judged off-topic stays on the operator's computer: out of the site, the search, the open data and every figure (the scope). The same rules place a paper in categories — modality, organism, population, subfield — that the Browse pages and the taxonomy list. The operator's own labels win over the rules; the cases the rules find ambiguous are meant for a local model, run only between 01:00 and 07:00, once it has been compared with those labels.

  3. Finding code links in the paper's own text

    Every link in the full text is collected with where it was seen: the availability statement, the methods, a table, the references, the supplementary material, the notes. Each mention is judged on its own — the authors' code, their data, or a third-party tool they used — with the reasons for the verdict, and the verdicts for the same repository are merged. A sentence saying the code is "available on request" is recorded as such. The sentences that decided a verdict stay in the private database: the site shows where a link was found, never the paper's words.

  4. Checking that the links answer

    Each repository that may hold code is visited where it lives. A git repository (GitHub, GitLab, Codeberg, Bitbucket, GIN and others) is asked for its current commit, then cloned without its history to list its files and read its scripts; the commit and its date are recorded. A Zenodo, OSF, figshare or Dryad record is read through its service's interface; supplementary files from PubMed Central's open-access copy. Software Heritage is asked whether it has archived the repository. A service that refuses robots (Code Ocean) is said to be one that "cannot be verified", never guessed. Repositories are checked again every 30 days, and every check is kept: a paper's page shows the history of its code's availability.

  5. The licence check

    The licence is read from the repository's own licence file, at its root (for an archive, a licence file in it or else its record's licence). Licences name each other in passing — the GPL mentions the Affero GPL — so the licence whose text comes first wins. A licence mentioned only in a README sentence is recorded, but it is not enough for the open dataset's copies. A repository without a licence is "all rights reserved": its files are never copied; the reader's browser fetches each one from its source, at the verified commit, and shows it only when its fingerprint is the one OSCR computed.

  6. Copies of scripts, only under verified licences

    When a repository's licence allows redistribution and its own licence file (or its record) confirms it, the text of its scripts goes to the open dataset. The copies are deduplicated (a file present in several repositories is stored once, under the SHA-256 of its text), compressed with zstd and stored as Parquet blocks in the public Hugging Face dataset OpenScientificCodeRegistry/Database, with one manifest per repository (its commit, its licence, and where each file lies). A published block never changes. The blocks are cut so that a browser can read a single script with HTTP range requests — about 78 KB for one script shown. The site's pages hide any email address written in the code. Which files are copied, which are shown from their source without a copy, and when the reader sends you to the source instead, is the code policy's subject.

  7. The Code ↔ Paper reader and its tracing maps

    For a paper whose code was read, the harvester pairs paragraphs of the paper with lines of the code (method lexical-v1): both are reduced to technical terms — word stems, identifiers, constants, tool names, figure numbers — and a pair is kept only when several rare terms agree, so that a shared common word counts for nothing. A wrong highlight costs the reader more than a missing one. The reader puts the paper and the code side by side, each pair in one colour on both sides. The paper's text is fetched by your browser from Europe PMC (or PubMed Central's copy) when you open the reader; OSCR never stores it. All of it — where the code is, the commit, the licence, the files and their digests, how the links were found, the matches — makes the paper's tracing map.

  8. An author's validation

    A map made by the machine is a proposal. An author of the paper, signed in with the ORCID iD the paper lists, can validate it — or first correct its links. The page carries a digest of the map it shows; the validation carries it back, so the map kept is exactly the one the author saw. If it changed in between, the author is asked to look again.

  9. A DOI, for validated maps only

    A validated map is deposited on Zenodo, the free repository run by CERN, and receives a DOI. The DOI is on the map — its links and metadata — never on the authors' code, which stays in their repository. The record is IsSupplementTo the paper and References the code at the validated commit; its creators are the validating author, with their ORCID iD, and the platform. A map the machine proposed never receives a DOI (the DOI policy).

Why it is built this way

Zero cost.
A registry that depends on a budget disappears with it. OSCR uses only free services and free plans, and is designed around their limits: the site is static files, which Cloudflare serves without limit; the few dynamic answers (the search, sign-in, requests) share a daily quota; the heavy work runs on the operator's computer. When a feature does not fit a free plan, it waits or takes another form.
No email address.
Sign-in asks ORCID, GitHub or Google for an identity, never for an address. No form accepts one: the texts you type lose any address before they are stored. Notifications stay in the site. The site shows no author's email address, even inside their code. This keeps readers and authors out of mailing lists, and keeps a whole class of personal data — and of mistakes — out of the registry. (The operator does keep, privately, the contact details that papers themselves publish for their authors: the privacy page says what, where and why.)
Licences respected.
A script's text is copied only when its licence allows it and the repository's own licence file confirms it. Everything else is never copied: the reader's browser fetches it from where its authors published it, at the verified commit, checks its SHA-256 fingerprint and hides its email addresses before showing it, or links to it there when its host does not let another site read it (the code policy). Abstracts and availability statements are shown in full only under the licences that allow it (the licence policy).
The paper's full text is never redistributed.
Neither the text nor the PDF of a paper leaves the operator's computer: the site links to the paper by its DOI, and the reader's browser fetches the text from Europe PMC itself. Evidence for a match is a handful of short technical terms, never a sentence (the full-text policy).
GitHub only as a last resort.
The code is read here, beside the paper: a paper's page opens on the reader, and the listing's code links lead to it. The forge is the source of truth and is always linked, but last — for what the reader cannot show, and to check the original. Nothing depends on one forge: GitLab, Codeberg, Zenodo, OSF and the others are read the same way. And the registry asks GitHub for no permission: sign-in with GitHub requests no scope, and the badge is added by the author in GitHub's own editor.
The machine proposes; the authors decide.
Every verdict is labelled with how it was reached. Matches are proposals with their evidence and score. A correction by an author becomes a new version of the record and survives every later scan. Only an author's validation earns a DOI.
Static first, the source prevails.
Pages are built ahead of time, so the site works without JavaScript for reading records, costs nothing to visit, and cannot be overloaded by readers. When the registry and the source disagree, the source is right.

In figures

From the catalogue of 29 September 2026, 19:57 UTC, the one this site was built from:

Open-access neuroscience papers read (in scope)22,128
… of them research articles15,623
Papers with a page: the authors' code, code on request or data only4,974
Papers with their authors' code2,882
Code repositories cited as the authors' code4,386
… answering at their last check4,208
… archived by Software Heritage293
Scripts read at the source211,077
Papers with Code ↔ Paper matches2,065
Matches between a paragraph and lines of code15,692
Authors with an ORCID iD, journals, institutions15,678, 720, 4,754
Tools found in the code, datasets cited228, 3,299
Publication dates covered6 November 2020 to 29 September 2026

Rates are computed on research articles only: reviews, conference abstracts, case reports and notices rarely have code of their own and would dilute them. 18% of the research articles read have their authors' code.

Who runs it

OSCR is run by one independent operator, who develops it, runs its harvester, and reviews, when available, what the published rules of its moderation cannot decide. No institution, company, publisher or funder is behind it, and it is affiliated with none of the services it reads or publishes to (Europe PMC, Crossref, DataCite, OpenAlex, Zenodo, Hugging Face, Cloudflare, ORCID, GitHub, Google). Its source code is public under the Apache-2.0 licence: github.com/yannbellec/Open-Scientific-Code-Registry-OSCR-. The operator's name and address are given on the privacy page.

How it is paid for

It is not: OSCR costs nothing to run beyond the operator's own computer and connection, has no advertising, no tracker and no paid service. It lives on free plans:

  • Cloudflare Workers, free plan: the site's files, served without limit, and 100,000 dynamic requests a day for the search, sign-in, requests and the pages of older papers; D1 databases for the search and the accounts.
  • Hugging Face: the public datasets, free.
  • Zenodo: the tracing maps' DOIs, free.
  • ORCID's public API and sign-in, free for non-commercial use, which OSCR is; GitHub's and Google's sign-in, free.
  • Europe PMC, Crossref, DataCite, OpenAlex (single lookups, within its free daily allowance), Software Heritage, the forges and archives: read politely, within their published limits.

Should a free plan end, the site's static pages would keep working; the feature that depended on it would wait.

Open data

Public
OpenScientificCodeRegistry/Database on Hugging Face: the authors' scripts under verified licences, each file under its repository's licence, as Parquet blocks with one manifest per repository. The source code, Apache-2.0. The tracing maps deposited on Zenodo, CC0-1.0. And the site itself: every record it shows is a public JSON file (Labs lists them).
Private for now
The full catalogue as tables (papers, repositories, scripts, matches, and a public SQLite database without any paper's sentence), prepared under CC0-1.0 for Hugging Face: it stays private until the operator decides to open it.
Private, always
The authors' contact details that papers publish, kept on the operator's computer and in a private Hugging Face dataset (why and how); the harvester's full database, with the sentences that decided each verdict; off-topic papers.

Accessing the data explains how to use each of them.

How to cite it

To cite the registry as a whole (its catalogue changes every night: give the date you used):

The OSCR contributors (2026). OSCR: Open Scientific Code Registry. Catalogue of 29 September 2026, 19:57 UTC. https://openscicode.org/
@misc{oscr2026,
  title        = {{OSCR}: Open Scientific Code Registry},
  author       = {{The OSCR contributors}},
  year         = {2026},
  howpublished = {\url{https://openscicode.org/}},
  note         = {Catalogue of 29 September 2026, 19:57 UTC; source code: \url{https://github.com/yannbellec/Open-Scientific-Code-Registry-OSCR-}}
}

A paper's page gives its own citation formats (APA-like text, BibTeX, RIS, CSL-JSON), and a validated map's DOI is the way to cite a map. The citation policy covers each case, the datasets included.

Limits

What OSCR does not do, or not yet — said plainly, since a registry is only as useful as its honesty:

  • Coverage. Only open-access papers whose full text Europe PMC holds, matched by its neuroscience query; papers without a full text there are read from their metadata only. The harvest goes back in time month by month: older years are still being read.
  • Accuracy. Measured on 25 September 2026: 10 of 11 "authors' code" verdicts right on papers never seen while tuning; the text alone finds the authors' exact repository for 72% of the papers of an independent benchmark (73 of 102). Some links are missed, a few misjudged: authors can correct them.
  • Matches are computed by a lexical method, not read by a person: proposals, with their evidence.
  • Classification is by rules; the model meant for the ambiguous cases is not chosen yet.
  • Quotas. The search and the other dynamic answers share Cloudflare's daily free quota; when it is spent, they say so until midnight UTC, while every static page keeps working.
  • Test services. While the platform is built, sign-in with ORCID may use ORCID's sandbox and map deposits go to Zenodo's sandbox, whose DOIs are not real; such validations are recorded as tests and never published.
  • Moderation. There is no human moderator on duty: published rules decide removal requests, submissions and claims, on the operator's computer. What they cannot decide waits for the operator, 30 days at most, and is then closed (how requests are decided) — except a request about personal data, which the operator answers within one month, as the GDPR requires, and which is never closed unanswered (your data and your rights).
  • Not built yet: unlinking an identity from an account, signing out of every browser at once, discussions, reproduction reports, and notifications beyond the account page.
  • Address. The site lives at its free workers.dev address until a domain of its own, before the public launch.

History

25 September 2026
The harvester's first measured results: precision on unseen papers, recall on an independent benchmark of papers linked to Zenodo software records.
26 September 2026
The project's first public commit, in English; the website deployed on Cloudflare; the decision to store the scripts' copies as deduplicated Parquet blocks; the Zenodo sandbox community for the maps.
27 September 2026
The operator's founding decisions (licences of statements, which papers get a page, the search, sign-in, notifications in the site only, classification, scope, hosting on Cloudflare Workers). The enriched record, the navigation by category, author, journal, institution, tool and dataset; the search; the full paper page; accounts with ORCID, GitHub and Google; the licence audit, and the first block of scripts published on Hugging Face.
28 September 2026
Contributions live: submission, author claims, corrections, map validation, the badge, removal requests. OpenAlex enrichment; a fixed budget of files whatever the catalogue's size; each paper's page opens on the Code ↔ Paper reader, with PubMed Central as a second source of the paper's text.
29 September 2026
The removal request page; the automatic moderator and its published rules; the scripts whose licence allows no copy shown from their source by the reader's browser; these information pages, the logo, and the preparation of the public launch.

Feedback and contact

OSCR has no email address, on purpose. To reach the operator: