OSCR

The authors' code: copies and sources

OSCR shows the code that the authors of a paper published with it. When the code's license allows redistribution, OSCR keeps a copy. When it does not, OSCR keeps no copy: your browser fetches each file from where its authors published it, checks that it is the file OSCR verified, and shows it. The authors keep every right to their code, and can remove it or change its license at any time.

The files OSCR copies, and why

A copy is kept of a file of the authors' code when the license of its repository allows redistribution and that license is verified. The copy serves two purposes: the file can be read beside the paper, in the Code ↔ Paper reader, even when its repository is slow or gone; and it is kept, with its license, in the open dataset of scripts of OSCR (on Hugging Face), where each file keeps the license of its repository, its address and the commit it was read at, for anyone to reuse under that license.

  • A copy is the text of the file as its authors published it, at the commit or version that was verified. Three things only differ: the email addresses are hidden (below); a Jupyter notebook is kept as the text of its cells, without their outputs; and a text longer than 200,000 characters is cut, which the reader says, with a link to the whole file at its source.
  • Each unique file is stored once, however many repositories hold it, and never changed once published: a new version of a file is a new file.
  • The license travels with the copy: the reader and the dataset name it, and link to the source.

The licenses that allow redistribution

A repository's files are copied under one of these licenses (their SPDX identifiers):

  • MIT, BSD-2-Clause, BSD-3-Clause, BSD-3-Clause-Clear, 0BSD, Apache-2.0, ISC, Unlicense, Zlib, BSL-1.0, Python-2.0, Artistic-2.0, MPL-2.0, GPL-2.0, GPL-3.0, LGPL-2.1, LGPL-3.0, AGPL-3.0, CECILL-2.1, EUPL-1.2, CC0-1.0, CC-BY-4.0, CC-BY-SA-4.0;
  • and, with their conditions kept, CC-BY-NC-4.0, CC-BY-NC-SA-4.0, CC-BY-NC-ND-4.0, CC-BY-ND-4.0: the non-commercial ones (NC) for non-commercial reuse only, as the dataset says; the no-derivatives ones (ND) as they are, which is how a copy is kept.

Code with no license at all is not free to redistribute: by default, its authors keep all their rights. The same goes for a license that forbids redistribution, and for one OSCR cannot recognize.

What "verified" means

  • A repository on a forge (GitHub, GitLab, Codeberg, Bitbucket…): its license is read from its own license file, at its root (LICENSE, COPYING…), at the commit that was read. The file's text is recognized as one of the licenses above; when a license names others in passing (the GPL names the LGPL), the one the text starts with counts.
  • An archive (a Zenodo record, figshare…): its own license file, at its root or one folder down (a release's name-v1.0/LICENSE), else the license of its record.
  • Never a license inferred from a sentence of a README, nor "some open license" without a license file, nor a guess. A repository whose license cannot be verified this way is treated as one without a license.

The license is read again each time the repository is verified: a license added later is taken into account then. The rule was audited before the first copy left OSCR's own machine, and is audited again before any change.

The files OSCR shows from their source

When a repository's license does not allow redistribution, OSCR keeps no copy of its files — not on this site, not in its open data, not in its dataset of scripts. Its reader still shows them: your browser fetches the file itself, from where its authors published it, at the version OSCR verified, and shows it with its line numbers, its colors and its matches with the paper. A notice above the file says where it comes from, at which commit or record, that no copy is kept, and that its rights remain with its authors.

Why this is not a copy

  • The file travels from its host to your browser, as when you open it there. OSCR's servers never receive it, keep it or serve it; the page holds no line of it.
  • OSCR publishes only facts about the file: its path, language, size, number of lines, the fingerprint (SHA-256) of its bytes, the commit or record it was verified at, and the line numbers of its matches with the paper.
  • The authors stay in control: a file they delete, a repository they make private or move, cannot be fetched any more, and the reader says so. OSCR never shows it from elsewhere.

Where your browser fetches a file

GitHub
its raw file, at the verified commit (raw.githubusercontent.com)
GitLab.com
its raw file, at the verified commit, through GitLab's API
Bitbucket, Codeberg
the raw file at the verified commit
Hugging Face
the raw file at the verified commit
Zenodo
the file of the record, whose files never change once it is published
Software Heritage
for the other forges (a university's own GitLab, Framagit…), which do not let another site's page read them, and for a file inside an archive of a Zenodo record: Software Heritage's archive of all public code, which gives a file by its fingerprint — the same bytes, or nothing. It is not used for a repository that OSCR's last check found gone from its source.

Some hosts do not let the page of another site read their files: OSF, and PubMed Central's collection of supplementary files. Their files are not shown here; the reader links to them, at their source.

Fetching a file tells its host, as any visit does, your address and that a page asked for it; no cookie and no page address are sent with the request. A file is fetched only when the reader shows it: the file a paper's page opens on, or one you choose.

The integrity check

When OSCR's machine verified the repository, it read each file and computed the SHA-256 fingerprint of its bytes. Your browser computes the same fingerprint of what it received, with its own cryptography (crypto.subtle), before it shows anything.

  • The same fingerprint: the file is the one OSCR verified, byte for byte. It is shown, and its matches with the paper are drawn.
  • A different fingerprint: the host sent another file. Nothing of it is shown, no match is drawn, and the reader says so in words, with a link to the file at its source.
  • No answer, an error, a file larger than 1 MB, or a file that is not text: the reader says which, and links to the source. Nothing is shown from elsewhere.

Email addresses

OSCR never displays an email address. In the authors' code, each address is replaced by [email hidden], on the same line, so that the line numbers and the matches with the paper still hold: in the copies before they leave OSCR's machine, and in a file shown from its source by your browser, with the same rule, once the file is checked. The SSH address of a git repository (which starts with git@), a Python decorator or a MATLAB function handle is not an email address, and stays.

For the authors: remove your code, or change its license

  • To remove it from OSCR, use the removal request page, linked from every paper's page ("Request removal"): ask for the copies of one repository, of one file, or of all the paper's scripts. A request applies to the display from the source too: what it names is neither copied nor fetched any more, from the next nightly publication. Signed in with the ORCID iD the paper lists, or as the repository's owner (or a public member of its organization) on GitHub, your request is applied without waiting; the moderation policy says how every request is decided.
  • To allow copies, add a license file at the root of the repository (for example the MIT license, or CC BY for an archive's record). OSCR reads it at its next verification of the repository.
  • To stop the display from the source without a request, delete the file or make the repository private: the fetch then fails, and nothing is shown.
  • A copy that has already been published in the open dataset stays in its block of files, which never changes; once withdrawn, no index of OSCR points to it any more.