The paper's full text
Neither the full text nor the PDF of a paper ever leaves the operator's computer. OSCR links to the paper by its DOI, and the Code ↔ Paper reader shows its text by having your browser fetch it from Europe PMC.
Never stored on the site, never served
The site, its search, its files, its open data and its tracing maps hold no paper's body text and no PDF. A paper's page shows its bibliographic record and, under an open licence only, its abstract and availability statements (the licence policy). Everything else about the paper is a link: to doi.org, to the publisher, to Europe PMC.
How the reader shows a paper
When you open a paper's reader, your browser asks Europe PMC for the paper's open-access XML, and, if Europe PMC fails, NCBI for PubMed Central's copy. The text goes from their servers to your screen; it does not pass through OSCR, which never sees it. The page's security policy names the services the browser may contact; for the paper's text, those two. If they are down, the reader shows links to the paper, not a copy.
Evidence without sentences
A match between a paragraph and lines of code refers to the paragraph by its number (its position among the paragraphs of the paper's XML) and its section's title, and gives as evidence a few short technical terms — no term longer than 60 characters is kept, so no sentence can pass. Where a link was found is said by its place ("the availability statement", "the references"), never by the sentence.
What the harvester keeps, and why
To find the code, the harvester must read the paper. On the operator's computer it keeps the full texts it read, in its cache (to read them again without asking Europe PMC twice), and, in its private database, the sentences that decided each verdict (so that a verdict can be checked and explained to an author who asks). They never leave that computer: the public export removes them, and the site's build drops, again, any evidence term long enough to be a sentence.
How it is checked
- The public database is made from the private one by removing every table and column that holds a paper's text.
- The site's build keeps only the known fields of a match, and drops any evidence term long enough to be a sentence.
- The pages' Content-Security-Policy lets a paper's page connect only to the services it names — Europe PMC and NCBI for the paper's text, and the hosts of the authors' code for a file shown from its source (the code policy) — and every other page to the site alone.
