Testing a new medicine produces a mountain of paperwork, and today almost all of it is copied into one company's software, where the copy becomes the official version. This console is a set of drawings of what the software could look like if the paperwork stayed with whoever wrote it and everyone else was given a way to find it and ask questions of it.
No experience of clinical trials is assumed on this page. The six screens after this one are written for people who do this for a living and they move fast; every term they use is defined in plain words in the vocabulary at the bottom of this page.
A trial is not run by one organisation. It is run by several, each of which keeps its own paperwork, and each of which is separately answerable for it. That plurality is the whole subject of this console, so it is worth being concrete about who is in the room. The colours are used consistently on every screen here.
The person who has agreed to be in the trial. They — or the clinician sitting with them — produce the consent form, the symptom diaries, and the observations recorded about their health. In this design they are treated as one of the record holders rather than as a subject the records are about.
The hospital or clinic where participants are actually seen. The senior doctor answerable for the trial there is called the investigator. A trial usually runs at many sites at once; this synthetic study has two, called Northgate and Rivermead.
The organisation that commissions the trial and answers to government regulators for it: a pharmaceutical company, a university, sometimes a hospital. Sponsors commonly hire a contract research organisation, or CRO, to run the day-to-day operations for them. Hiring one does not move the responsibility.
An independent board that has to approve the trial before it may start, and review it again when anything material changes. In the United States it is called an Institutional Review Board, or IRB, and that is the abbreviation used on the other screens.
Makes, labels and ships the drug being tested, and keeps the records showing that each batch was made, stored and transported properly. If those records are missing, the results built on that batch are hard to defend.
A government body — the Food and Drug Administration in the United States, the European Medicines Agency in Europe, and their equivalents elsewhere — which may arrive and check that the trial was run the way it was described. They hold no records of their own, but much of the paperwork exists because they might ask.
The records that matter here are the ones known as essential records: the documents that let someone who was not present conclude that the trial was run properly and that its results can be believed. Ethics approvals, signed consent forms, evidence that the staff were qualified, drug shipping and temperature logs, records of monitoring visits, and the protocol — the document setting out exactly what the trial will do — are all essential records.
Assembled together, that collection is called the Trial Master File, almost always shortened to TMF. When it is held in software rather than in filing cabinets, which it now always is, people write eTMF — the "e" is simply "electronic". An eTMF is therefore a product category, and a large industry sells it.
The rules for all of this are an international guideline called Good Clinical
Practice. Its current version is written by the International Council for Harmonisation
and is known by its document number, ICH E6(R3). Wherever the other screens show
something like C.2.3 or 3.9.1, that is only a paragraph address inside
that guideline, quoted so you can look it up and disagree. The full text is reproduced in this
bundle under reference/ich-e6r3/.
In current practice, every party copies its records into the sponsor's system, and the sponsor's copy becomes the version that counts. Everyone else — the site, the ethics committee, and above all the participant — gets a guest account, or simply hands things over. This is so standard that it is easy to assume the rules require it.
They do not. The guideline says essential records should be maintained in or referred to from repositories held by the sponsor and by the investigator and institution — repositories, plural, and a pointer is enough. It says each party should keep a record of where the records are. And it says the original should generally stay with the party who produced it. Those three sentences describe a set of separately held collections with an index over them, which is not what the industry built.
Copies flow one way, into a single system. The sponsor's copy becomes the official one, and the sponsor's software vendor sits underneath all of it.
Once the copy is the authoritative one, the party who wrote the record can no longer answer for it — and there is nothing left to compare the copy against.
Originals stay put. A shared index records only where each one lives, and each party assembles the view its own responsibilities require, on demand.
Reading anybody's record spends a narrow, time-limited permission that is logged. The person who holds the record can see that it was read, and can withdraw the permission.
Everything on the six screens after this one follows from these, and nothing else about them is subtle.
On the Readiness screen, two of the six record holders have not responded. Rather than scoring the four who did and presenting that as the study's readiness, the calculation declines to produce a number at all, and names who is missing and how much of the picture they held. Switch them back on and the score appears. It is the clearest difference between this and a system that holds one copy of everything, because a system holding the only copy always believes its view is complete.
Both columns are asserted, not measured. The right-hand one is not a list of risks that careful engineering would remove; it is what the design gives up on purpose.
Fewer copies to keep in step. The record is in one place, held by the party who can actually answer questions about it.
Contradictions between two parties become visible. If the sponsor's file says one thing and the participant's own record says another, that disagreement is something you can look for. A system holding a single copy of every answer has nothing to compare against.
Records survive institutional change. A participant who moves city, a clinic that closes, a sponsor that is acquired — the records do not have to be migrated, because they never moved in the first place.
Being inspected need not mean handing over everything. An inspector needs to establish that the evidence exists and is consistent, which can be answered without shipping copies of every participant's file.
You lose the single confident number. Completeness arrives bounded by how many parties answered, and the readiness score sometimes refuses to exist. If you need one percentage for a board meeting regardless of what stands behind it, this is worse for you, and noticing that is not a misunderstanding.
Corrections get slower. Nobody can quietly fix someone else's record. They ask, and the holder answers, which takes longer and leaves a clearer trail.
You depend on other people replying. A site that does not respond degrades your view, and this design's answer is to tell you so rather than to paper over it.
Nobody obviously owns the index. If the shared list of where records live is wrong, it is unclear who is accountable. This is the most serious unanswered objection in the bundle and it is not solved here.
Each screen takes one job an eTMF product already does and asks what it would have to say differently if the records were not all in one place. They can be read in any order.
"What kind of document is this, where does it belong, and who may put it there?"
Software reads an incoming document, guesses what it is, and says how sure it is. That confidence figure decides whether a person needs to look, not whether the guess is right. The screen adds a column an ordinary system has no need for: whether you are permitted to act at all, given who holds the record.
"What paperwork should exist that we have not got?"
The interesting part is where the target number comes from. "Eleven staff qualification records are expected, because the delegation log names eleven people" can be checked and argued with. "Twelve documents, per the standard list" cannot. It also allows a record to be legitimately absent, with a reason and someone's name against it.
"Is this document actually any good?"
QC is short for quality control. Two kinds of finding are kept apart and never averaged: facts a rule settles, such as a missing signature, and opinions a model offers, such as whether a scan looks complete. A check also stays marked candidate until a human has agreed what it should do, and candidate checks are kept out of the scoring.
"Do two parties' records contradict each other?"
The sponsor's file records consent version 2.1; the participant's own record says 3.0. Neither file is wrong in itself, and both were filed correctly — the fault exists only in the relationship between them. This is the screen a single central system cannot draw, rather than merely draws badly.
"Could we survive an inspection tomorrow?"
One overall score, assembled from five inputs whose weightings are shown and adjustable, with the weakest input named and the remaining work ordered by how much each action would actually move the score. And, when someone has not answered, no score at all.
"Which of these software vendors is any good?"
Several vendors read the same records and each publishes how often its suggestions were rejected by the people reviewing them. A high rejection rate is not necessarily a bad product; it may be a rule nobody has agreed yet. Not publishing the figure at all is the problem.
They are invented. COMMONS-STUDY-001 is a made-up study, written by hand to make each pattern legible. The 412 documents, the 0.86 threshold, the rejection rates, the two unresponsive parties — all authored, none measured. This is a drawing of an instrument, not a reading from one, and no part of it has been built, run, reviewed or tried against a real trial.
Two things in the wider bundle are mechanically checked, and neither of them is on this
page: that every addressable paragraph of the guideline has been given an explicit disposition,
and that every place the encoding claims "this paragraph is handled there" points at a document
that exists. Those checks are what ruby tools/validate_tsdg.rb runs. Everything about
whether any of this would work in a real study is untested, and the bundle names the specific bets
it is making so that losing them is visible.
Every abbreviation and piece of shorthand used on the other screens, in plain words. The other screens do not stop to explain these; this is where they are explained.
Trial.Records.Classify@v1.
There are 77 of them in this bundle and none of them are running code.It is not a product, and there is nothing to buy, pilot or evaluate. Nothing described here has been built. No record has been filed, no party has joined, and no software runs any of it.
It is not advice, and it is not a claim about the rules. The reading of the guideline set out here was made for engineering purposes. It is not a regulatory interpretation, not a statement about whether anything complies with anything, and not legal advice. The body that writes the guideline has not reviewed it. Do not change how a trial is run on the strength of this.
Nobody was consulted. No participant, patient organisation, ethics committee, sponsor or regulator was asked, and none is represented as supporting any of it. Where a patient community holds a view about who should hold their records, that view outranks this design.
The way these screens work is somebody else's design. The shape of the six views is drawn from published work by IntuitionLabs, credited in full behind About & attribution in the bar above. What is different here is what the screens are built on top of, and therefore what they are permitted to claim.
A model reads a record, proposes a kind and a destination, and stops. The custodian files.
Adopted: confidence as a routing signal rather than a truth claim; consequence and confidence as independent axes. Added here: a third axis — whose custody the act writes into — and a hard separation between proposing and filing.
Horizontal is the model's confidence. Vertical is the consequence of acting on it. The shaded region is the only one where the substrate acts unattended. A record at 0.97 confidence still goes to a person if the act writes into someone else's custody — which is the axis a single-tenant system does not have.
Every row states what was found, what evidence supports it, what may be done, and — the column that only exists here — whether you can act at all. Where the responsible custodian is another party, the item marks itself unactionable and names them.
| Record | Proposed kind | Conf. | Route | Custody destination | Evidence | Disposition |
|---|
It will not file. Trial.Records.Classify@v1 emits a proposal with a
custody destination; the write is Trial.Graph.AssertRecordNode@v1 executed under
the custodian's key. There is no configuration in which the substrate holds a key that can
file into a participant's or a site's records.
It will not show the confidence as a quality badge. The number decides which path the record takes. It is not an estimate of whether the answer is right, and displaying it as one would be a category error — so the route and its reason are what the reviewer reads.
What records should exist, computed from this study's own records, with a denominator you can argue with.
Adopted: expected records derived from a named source and a countable set, not a template checklist; derived and template coverage never summed. Added here: the custodian expected to produce each record, an explicit third terminal state, and locator index coverage as a third denominator that bounds the other two.
These are never added together. The first is derived from the study; the second is a conventional expected-document list, reported only for continuity with existing practice; the third bounds both, because a percentage over custodians who did not answer is a percentage over nothing.
Three states, not two. Explained means ICH E6(R3) C.3.3 applies — the record legitimately was not produced — and it carries a reason and a named attesting party. A two-state model insists every expected record must exist, which C.3.3 says plainly is not so, and the result is a queue that is permanently noisy.
resolved explained (C.3.3) open custodian did not answer
Each row's denominator has a provenance. "Three CVs are expected, derived from the delegation log at page 2, rows 4 through 11" can be checked, disputed and corrected. "Twelve documents per the standard TMF index" cannot be disputed at all, which is why it is not used here.
| Expected record | Derived from | Denom. | Custodian | Matching criteria | State |
|---|
Every derivation rests on a claimed morphism in registry/morphisms.yaml, and four
are declared: a delegation log implies qualification records, a protocol implies its expectation
set, an IRB review implies approvals and reconsent, and a CTQ declaration constrains which rank
first. Each is a hand-authored claim about how real studies work, made without having
examined one. If a delegation log does not reliably imply the qualification records the
morphism claims, this whole screen degrades to a checklist with better citations. That is
H-CTM5, and it is untested.
Deterministic failures and model judgments are reported separately, and a check that nobody has agreed to is labelled as such.
Adopted, without modification: the candidate-versus-running check state, and the refusal to blend rule-based and model-based results into one score. These are the two disciplines this bundle admires most in IntuitionLabs' published design. Added here: the locator is a resolvable handle rather than only a page reference, because the reviewer may not hold the record.
A signature block is absent or present. A date is after another date or it is not. These either pass or fail, and a failure is never softened by a model that disagrees.
Legibility, whether a version looks superseded, whether a scan is complete. Useful, reviewable, and epistemically different from the column on the left — so it is not averaged with it.
A check is candidate until its source rule, threshold, exceptions and expected action are agreed with the sponsor. A model can detect that a CV is two years old; only the sponsor determines whether that requires replacement. Candidate findings are shown — hiding them would waste the detection — and they are excluded from readiness scoring, because a score built on rules nobody agreed to is a score nobody can defend.
| Check | Kind | State | Findings | Rejection rate | In readiness? | What promotion needs |
|---|
No extracted value without a locator. Here the locator is a page reference plus a resolvable handle, and resolving it consumes a scoped grant and is logged. The reviewer sees where to look; the custodian sees that they looked. Click a handle to spend a grant.
| Field | Value | Source record | Locator | Custodian | Grant |
|---|
The finding that lives in the relation between two records, where each record is internally valid and correctly filed.
Adopted: reconciliation across systems as the differentiating capability, and the four named checks — consent version against subject-level events, active site status against filed signature pages, monitoring visit timeliness against committed cadence, and eConsent completions against audit certificates. Added here: the argument that a single system of record cannot form this finding at all, and a permitted disposition that disputes the check itself.
Each arc is a contradiction between two custodians' records. Neither record is wrong on its own; the problem exists only in the pair. A single system of record holds one version of every answer by construction, so it has nothing to compare against — it cannot draw this diagram slowly or partially. It cannot draw it.
Every row states what one side holds, what the other says, which direction the difference runs, and why it matters. Without the consequence note this is a mismatch list; with it, it is a work queue. The disposition set always includes disputing the rule, which routes back to the check catalogue rather than closing the item.
| Check | Side A | Side B | Direction | Consequence | Sev. | Disposition |
|---|
The eConsent reconciliation compares completions against an audit certificate, and the
certificate distinguishes integration-user actions from human signatures. A service
account's action is never relabelled as the human's — the audit trail has to survive someone
asking who actually signed. Generalised here to all connector activity via
Trial.Ecosystem.BindConnector@v1.
A composite that never travels without its weights and inputs, and refuses to exist over a partial view.
Adopted: a weighted composite with visible, adjustable weights; the weakest contributing input named alongside the score; remediation ranked by the score movement each action would produce. Added here: weights carry provenance, zero weights are refused, each action names its responsible custodian, and the whole computation refuses when the index is incomplete.
Toggle a custodian off to simulate a non-response. This is the control worth playing with, because it is where the substrate differs most: a partial view that knows it is partial supports a correct decision; one that reports a confident number does not.
Not by severity and not by count — by the arithmetic of what each action would move. Each bar names the custodian who must act, and actions blocked by a non-delegable act are marked unworkable and listed separately, because ranking something nobody is permitted to do is ranking noise.
The five weights below were chosen to make the demonstration legible. They are not
calibrated against anything, and every output that carries them says so.
Trial.Readiness.WeightInputs@v1 requires a weight profile to record who set it,
when, on what rationale, and against which critical-to-quality factors — and refuses a zero
weight, because a zero silently removes a dimension while appearing to include it.
Several vendors reading one graph, each publishing the rate at which its recommendations were rejected.
Adopted: recording rejections as carefully as adoptions, because the rejection rate per check is how an automation's quality is actually measured. Added here: publishing those rates as a market mechanism, and the whole conformance surface — which is this bundle's own argument, since ICH E6(R3) is silent on ecosystem structure and says nothing about credit.
An offering competes on measured quality rather than on marketing. A high rejection rate is not a scandal — it is a check whose rule has not been agreed yet, or a model that is wrong about this study. What would be a scandal is not publishing it.
| Offering | Reads | Proposes | Adopted | Rejected | Rejection rate | Conformance | Attribution owed |
|---|
This is the constraint the other screens were re-expressed under, and it is the boundary of the whole ecosystem argument. A vendor's read capability and its write capability are separate grants, and no vendor holds a write grant into another party's custody. The registry column above reads zero because it cannot read anything else.
IntuitionLabs' design assumes a repository the software can write into, which was the reasonable assumption before a commons was available. That is a difference in substrate, not a criticism.