Model card

What has been tested, what the tests found, and where the model is known to be weak — assembled from the published validation artifacts.

Read this first. Most panels below are internal-consistency diagnostics: they check that the model's machinery is honest about its own uncertainty and that it extracts the signal its inputs contain. They are not evidence that the model predicts real patient outcomes. The genuinely external checks are labelled as such. Every number is read live from the JSON the generating script wrote, so this page cannot drift from the analyses it describes.

How good could any model be here?

Parameter recovery: simulate centers whose truth is known, run the pipeline, and see how well it recovers the ordering.

Are the intervals honest?

A 95% interval should contain the truth 95% of the time. Measured against held-out later SRTR releases.

Is the Bayesian machinery calibrated?

Simulation-based calibration: draw parameters from the prior, simulate data, refit, and check the rank statistics are uniform. A failure here would mean priors, likelihood and sampler disagree.

Does the ranking match the registry?

Per-center Spearman correlation between predicted access and observed SRTR transplant rates.

Pediatric

Children are modeled from pediatric SRTR cohorts and validated against SRTR's own risk-adjusted pediatric tier — the one pediatric ground truth that is not an input to the pipeline.

Would pooling across releases help?

A design question tested rather than assumed: does shrinking estimates across SRTR releases beat using the latest one?

When was each analysis last run?

Every published artifact, newest first. A stale date is itself a finding.

How much does the ranking depend on the weights?

The score is a weighted sum of eight categories. Those weights are a judgement, not a measurement — so the question is whether a different reasonable judgement gives a different answer.

What is the ranking actually made of?

A weight only moves the ranking through a sub-score that varies between centers. A category scoring the same everywhere adds the same constant to every total and cannot reorder anything — however large its weight. So the weights the scoring panel displays are not the same thing as what the ordering is built from.

Which of your answers change the recommendation?

The tool asks for several things about you and returns a ranked list of centers. Those are two different questions: whether an answer changes your numbers, and whether it changes which center comes out on top. They do not have the same answer, and the difference is not currently obvious from the interface.

Known limitations

The full register lives in the repository; these are the ones that most change how the output should be read.