Skip to content

FAIRSCAPE

FAIRSCAPE packages research data, software and computations as RO-Crates with deep provenance: what was used, what produced it, and who ran it. It renders them as datasheets and evidence graphs, grades them for AI-readiness, and publishes them. It implements the Bridge2AI AI-readiness principles and ethical FAIRness1, and produces human- and machine-readable datasheets, including Croissant and Croissant RAI metadata.

FAIRSCAPE is a set of small Python packages. Each one does one step, and you can stop after any of them:

You already have

  • Datasheet for Datasets
  • Snakemake
  • Cromwell
  • Galaxy
  • MLflow
  • REDCap
  • Frictionless
  • Croissant
  • Python

1 Create

ro-crate-metadata.json

one JSON-LD file: every dataset, software and computation, linked by provenance

validated by fairscape-models

2 View & assess

scored by the AI-Readiness grader

3 Publish

  • Dataverse
  • Zenodo
  • Figshare
  • DataCite DOI

or host it yourself with fairscape-lite

What a crate looks like

Every crate below was made from a real workflow run. This one is the Snakemake variant-calling tutorial, run for real, then converted, rendered and graded with the commands in the quick start. Drag the graph, click a node.

Evidence graph: variant calling on three sequenced samples Open full page
13 computations, 22 datasets and 8 pieces of software, linked by what used and produced what. Built by fairscape-artifacts from the crate the Snakemake reporter wrote.

1. Create an RO-Crate

2. View and assess it

3. Publish it

Use Cases

FAIRSCAPE was first built to support fully provenanced, complex computations in predictive analytics for clinical research2. It now supports the NIH Bridge2AI program’s Functional Genomics Grand Challenge3, as well as ongoing clinical work in predictive analytics. We think it is useful for building pre-model explainability into AI applications4, and we plan to extend it in that direction. Its HTML datasheets extend the basic ideas of Gebru et al. 20215 to cover the extra metadata that biomedical AI-readiness requires.

Funding

FAIRSCAPE was developed with funding from the U.S. National Institutes of Health awards OT2OD032742, OT2OD032701, 5R01HD072071-10; and from the University of Virginia’s Frederick Thomas Fund.

Footnotes

  1. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3, 160018. doi:10.1038/sdata.2016.18. ↩

  2. Niestroy JC, Moorman JR, Levinson MA, et al. Discovery of signatures of fatal neonatal illness in vital signs using highly comparative time-series analysis. npj Digit Med. 2022;5(1):6. doi:10.1038/s41746-021-00551-z. ↩

  3. Clark T, Schaffer LV, Obernier K, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. doi:10.1101/2024.05.21.589311. Published online May 6, 2024. ↩

  4. Clark T, Caufield H, Parker JA, et al. AI-readiness for Biomedical Data: Bridge2AI Recommendations. Published online October 25, 2024. doi:10.1101/2024.10.23.619844. ↩

  5. Al Manir S, Levinson MA, Niestroy J, Churas C, Parker JA, Clark T. FAIRSCAPE: An Evolving AI-readiness Framework for Biomedical Research. doi:10.1101/2024.12.23.629818. Published online May 6, 2025. ↩