Rubric for Review of AI-readiness Evaluation Criteria, v1.8 (2026-09-10) · generated 2026-09-28 17:18 UTC · network checks disabled
Variant calling on three sequenced samplesark:59853/rocrate-variant-calling-on-three-sequenced-sampl-68f3c95 · ark:59853/rocrate-variant-calling-on-three-sequenced-sampl-68f3c95 · published 2026-09-11T13:35:02.543766-04:00 · RO-Crate
Deposit datasets in a searchable FAIR-compliant data repository.
Questions (2)
Does the dataset have a resolvable PID: ARK, CSTR, DOI, IGSN, PURL, URN, HDL?
Does the dataset have a compact ID whose prefix is registered, for example, with identifiers.org (e.g., IDs such as ena.embl:ERP001234, geo:GSE68086, dbgap:phs000001)?
2PID or prefix-registered compact ID is present, and resolves to a specific version of the dataset, deposited in a sustainable repository.
1Only a PID or compact ID is given, but the dataset is not in a sustainable repository. Alternatively, the data was deposited in a registered repository without a resolvable PID.
0Neither is present.
A sustainable repository assigns persistent identifiers, maintains descriptive metadata that remains resolvable independently of the deposited object, and has resources to persist long-term (e.g., NIH and EBI databases, DDBJ, ICPSR, the NIH GREI generalist repositories, Software Heritage). Purely local or unmanaged resources — departmental servers, S3/GCS accounts, Google Drive, Box — are excluded.
Gating threshold — 0.a must score 2 (because a dataset not findable in a sustainable repository cannot be retrieved or evaluated at all). N/A is not permitted.
Publisher is unmanaged storage (excluded from 'sustainable repository' by the rubric glossary)
✗ No
Publisher found in re3data
— not checkedquery: None
Automated estimate1
PID present (scheme: ARK)
PID present but publisher is not a recognized sustainable repository
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
0.b Accessible
Gate: must score above 0
Descriptive metadata should always be available and accessible, even if the dataset is restricted, unavailable, or de-accessioned. Ensure metadata conforms to standards such as DCAT (Data Catalog Vocabulary), Datacite schema, and/or schema.org.
Questions (2)
Is descriptive metadata available via PID or compact ID lookup independently of the dataset, i.e., even if the dataset is restricted or de-accessioned?
Does descriptive metadata conform to a standard metadata vocabulary (for example DCAT 3, Datacite schema, schema.org, and/or bioschemas.org)?
2Both are true.
1Metadata is always available but does not conform to a standard schema.
0Metadata is not accessible.
Gating threshold — must score above 0. N/A is not permitted.
schema.org / EVI here means the metadata follows a standard vocabulary
Standard vocabulary references found in metadata
✗ No
No automated estimate — this criterion needs
human judgment.
Your score:
0.c Interoperable
Gate: must score above 0
Wherever possible, provide data and metadata using formally defined specifications for digital objects.
Questions (2)
Is metadata about the dataset presented using one of the following formal interoperable specifications (e.g., JSON-LD, DataCite XML, RDF), referencing at least one standard vocabulary (such as SNOMED, LOINC, UBERON, OBI, etc.)?
Is metadata about the dataset provided using a machine-readable schema such as JSON + JSON Schema, XML + XSD, Parquet with embedded schema, or Frictionless Data?
2Formal interoperable specifications for metadata are provided.
1A machine-readable metadata schema is provided.
0Neither is provided.
Gating threshold — must score above 0. N/A is not permitted.
Metadata is JSON-LD (formal interoperable specification)
metadata is JSON-LD but references no standard vocabulary
1 machine-readable schema entities
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
0.d Reusable
Gate: must score above 0
Attach a clear and accessible data usage license that allows the responsible use of AI/ML applications or link to a Data Use Agreement (DUA).
Questions (2)
Is a data usage license (e.g., Creative Commons, public domain) or a DUA programmatically linked in the metadata (for example, via schema.org:license)?
Does the explicit text of the linked license or DUA permit (or at minimum, not prohibit) the computational reuse of the data for AI/ML applications?
2A machine-readable license or DUA is present, and it does not prohibit AI/ML use.
1A license or DUA is present, but it is not machine-readable (e.g., buried in a PDF abstract), and does not prohibit AI/ML use.
0No license/DUA is linked in the metadata, or AI/ML is explicitly prohibited.
Gating threshold — must score above 0. N/A is not permitted.
License is machine-readable (an IRI linked in the metadata, not prose)
✓ Yes
License link resolves
— not checkednetwork checks disabled
Conditions of access / DUA terms
not provided
AI/ML language in license or use terms
No specific mention of AI/ML detected
Automated estimate2
machine-readable license linked in the metadata
no AI/ML prohibition language found in license or use terms
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
1. ProvenanceGating
1.a Transparent
Gate: must score above 0
Identify data sources traceable to a reasonable ground truth, e.g., clinical data from EHR at a given hospital, clinical trials, or laboratory data.
Questions (1)
Does the metadata identify the data's source specifically enough to reach a real-world, trustworthy ground truth? E.g., a privacy-protected dataset from: a named EHR/hospital or consortium, drug trial, or named lab; for experimental labs, the ground truth should include samples, instruments, reagents, and generated dataset.
2A source is named and traceable to a ground truth, in a structured metadata field. To fully ground a source, you must specify the exact data artifact (e.g., a specific dataset) instead of just listing the providing hospital or laboratory.
1The source is named but either not to a fully grounded truth, or not in a structured metadata field.
0No source is identified.
Gating threshold — must score above 0. N/A is not permitted.
Ground-truth elements present as structured entities
samples: no
instruments: no
experiments: no
datasets_with_provenance: yes
Sample entities
0
Instrument entities
0
Experiment entities
0
Example sample entity
none found
Example instrument entity
none found
Automated estimate1
12 datasets carry provenance links
ground-truth elements missing: samples, instruments, experiments
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
1.b Traceable
Gate: must score above 0≤ 1.a
Identify important data transformation steps, with links to software, at an appropriate level of detail, using a machine-readable representation, such as W3C PROV-O or EVI, to enable step tracing.
Questions (3)
Does your documentation include the generation metadata (e.g., instrument details), the key processing steps used to create the final analytical dataset, and links to all software used in data preparation?
Have you utilized valid ontologies such as W3C PROV-O, EVI, or CPM (https://www.iso.org/standard/87714.html), etc.?
Are any known gaps in the provenance records (e.g., missing original collection circumstances or chain-of-custody breaks for retrospective data) explicitly documented and disclosed to downstream users?
2Steps are provided in a machine-readable provenance record, and it includes software links. These provenance records are either fully complete or any known chain-of-custody gaps are explicitly documented and disclosed.
1Transformation steps are documented in enough detail to trace a sequence of steps, but the record falls short on one or more of the following: it is prose rather than machine-readable; software links are absent; known provenance gaps are undisclosed.
0No transformation steps were provided.
Scoring dependency rule — 1.b ≤ 1.a. The score for 1.b (Traceable) may never exceed the score for 1.a (Transparent). A transformation pipeline cannot be fully traced if the fundamental data source is unidentified.
Gating threshold — must score above 0. N/A is not permitted.
Provenance-gap / chain-of-custody disclosure language found in the prose metadata
✗ Noa 2 requires records to be complete OR known gaps explicitly disclosed — absence of this language is fine if the record is complete
Evidence graphs (visual provenance per sub-crate)
none
Automated estimate2
13 machine-readable transformation steps
every computation links its software
completeness of the provenance record (or disclosure of known gaps) is asserted, not verified
final score is capped at 1.a's score (rule 1.b ≤ 1.a)
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
⚠ Dependency rule
1.b ≤ 1.a: this score exceeds 1.a's
score, which the rubric does not allow — the aggregate will cap it.
1.c Interpretable
Gate: must score above 0
Make software for key data transformation and analysis steps available in a sustainable repository.
Questions (1)
Is the software for key transformation/analysis steps deposited where it will persist and is citable?
2Software is archived in a sustainable repository with a PID, that is, an immutable snapshot with a persistent identifier (e.g., Zenodo, Software Heritage, Dataverse, an institutional archive; includes a GitHub release archived to such a repository). For proprietary commercial software, a URI is provided to the provider's website.
1Software is available but only in mutable code hosting with no PID (e.g., a GitHub/GitLab/BitBucket repo alone).
0Software is not available.
Gating threshold — must score above 0. N/A is not permitted.
Software entities
8
Archived with a PID (Zenodo / Software Heritage / DOI / PyPI)
0 of 8
Mutable code hosting only (GitHub and the like)
1 of 8
Provider/vendor website URI only
0 of 8 counts as 2 only for proprietary commercial software (reviewer's call)
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
1.d Key actors identified
Gate: must score above 0
Identify people and organizations responsible for obtaining and processing the data, the samples, and the subject groups from which the data were produced. Reference these parties along with other dataset metadata.
Questions (1)
Does the metadata identify the key actors? "Key actors" are the people and organizations who obtained and processed the data/samples and who recruited the participants.
2Key actors are identified with resolvable PIDs (Open Researcher and Contributor ID (ORCID) for people, Research Organization Registry (ROR) for organizations).
1Key actors are named in free text, without PIDs.
0Key actors are not identified.
Gating threshold — must score above 0. N/A is not permitted.
Key actors identified
✓ Yes
Authors identified with a PID (ORCID)
0.0%(0 of 1)
Authors with PIDs (first 10)
none
Authors named in free text only
Example Researcher
Principal investigator
not provided
Contact
not provided
Organizations (root isPartOf)
none
Organizations identified with ROR PIDs
✗ No
Automated estimate1
0 of 1 authors carry a PID
1 named in free text only
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
2. Characterization
2.a Semantics
Use comprehensive descriptive metadata for datasets, including a detailed abstract, dataset keywords, and subject-specific vocabularies (e.g., MeSH) to enable detailed search and discovery.
Questions (1)
Does the metadata support search and discovery through a substantive abstract, keywords, and subject-specific controlled-vocabulary terms?
2The abstract, keywords, and controlled-vocabulary terms are present and from recognized vocabularies (e.g., MeSH, or another subject-specific ontology/thesaurus).
1Abstract and keywords are present, but not controlled-vocabulary terms (free-text keywords only).
0No abstract, or effectively no descriptive metadata is present.
Scores the subject and discovery terms only, not the vocabulary bindings already scored in 0.c (Interoperable) and 2.c (Standards).
Dataset abstract/description
Reads from three samples aligned to the reference with bwa-mem, sorted and indexed with samtools, and jointly called with bcftools; a Snakemake run captured as an EVI RO-Crate.176 chars
Keywords
snakemake
variant calling
bwa
bcftools
genomics
Controlled-vocabulary terms present
✗ No0 MeSH terms of 0 DefinedTerms
Controlled-vocabulary terms
none
Automated estimate1
abstract and keywords present, no controlled-vocabulary terms
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
2.b Statistics
Provide domain-appropriate statistical characterizations of key features of the dataset to inform analysis planning. Ensure missing values are encoded consistently.
Questions (3)
Are descriptive statistics of key variables provided?
Are the chosen statistics domain-appropriate?
Are missing values encoded consistently?
2Missing values are encoded consistently, and descriptive statistics of key variables are provided. Where applicable, the dataset provides global record and variable counts, detailed per-variable descriptive statistics (central tendency, spread, min/max, and valid counts for numeric data; frequencies for categorical data), and strictly utilizes a single, documented convention for representing missing values (not mixed; e.g., blanks/NA/-999/0).
1Missing values are encoded consistently, but no descriptive statistics are present.
0Missing values are not consistently encoded.
Score on missing-value encoding consistency, expressed in terms appropriate to the modality (e.g., absent channels, dropped leads, corrupt or omitted slices, or gaps in a time series). For non-tabular data, per-variable descriptive statistics are not applicable (N/A). This criterion is N/A when the data has no representable missingness.
Entities carrying a summary-statistics link (hasSummaryStatistics)
nonefor non-tabular data, per-variable statistics are N/A but missing-value consistency is still scored in modality terms (absent channels, dropped leads, corrupt slices, time-series gaps); the criterion is N/A only if the data has no representable missingness at all
No automated estimate — this criterion needs
human judgment.
Your score:
2.c Standards
Gate: must score above 0
Provide a machine-readable data dictionary or schema for each dataset, linked to the dataset metadata, and referencing any relevant standards.
Questions (3)
Is each file governed by a formal open specification with a built-in or paired external schema that names and types its data elements as required for downstream usage? (e.g., RDF/JSON-LD, OME-TIFF, OME-NGFF, NWB, WFDB, DICOM, NIfTI, Parquet, CSVW, Frictionless Table Schema, JSON Schema, XSD.)
Is at least one data element bound to a standard vocabulary term through (a) a dataset-level JSON-LD sidecar, or (b) through the format's inherent mechanism? Examples of inherent mechanisms include vocabulary IRIs in the data itself (where the data is RDF/JSON-LD), DICOM Code Sequences, CSVW propertyUrl or aboutUrl, Frictionless rdfType, Define-XML NCI Thesaurus codes, OME annotations, NWB ontology links, etc.
Is the standard a true schema/model that dictates structure and content, or merely a generic file container?
2For every format class present in the deposit, a formal schema and at least one populated standard vocabulary binding — as appropriate to the datatype — have been provided.
1For every format class present in the deposit, a formal schema is available, but no standard vocabulary binding has been provided.
0No formal schema has been provided. E.g., HDF5 with no NWB or other schema layer, such as pickle, .mat, npy, or undocumented vendor binaries.
A standard vocabulary is one registered in NCBO BioPortal, EBI Ontology Lookup Service, or OBO Foundry. A standard vocabulary binding must be populated in the deposited instance (capability alone scores 1). Scoring may be done per format class, and a sidecar covering a class covers every file in it. The formats named here are examples, not an enumeration.
Gating threshold (the "Standards" gate) — must score above 0. N/A is not permitted.
1 schema entities but no standard-vocabulary binding found
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
2.d Potential Sources of Bias
Describe known sources of bias in the data and assumptions made in collecting, processing, or interpreting the data. Include any known explanations regarding missing values, including reasons for missingness, and the degree to which data represents a state vs. a control (e.g., disease state vs. healthy state).
Questions (5)
Does the metadata describe known biases and assumptions?
Does the metadata describe, where applicable, reasons for missingness?
Does the metadata describe the state-vs-control balance?
Are demographic gaps and limitations in cohort representativeness (e.g., missing geographic or socioeconomic populations) explicitly disclosed within the dataset metadata or linked documentation?
If the data is derived from clinical settings, does the documentation describe how admission patterns or site selection might introduce systematic biases?
2Substantive and specific documentation details known biases and assumptions throughout collection, processing, and interpretation, including specific justifications for missing data and case-vs-control representation where applicable.
1Bias/assumptions are addressed but only in general terms; or applicable missingness/state-control elements were omitted.
0No description of bias or assumptions is given, or only a non-substantive or generic placeholder is offered (e.g., "no known bias" with no basis).
"Substantive" means that it engages the actual data (specific bias sources, assumptions, or reasons are named); it is not a boilerplate disclaimer. "Reasons for missingness" does not apply (N/A) if the dataset has no missing values. "State-vs-control" is not applicable (N/A) if the dataset has no such structure.
Missingness explanation present and substantive-length
✗ No
Completeness statement
not provided
Limitations (rai:dataLimitations)
not provided
Demographic / cohort-representativeness language found
✗ No
Case-vs-control language found
✗ No
Clinical admission-pattern / site-selection language found
✗ Noonly applicable to clinically-derived data
Automated estimate0
no bias, missingness, limitations, or completeness description anywhere in the metadata
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
2.e Data quality
Have quality control procedures been applied? If so, provide a link to a description.
Questions (3)
Has data-type specific quality control (QC) been applied?
Is QC indicated? Does the link resolve? Is the description real and not just a placeholder?
Is a link provided to the specific protocol or software that was used?
2QC was applied and appropriate for the domain.
1QC was applied and documented, but it is incomplete for the domain.
0No QC was applied, or the description link doesn't resolve, or the description is inadequate.
Adequacy (1 vs 2) is a domain-expert determination.
Collection / QC description (rai:dataCollection)
not provided
Missing-data handling
not provided
QC language found in metadata
✗ No
Links to QC protocol/software found in the collection description
nonethe rubric asks for a link to the specific protocol or software used
Automated estimate0
no QC description and no quality-control language anywhere in the metadata
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
3. Pre-model Explainability
3.a Data documentation template
Machine-readable metadata and a linked human-readable document should support a domain-appropriate extension of the original Gebru Datasheets concept for pre-model explainability, as detailed in the present article. Healthsheets information may be a helpful supplement.
Questions (3)
Is there datasheet-style documentation of the dataset as machine-readable metadata?
Is there a linked human-readable document, such as Datasheets for Datasets (Gebru et al.)?
Does the datasheet substantively address its dimensions (composition, collection, processing, limitations/uses)?
2Both machine-readable metadata and a linked, human-readable datasheet-style document that substantively covers dataset composition, collection, processing, and known limitations/uses are provided.
1One is present but not the other, or one or both are present but are thin or partial.
0Neither is provided.
Scoring rule — score only whether the datasheet exists and covers its dimensions, and do not re-score bias documentation (2.d Potential Sources of Bias) or use-case guidance (3.b Fit for purpose) here.
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
3.b Fit for purpose
Identify appropriate and inappropriate use cases for the data set. Link to any previously published analyses using this data.
Questions (2)
Does the dataset identify the appropriate and inappropriate use cases, and are there any links to previously published analyses using it?
Were the use cases provided in a structured manner (i.e., Croissant RAI or equivalent)?
2Appropriate and inappropriate use cases are identified, ideally in structured machine-readable form (e.g., Croissant RAI recommended/prohibited-use fields), and links to prior analyses are provided where such analyses exist.
1Use cases are identified but only partially (e.g., only appropriate uses are listed, or they are provided in prose with no structured form), or links to prior analyses are missing where such analyses exist.
0No use-case guidance is provided.
"Link to prior analyses" is not applicable (N/A) if the dataset is newly released with no published analyses; its absence doesn't count against the score.
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
3.c Verifiable
Provide a mechanism for ensuring the integrity of each raw or processed dataset, such as a cryptographic hash.
Questions (1)
Is there an integrity mechanism (e.g., cryptographic hash) for each raw or processed dataset, recorded in the metadata that describes it?
2An integrity mechanism is provided in the metadata for each dataset (e.g., SHA-256 per file/dataset, or a Merkle tree/manifest of hashes covering all components), so metadata and data verifiably reference each other.
1An integrity mechanism is provided, but it is incomplete (i.e., some datasets/files are covered, others are not).
0.0%(0 of 30)counts md5/sha256 recorded on the metadata entities themselves, which is what the rubric requires (metadata and data verifiably reference each other)
Embargoed datasets excluded from the denominator
0
Example entity with a checksum
none found
Automated estimate0
no checksums found
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
4. EthicsGating
4.a Ethically acquired
Gate: must score above 0
Describe ethical data acquisition consistent with accepted principles (e.g., Belmont Report, Menlo Report, CARE Principles) sufficient to evaluate for intended use, along with a management plan.
Questions (6)
Is ethical data acquisition consistent with accepted bioethics principles sufficient to evaluate for the intended use?
Are the following elements present: IRB/REC protocol, data governance committee chair, and ethics reviewer and institution?
For Indigenous-sourced data, is CARE principles conformance demonstrated, for example, by describing the composition of the governance committee with tribal representation, etc.?
Is acquisition legitimate — obtained without theft, coercion, or deception — with required permissions/consent secured?
For retrospective data, is the basis for secondary AI/ML use (e.g., a documented IRB waiver of authorization) explicitly stated in the metadata?
For prospective data, does the documented consent language explicitly cover downstream AI/ML training and potential commercialization?
2All of the following are present: acquisition is described with the source, including a permission or consent basis that specifically covers downstream AI/ML use (for prospective data), or explicitly details the specific basis for a waiver/exemption (for retrospective data); confirmation of legitimacy of data collection is included; and the DMP/DMSP is linked.
1Acquisition is described, but it is incomplete or generic.
0No acquisition description was offered.
Gating threshold — must score above 0. N/A is not permitted.
Collection description (rai:dataCollection)
not provided
Ethics reviewers (ethicalReview)
not provided
Human subjects research
not provided
Human subjects exemption
not provided
Informed consent
not provided
At-risk populations
not provided
IRB / ethics-review references found
✗ No
Consent/ethics text mentions AI/ML or commercialization (prospective data: consent must explicitly cover downstream AI/ML use)
✗ No
Waiver / exemption / secondary-use language found (retrospective data: the basis for secondary AI/ML use must be explicit)
✗ No
Management plan (rai:dataReleaseMaintenancePlan)
not provided
Automated estimate0
no acquisition, consent, or ethics-review description anywhere in the metadata
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
4.b Ethically managed
Gate: must score above 0
Data management must align with ethical principles throughout the health AI lifecycle. Indicate privacy-protection processing, if any, sufficient to evaluate ethical status for intended use.
Questions (5)
Is data sensitivity properly described in alignment with disclosure risks?
Is data sensitivity aligned with intended use?
Is privacy protection described in a way aligned with the above?
Has a formal Privacy Impact Assessment (PIA; or equivalent risk assessment) been conducted for high-sensitivity data types (e.g., biometrics, longitudinal EHR)?
Is there a documented procedure or commitment to periodically reassess re-identification risks as new data linkage technologies emerge?
2DMP/DMSP content substantively describes management appropriate to that sensitivity. A PIA or equivalent risk assessment for high-sensitivity data exists. Alternatively, the data is documented as low-sensitivity with no privacy processing required, and that determination is justified. A documented plan for periodic reassessment of re-identification risks is detailed for sensitive data.
1Privacy protection is inadequate, and/or formal PIAs or reassessment plans are absent for sensitive data.
0There is no management description where sensitivity requires it.
Gating threshold — must score above 0. N/A is not permitted.
Confidentiality level
not provided
Personal/sensitive information statement
not provided
Data governance committee
not provided
Ethics review
not provided
Management plan content (judge adequacy for the declared sensitivity)
not provided
Privacy Impact Assessment / risk-assessment language found (required for high-sensitivity data)
✗ No
Periodic re-identification-risk reassessment language found (required for sensitive data)
✗ No
Automated estimate0
no sensitivity classification, sensitivity statement, or management description
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
4.c Ethically Disseminated
Gate: must score above 0
Specify a licensing agreement and/or data use agreement (DUA) on as open terms as ethical and sustainability considerations permit. Specify contact information for a data access committee.
Questions (3)
Assuming a license and/or DUA exists, are the permitted and prohibited uses ethically justified (i.e., "as open as possible, as closed as necessary" based on privacy/population risks)?
Where the modality warrants it, do the DUA(s) include explicit ethical prohibitions relevant to the specific data modality (e.g., prohibiting biometric synthesis, identity cloning, or discriminatory use) that go beyond minimum legal requirements?
If the dataset requires controlled access, is active information for a Data Access Committee (DAC) provided?
2The dataset is ethically cleared for fully open access under a license that clearly defines permitted and prohibited ethical boundaries. Or, for controlled-access datasets, the qualitative terms of the DUA do not unjustifiably restrict use, explicit modality-specific ethical prohibitions are clearly defined where required, and active contact information for a Data Access Committee (DAC) is present. Modality-specific prohibitions are required where the data itself can be used to synthesize, impersonate, or re-identify an individual beyond what a general re-identification prohibition addresses; for example, raw voice or speech, personal genomes, face or full-head imaging, dense longitudinal geolocation.
1The terms of access are more restrictive than ethically warranted; use boundaries are vague; access is controlled, and the DAC contact is missing. Alternatively, the DUA is well-scoped with an active DAC contact but lacks modality-specific ethical prohibitions where the modality warrants them.
0Access is controlled without any defined ethical framework or oversight mechanisms, or the dataset is openly released with no defined ethical terms.
Gating threshold — must score above 0. N/A is not permitted.
High-risk modality signals in the metadata (voice, genomes, face imaging, geolocation, …)
✓ Yesgenomics
Access-committee / dataset contact
none listedrequired (and active) where access is controlled
No automated estimate — this criterion needs
human judgment.
Your score:
4.d Secure
Gate: must score above 0
Specify security requirements for storing and accessing this data, e.g., "public", "controlled access only", etc.
Questions (1)
Are security requirements for storing and accessing the dataset specified, following standards such as HL7's Confidentiality security labels (http://terminology.hl7.org/ValueSet/v3-Confidentiality)?
2Security level metadata is provided, indicating whether or not any level of protection is required to safeguard sensitive information contained in the dataset. These security requirements are specified using a standard, such as the HL7 security labels.
1Security level is specified in prose, without reference to a standard vocabulary.
0No security level metadata is provided, or the declared security level is not enforced, meaning controlled-access data is retrievable without authorization or authorized users are blocked.
Gating threshold — must score above 0. N/A is not permitted.
Confidentiality level
not provided
Value is an HL7 v3-Confidentiality code
✗ Noprose value — compare against U/L/M/N/R/V display names
De-identification statement
not provided
Personal/sensitive information statement
not provided
Automated estimate0
no security-level metadata
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
5. Sustainability
5.a Persistent
Ensure unprocessed data is preserved in an archive that complies with privacy laws and retention guidelines, enabling future reprocessing and publishing of revised data.
Questions (1)
Is all unprocessed/raw data preserved in an archive that supports retention and future reprocessing?
2Raw data is preserved in a sustainable repository, with a persistent identifier and a retention commitment.
1Raw data is preserved but only in non-sustainable storage (no PID, or mutable hosting with no retention commitment).
0Raw data is not preserved (only processed/derived data are available).
nonea 2 requires a sustainable repository with a PID and a retention commitment
Unmanaged storage detected (excluded from 'sustainable' by the rubric glossary)
none
Datasets with a contentUrl (data parked somewhere)
86.4%(19 of 22)0 of these are remote archive links; the rest live inside the crate deposit
Automated estimate1
PID present (scheme: ARK)
19 datasets have a contentUrl but no recognized archive detected
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
5.b Domain-appropriate
Ensure single-domain raw or processed data are deposited in a FAIR, domain-appropriate specialist repository if available.
Questions (2)
For single-domain data, is it deposited in the repository the domain community treats as appropriate, where one is available for this data?
For multi-domain data, is it deposited in a widely recognized generalist repository (e.g., NIH GREI participants)?
2A specialist domain repository was used where one exists and accepts this data type, or no specialist repository exists, and a recognized generalist repository was used.
1Data is deposited in a generalist repository even though a specialist, domain-appropriate repository that accepts this type of data exists and is accepted by the community.
0Data was not deposited in any domain-appropriate repository.
A domain-appropriate repository is a sustainable repository recognized by the relevant research community as appropriate for a specific data type.
Publisher repository
None
Specialist (domain) hosts
none
Generalist repository hosts
none
Unrecognized hosts
none
Dataset files on recognized repository hosts
— no denominator
Domain hint — collection types
none
Domain hint — keywords
snakemake
variant calling
bwa
bcftools
genomics
Automated estimate0
no recognized repository host detected
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
5.c Well-governed
Implement a data governance model before data capture, including a repository that facilitates data stewardship in the future and governance that accounts for maintenance, terms, policy changes, and fairness.
Questions (2)
Is a formal Data Management Plan (DMP) or long-term governance explicitly linked (i.e., via rai:dataReleaseMaintenancePlan)?
Does the linked governance plan actively account for future repository stewardship, data maintenance, and policy changes?
2A comprehensive DMP/governance plan is linked and explicitly details long-term repository stewardship, maintenance, and policy handling.
1A DMP/governance plan is linked, but it lacks detail regarding future maintenance, policy changes, or long-term stewardship.
0No formal DMP or governance model is specified.
Governance / maintenance plan (rai:dataReleaseMaintenancePlan)
not provided
DMP link extracted from the plan
none
Governance plan present
✗ No
Stewardship / maintenance / policy language in the plan (the 1-vs-2 question)
✗ No
Data governance committee
not provided
Responsible party
not provided
Context — terms of access (conditionsOfAccess; no longer scored under 5.c)
not provided
Automated estimate0
no DMP / governance plan in the metadata
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
5.d Associated
Document project-level connections between data components and elements in a machine-readable manner.
Questions (4)
Will all data users be able to link to all data required for analysis, as an associated set that persists?
Are all declared components retrievable?
Are the associations durable (archived with the data)?
Is this information provided as machine-readable or prose?
2All necessary components are accessible and associated in persistent machine-readable form (single-modality: sole component accessible and self-contained; multimodal: components linked machine-readably in an archived record, e.g. RO-Crate, Frictionless Data).
1All components are accessible, and associations are durable, but they are only documented in human-readable prose (persists with the deposit, not machine-actionable).
0Necessary components are not all accessible, or associations are only in a non-durable location (project page/README not archived with the data), or associations are absent.
Entities carrying at least one machine-readable provenance link
56.8%(25 of 44)
Sub-crates linked from the parent and present
0 of 0
hasPart references on the root
44
Evidence graphs (machine-derived association views)
none
Automated estimate2
components associated machine-readably in the archived RO-Crate (hasPart + provenance links)
accessibility of every component not verified — downgrade if pieces are missing
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
6. Computability
6.a Standardized
≤ 2.c
Datasets adhere to established, documented standards, including metadata standards, and their adherence can be validated deterministically.
Questions (2)
Can the adherence to standards, as described in 2.c, be validated deterministically?
Does a programmatic validator (e.g., script, CLI tool, or automated pipeline) exist to deterministically verify this adherence without human intervention?
2Given a score of 2 in criterion 2.c: deterministic programmatic validation tools (e.g., OMOP DQD, pyshacl, JSON Schema validators, etc.) exist for the declared standard, and the dataset architecture allows these tools to execute successfully.
1Given a score of 1 in criterion 2.c: a formal schema is present, enabling structural validation. However, because no vocabulary bindings are populated, semantic conformance cannot be deterministically validated.
0No declared or documented standard is used. The data structures are arbitrary, proprietary, or undocumented, making programmatic validation impossible.
Scoring dependency rule — 6.a ≤ 2.c. The score for 6.a (Standardized validation) may never exceed the score for 2.c (Standards definition). Because 6.a evaluates the deterministic validation of the standards defined in 2.c, validation depth is limited by the completeness of the underlying schema and semantic bindings. A dataset cannot be fully validated (6.a = 2) if the standard itself is incomplete or missing (2.c < 2).
Populated standard vocabulary bindings (2.c evidence — without them 6.a caps at 1)
✗ No
File formats present
application/x-bam: 6
text/x-fastq: 3
application/octet-stream: 3
application/yaml: 1
text/x-fasta: 1
amb: 1
ann: 1
bwt: 1
application/x-ns-proxy-autoconfig: 1
sa: 1
text/x-vcf: 1
text/tab-separated-values: 1
image/svg+xml: 1
Automated estimate1
formal schema/standard declared, so structural validation is possible
no populated standard-vocabulary bindings — semantic conformance cannot be deterministically validated (the 1-rule)
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
⚠ Dependency rule
6.a ≤ 2.c: this score exceeds 2.c's
score, which the rubric does not allow — the aggregate will cap it.
6.b Computationally accessible
Provide a mechanism to access data either through established exchange protocols or a well-documented API.
Questions (3)
Is there a programmatic mechanism to access the data at scale (e.g., cloud object storage, standard API)?
Does the access mechanism utilize standard interoperability protocols (e.g., GA4GH DRS, HL7 FHIR APIs, other REST API endpoints, etc.)?
Are the endpoints, payloads, and authentication methods formally documented (e.g., via an OpenAPI/Swagger specification)?
2Programmatic access is provided via formally documented APIs or standardized interoperability exchange protocols (e.g., cloud-native S3/GCS Buckets, GA4GH DRS URIs, HL7 FHIR Bulk Data Access, etc.). For access-controlled data, the documentation clearly details a straightforward, executable process for obtaining access.
1Programmatic access exists, but it introduces friction for automated pipelines. The access mechanism may rely on poorly documented or non-machine-readable API documentation (e.g., a plain-text README), require custom web scraping, or use non-standard authentication that requires human-in-the-loop intervention.
0Programmatic access is absent. The data is either entirely inaccessible, restricted to manual downloads (e.g., clicking a link or submitting a web form), or requires direct human intervention to complete the transfer.
Datasets with a distribution link (contentUrl)
86.4%(19 of 22)
Datasets with a REMOTE distribution link (http/ftp/s3-style)
0.0%(0 of 22)the rest are file:// paths inside the crate or embargoed placeholders
Protocols used by distribution links
none: 19 files file = paths inside the crate; none = placeholder values such as 'Embargoed'
Distribution hosts
none
Access instructions (conditionsOfAccess)
not provided
Automated estimate0
no dataset has a remote distribution link — no programmatic access mechanism visible
⚠ Estimated from the extracted evidence only —
it may be wrong. Review the evidence above before accepting.
Your score:
6.c Portable
Maximize portability across computational resources where possible. If data use requires specific resources, provide machine-readable documentation defining these resources.
Questions (3)
Can the data be processed or utilized without being locked into a specific, proprietary, or localized compute environment?
Are software dependencies (libraries, OS versions) and hardware requirements (GPUs, memory) explicitly documented?
Is this documentation provided in a machine-executable format (e.g., containers, workflow languages) that automates environment setup?
2The data requires no specialized compute environment, or all required resources and software dependencies are defined using standard, machine-executable deployment formats. These formats include containerization (e.g., Dockerfiles, Singularity images), environment managers (e.g., Conda environment.yml), or standard workflow languages (e.g., CWL, WDL, Nextflow), which allow the execution environment to be programmatically spun up across various infrastructures (e.g., cloud, HPC, local).
1Hardware constraints and software dependencies are documented in plain text (e.g., a README.md listing required Python versions or RAM limits), but lack machine-executable formats. Consequently, a user must manually provision the hardware and install dependencies to interact with the data.
0Compute requirements and software dependencies are undocumented. Alternatively, the data is tightly coupled to a proprietary system, an unshareable internal cluster, or hardcoded local file paths, making it virtually impossible for external users to port to their own environments.
No automated estimate — this criterion needs
human judgment.
Your score:
6.d Contextualized
Include any considerations regarding splits of the data, including any information withheld at any point of data collection and processing. Provide examples of data components to aid understanding of their general structure and content.
Questions (3)
Are machine-learning data splits (train/test/validation) explicitly defined, or is the rationale for a contiguous dataset documented?
Is any withheld information (such as de-identified/redacted fields, censored variables, or blinded evaluation holdouts) clearly described?
Are representative examples or synthetic datasets provided that strictly mirror the structure of the full dataset?
2The dataset composition is fully transparent. Any predefined splits (train/test/validation) are explicitly defined, including stratification strategies or random seeds, or the dataset is explicitly documented as not requiring splits. Redacted, censored, or blinded data — including information withheld during collection or processing — is clearly detailed so users understand exactly what is missing. Critically, machine-readable example data (e.g., a synthetic set or representative subset) is provided, enabling users to write and test code against the exact schema without needing access to the full cohort.
1Context is provided but incomplete. For example, splits are mentioned, but the underlying methodology (e.g., stratification) is missing; or withheld data is noted vaguely (e.g., "some patient data removed") without specifying the exact fields affected. Alternatively, structural examples are provided in a non-actionable format (e.g., a PDF screenshot of a table rather than a machine-readable JSON/CSV snippet).
0No context is provided regarding how the data might be split for training and evaluation. No documentation details what information was withheld or redacted during processing, and no structural examples or synthetic datasets are provided to help users understand the data format.
"Splits" is not applicable (N/A) where the dataset has no defined splits; the withholding and its example components still apply. Score on the applicable elements.
Datasets with split-like names (train/test/validation/holdout)
Points and percentages cover graded criteria only;
N/A answers are excluded from the denominator (N/A is not permitted on
a gating criterion). The overall AI-readiness score is
the unweighted average of the section percentages. Gate thresholds:
0.a must score 2; every other FAIRness, Provenance, and Ethics
criterion, and 2.c (Standards), must score above 0 — a gate failure
marks the section and the overall result "Gating FAIL" (the score is
still computed). Dependency rules 1.b ≤ 1.a and 6.a ≤ 2.c are checked
as you grade. Sections not yet graded plot at 0 on the chart.