AI-Readiness Human Review

Rubric for Review of AI-readiness Evaluation Criteria, v1.8 (2026-09-10) · generated 2026-09-15 18:34 UTC
EndoTag AP-MS Profiling of Chromatin Modifier Interactome Rewiring in MDA-MB-468 Cells Upon Paclitaxel Perturbation ark:59853/rocrate-endotag-ap-ms-in-mda-mb-468-cells-paclitaxel · None · published 2026-05-22T13:23:35.699025+00:00
Graded 0 of 28
By section
210ab score chips — hollow = not yet graded; click any chip to jump to that criterion. * = gating section.
Crate inventory — 854 entities (Dataset: 557, Experiment: 118, Sample: 117, BioChemEntity: 59, Schema: 2, Instrument: 1)
Sub-crate Location Entities Links

0. FAIRnessGating

0.a Findable

Gate: must score 2
Deposit datasets in a searchable FAIR-compliant data repository.
Questions (2)
  • Does the dataset have a resolvable PID: ARK, CSTR, DOI, IGSN, PURL, URN, HDL?
  • Does the dataset have a compact ID whose prefix is registered, for example, with identifiers.org (e.g., IDs such as ena.embl:ERP001234, geo:GSE68086, dbgap:phs000001)?
2PID or prefix-registered compact ID is present, and resolves to a specific version of the dataset, deposited in a sustainable repository.
1Only a PID or compact ID is given, but the dataset is not in a sustainable repository. Alternatively, the data was deposited in a registered repository without a resolvable PID.
0Neither is present.
A sustainable repository assigns persistent identifiers, maintains descriptive metadata that remains resolvable independently of the deposited object, and has resources to persist long-term (e.g., NIH and EBI databases, DDBJ, ICPSR, the NIH GREI generalist repositories, Software Heritage). Purely local or unmanaged resources — departmental servers, S3/GCS accounts, Google Drive, Box — are excluded.
Gating threshold — 0.a must score 2 (because a dataset not findable in a sustainable repository cannot be retrieved or evaluated at all). N/A is not permitted.
Identifier
ark:59853/rocrate-endotag-ap-ms-in-mda-mb-468-cells-paclitaxel
Persistent identifier present
✓ Yes scheme: ARK
Identifier resolves
✗ No URLError
Publisher
MassIVE
Publisher is a recognized sustainable repository
✗ No
Publisher is unmanaged storage (excluded from 'sustainable repository' by the rubric glossary)
✗ No
Publisher found in re3data
✓ Yes Mass Spectrometry Interactive Virtual Environment, CosmoHub, FILER: Functional genomics repository
Automated estimate 1
  • PID present (scheme: ARK) but it did not resolve when checked
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

0.b Accessible

Gate: must score above 0
Descriptive metadata should always be available and accessible, even if the dataset is restricted, unavailable, or de-accessioned. Ensure metadata conforms to standards such as DCAT (Data Catalog Vocabulary), Datacite schema, and/or schema.org.
Questions (2)
  • Is descriptive metadata available via PID or compact ID lookup independently of the dataset, i.e., even if the dataset is restricted or de-accessioned?
  • Does descriptive metadata conform to a standard metadata vocabulary (for example DCAT 3, Datacite schema, schema.org, and/or bioschemas.org)?
2Both are true.
1Metadata is always available but does not conform to a standard schema.
0Metadata is not accessible.
Gating threshold — must score above 0. N/A is not permitted.
Identifier used for lookup
not provided
Descriptive metadata fetchable via PID alone
— not checked no identifier
@context vocabularies of the crate metadata
schema.org / EVI here means the metadata follows a standard vocabulary
Standard vocabulary references found in metadata
✓ Yes Cellosaurus (117 refs), UniProt (57 refs), OBO Foundry PURL (2 refs)
No automated estimate — this criterion needs human judgment.
Your score:

0.c Interoperable

Gate: must score above 0
Wherever possible, provide data and metadata using formally defined specifications for digital objects.
Questions (2)
  • Is metadata about the dataset presented using one of the following formal interoperable specifications (e.g., JSON-LD, DataCite XML, RDF), referencing at least one standard vocabulary (such as SNOMED, LOINC, UBERON, OBI, etc.)?
  • Is metadata about the dataset provided using a machine-readable schema such as JSON + JSON Schema, XML + XSD, Parquet with embedded schema, or Frictionless Data?
2Formal interoperable specifications for metadata are provided.
1A machine-readable metadata schema is provided.
0Neither is provided.
Gating threshold — must score above 0. N/A is not permitted.
Metadata is JSON-LD (formal interoperable specification)
✓ Yes
@context vocabularies of the crate metadata
schema.org / EVI here means the metadata follows a standard vocabulary
Standard vocabulary references found in metadata
✓ Yes Cellosaurus (117 refs), UniProt (57 refs), OBO Foundry PURL (2 refs)
Subject terms on the root (about)
none
Machine-readable schema entities (EVI:Schema)
2
Datasets linked to a schema
2
Parquet-format datasets (embedded schema)
0
Example schema entity
summary.tsv schema
{
  "@id": "ark:59853/schema-summary-tsv-schema-20260522-092748",
  "@context": {
    "@vocab": "https://schema.org/",
    "EVI": "https://w3id.org/EVI#"
  },
  "@type": "EVI:Schema",
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "name": "summary.tsv schema",
  "description": "Inferred schema for summary.tsv, sampled from MassIVE MSV000101915 via FTPS on 2026-05-22.",
  "type": "object",
  "separator": "\t",
  "header": true,
  "required": [
    "Metadata_file",
    "Uploaded_file",
    "File_descriptor",
    "Conditions",
    "Conditions_url_params",
    "…(+4 more)"
  ],
  "properties": {
    "Metadata_file": "{…}",
    "Uploaded_file": "{…}",
    "File_descriptor": "{…}",
    "Conditions": "{…}",
    "Conditions_url_params": "{…}",
    "Bio_replicates": "{…}",
    "Bio_replicates_url_params": "{…}",
    "Tech_replicates": "{…}",
    "Tech_replicates_url_params": "{…}"
  },
  "additionalProperties": true,
  "conformsTo": [
    {
      "@id": "https://json-schema.org/draft/2020-12/schema"
    }
  ]
}
Automated estimate 2
  • metadata is JSON-LD
  • standard vocabularies referenced: Cellosaurus, UniProt, OBO Foundry PURL
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

0.d Reusable

Gate: must score above 0
Attach a clear and accessible data usage license that allows the responsible use of AI/ML applications or link to a Data Use Agreement (DUA).
Questions (2)
  • Is a data usage license (e.g., Creative Commons, public domain) or a DUA programmatically linked in the metadata (for example, via schema.org:license)?
  • Does the explicit text of the linked license or DUA permit (or at minimum, not prohibit) the computational reuse of the data for AI/ML applications?
2A machine-readable license or DUA is present, and it does not prohibit AI/ML use.
1A license or DUA is present, but it is not machine-readable (e.g., buried in a PDF abstract), and does not prohibit AI/ML use.
0No license/DUA is linked in the metadata, or AI/ML is explicitly prohibited.
Gating threshold — must score above 0. N/A is not permitted.
License
License is machine-readable (an IRI linked in the metadata, not prose)
✓ Yes
License link resolves
✓ Yes
Conditions of access / DUA terms
Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.
AI/ML language in license or use terms
No specific mention of AI/ML detected
Automated estimate 2
  • machine-readable license linked in the metadata (CC0 1.0)
  • no AI/ML prohibition language found in license or use terms
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

1. ProvenanceGating

1.a Transparent

Gate: must score above 0
Identify data sources traceable to a reasonable ground truth, e.g., clinical data from EHR at a given hospital, clinical trials, or laboratory data.
Questions (1)
  • Does the metadata identify the data's source specifically enough to reach a real-world, trustworthy ground truth? E.g., a privacy-protected dataset from: a named EHR/hospital or consortium, drug trial, or named lab; for experimental labs, the ground truth should include samples, instruments, reagents, and generated dataset.
2A source is named and traceable to a ground truth, in a structured metadata field. To fully ground a source, you must specify the exact data artifact (e.g., a specific dataset) instead of just listing the providing hospital or laboratory.
1The source is named but either not to a fully grounded truth, or not in a structured metadata field.
0No source is identified.
Gating threshold — must score above 0. N/A is not permitted.
Datasets carrying provenance links (EVI/PROV terms)
100.0% (557 of 557)
Example dataset with provenance links
exLewis009053.raw
{
  "@id": "ark:59853/dataset-msv000101915-7c441f304dddb5e2",
  "name": "exLewis009053.raw",
  "@type": [
    "prov:Entity",
    "https://w3id.org/EVI#Dataset"
  ],
  "author": "Richa Tiwari",
  "description": "RAW from MassIVE MSV000101915 at raw/APMS_RawFiles/Batch02_RawFiles/exLewis009053.raw.",
  "version": "1.0",
  "associatedPublication": "",
  "contentUrl": "ftp://massive-ftp.ucsd.edu/v13/MSV000101915/raw/APMS_RawFiles/Batch02_RawFiles/exLewis009053.raw",
  "isPartOf": [],
  "usedByComputation": [],
  "fairscapeVersion": "1.1.3",
  "prov:wasAttributedTo": [
    "Richa Tiwari"
  ],
  "additionalType": "Dataset",
  "datePublished": "2026-05-22T13:23:35.699025+00:00",
  "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "Endotagging",
    "AP-MS",
    "…(+6 more)"
  ],
  "format": "raw",
  "generatedBy": [
    {
      "@id": "ark:59853/experiment-msv000101915-morf4l1-dmso"
    }
  ],
  "derivedFrom": [
    {
      "@id": "ark:59853/sample-msv000101915-morf4l1-dmso"
    }
  ],
  "contentSize": "879 MB"
}
Ground-truth elements present as structured entities
  • samples: yes
  • instruments: yes
  • experiments: yes
  • datasets_with_provenance: yes
Sample entities
117
Instrument entities
1
Experiment entities
118
Example sample entity
MDA-MB-468 ATM-tagged + DMSO
{
  "@id": "ark:59853/sample-msv000101915-atm-dmso",
  "name": "MDA-MB-468 ATM-tagged + DMSO",
  "@type": [
    "prov:Entity",
    "https://w3id.org/EVI#Sample"
  ],
  "author": "Richa Tiwari",
  "description": "MDA-MB-468 cells with endogenous ATM EndoTag knock-in, treated with DMSO, used for AP-MS pulldown of the ATM interactome.",
  "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "Endotagging",
    "AP-MS",
    "…(+3 more)"
  ],
  "cellLineReference": {
    "@id": "https://www.cellosaurus.org/CVCL_0419"
  },
  "isPartOf": [],
  "fairscapeVersion": "1.1.3",
  "bait": {
    "@id": "ark:59853/protein-msv000101915-atm"
  },
  "usedTreatment": {
    "@id": "ark:59853/treatment-msv000101915-dmso"
  },
  "additionalProperty": [
    "{…}",
    "{…}",
    "{…}",
    "{…}",
    "{…}"
  ],
  "derivedTo": [
    {
      "@id": "ark:59853/dataset-msv000101915-6ef6e7d957f01950"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-ebb95e949d689b9d"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-7f0b86bbab8614df"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-44b97023d35b4bf7"
    }
  ]
}
Example instrument entity
Orbitrap Exploris 480
{
  "@id": "ark:59853/instrument-msv000101915-1",
  "name": "Orbitrap Exploris 480",
  "@type": [
    "prov:Entity",
    "https://w3id.org/EVI#Instrument"
  ],
  "manufacturer": "Thermo Fisher Scientific",
  "model": "Orbitrap Exploris 480",
  "description": "Orbitrap Exploris 480 mass spectrometer used for proteomics analysis in MassIVE dataset MSV000101915.",
  "usedByExperiment": [
    {
      "@id": "ark:59853/experiment-msv000101915-1"
    }
  ],
  "isPartOf": [],
  "fairscapeVersion": "1.1.3",
  "metadataType": "https://w3id.org/EVI#Instrument"
}
Automated estimate 2
  • 557 datasets carry provenance links
  • all ground-truth elements present as structured entities (samples, instruments, experiments)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

1.b Traceable

Gate: must score above 0 ≤ 1.a
Identify important data transformation steps, with links to software, at an appropriate level of detail, using a machine-readable representation, such as W3C PROV-O or EVI, to enable step tracing.
Questions (3)
  • Does your documentation include the generation metadata (e.g., instrument details), the key processing steps used to create the final analytical dataset, and links to all software used in data preparation?
  • Have you utilized valid ontologies such as W3C PROV-O, EVI, or CPM (https://www.iso.org/standard/87714.html), etc.?
  • Are any known gaps in the provenance records (e.g., missing original collection circumstances or chain-of-custody breaks for retrospective data) explicitly documented and disclosed to downstream users?
2Steps are provided in a machine-readable provenance record, and it includes software links. These provenance records are either fully complete or any known chain-of-custody gaps are explicitly documented and disclosed.
1Transformation steps are documented in enough detail to trace a sequence of steps, but the record falls short on one or more of the following: it is prose rather than machine-readable; software links are absent; known provenance gaps are undisclosed.
0No transformation steps were provided.
Scoring dependency rule — 1.b ≤ 1.a. The score for 1.b (Traceable) may never exceed the score for 1.a (Transparent). A transformation pipeline cannot be fully traced if the fundamental data source is unidentified.
Gating threshold — must score above 0. N/A is not permitted.
Transformation steps (Computation + Experiment entities)
118 0 computations, 118 experiments
Steps with inputs and outputs declared
100.0% (118 of 118)
Computations linked to software
— no denominator
Example transformation step
EndoTag AP-MS LFQ — ATM + DMSO (Paclitaxel study)
{
  "@id": "ark:59853/experiment-msv000101915-atm-dmso",
  "@type": [
    "prov:Activity",
    "https://w3id.org/EVI#Experiment"
  ],
  "name": "EndoTag AP-MS LFQ — ATM + DMSO (Paclitaxel study)",
  "experimentType": "AP-MS LFQ",
  "runBy": "Richa Tiwari",
  "datePerformed": "2026-05-22T15:24:41.220165+00:00",
  "description": "AP-MS LFQ measurement of MDA-MB-468 cells with endogenous ATM EndoTag under DMSO treatment, part of the Paclitaxel perturbation study (MSV000101915). Captured 4 replicate run(s); batches: Batch 01.",
  "usedInstrument": [
    {
      "@id": "ark:59853/instrument-msv000101915-1"
    }
  ],
  "usedSample": [
    {
      "@id": "ark:59853/sample-msv000101915-atm-dmso"
    }
  ],
  "usedTreatment": [
    {
      "@id": "ark:59853/treatment-msv000101915-dmso"
    }
  ],
  "usedStain": [],
  "isPartOf": [],
  "fairscapeVersion": "1.1.3",
  "prov:wasAssociatedWith": [
    "Richa Tiwari"
  ],
  "additionalType": "Experiment",
  "additionalProperty": [
    "{…}",
    "{…}",
    "{…}",
    "{…}",
    "{…}"
  ],
  "bait": {
    "@id": "ark:59853/protein-msv000101915-atm"
  },
  "generated": [
    {
      "@id": "ark:59853/dataset-msv000101915-6ef6e7d957f01950"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-ebb95e949d689b9d"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-7f0b86bbab8614df"
    },
    {
      "@id": "ark:59853/dataset-msv000101915-44b97023d35b4bf7"
    }
  ]
}
Provenance-gap / chain-of-custody disclosure language found in the prose metadata
✗ No a 2 requires records to be complete OR known gaps explicitly disclosed — absence of this language is fine if the record is complete
Evidence graphs (visual provenance per sub-crate)
none
Automated estimate 1
  • 118 machine-readable transformation steps
  • software linked on 0 of 0 computations — consider 2 if the gap is negligible
  • final score is capped at 1.a's score (rule 1.b ≤ 1.a)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

1.c Interpretable

Gate: must score above 0
Make software for key data transformation and analysis steps available in a sustainable repository.
Questions (1)
  • Is the software for key transformation/analysis steps deposited where it will persist and is citable?
2Software is archived in a sustainable repository with a PID, that is, an immutable snapshot with a persistent identifier (e.g., Zenodo, Software Heritage, Dataverse, an institutional archive; includes a GitHub release archived to such a repository). For proprietary commercial software, a URI is provided to the provider's website.
1Software is available but only in mutable code hosting with no PID (e.g., a GitHub/GitLab/BitBucket repo alone).
0Software is not available.
Gating threshold — must score above 0. N/A is not permitted.
Software entities
0
Archived with a PID (Zenodo / Software Heritage / DOI / PyPI)
0 of 0
Mutable code hosting only (GitHub and the like)
0 of 0
Provider/vendor website URI only
0 of 0 counts as 2 only for proprietary commercial software (reviewer's call)
No download/repository link
0 of 0
Archived software
none
Code-hosted software
none
Provider-site software
none
Software without links
none
Automated estimate 0
  • no software entities in the crate
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

1.d Key actors identified

Gate: must score above 0
Identify people and organizations responsible for obtaining and processing the data, the samples, and the subject groups from which the data were produced. Reference these parties along with other dataset metadata.
Questions (1)
  • Does the metadata identify the key actors? "Key actors" are the people and organizations who obtained and processed the data/samples and who recruited the participants.
2Key actors are identified with resolvable PIDs (Open Researcher and Contributor ID (ORCID) for people, Research Organization Registry (ROR) for organizations).
1Key actors are named in free text, without PIDs.
0Key actors are not identified.
Gating threshold — must score above 0. N/A is not permitted.
Key actors identified
✓ Yes
Authors identified with a PID (ORCID)
0.0% (0 of 1)
Authors with PIDs (first 10)
none
Authors named in free text only
Richa Tiwari
Principal investigator
Trey Ideker
Contact
tideker@health.ucsd.edu
Organizations (root isPartOf)
  • ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,
  • ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,
  • ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release
Organizations identified with ROR PIDs
✗ No
Automated estimate 1
  • 0 of 1 authors carry a PID
  • 1 named in free text only
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

2. Characterization

2.a Semantics

Use comprehensive descriptive metadata for datasets, including a detailed abstract, dataset keywords, and subject-specific vocabularies (e.g., MeSH) to enable detailed search and discovery.
Questions (1)
  • Does the metadata support search and discovery through a substantive abstract, keywords, and subject-specific controlled-vocabulary terms?
2The abstract, keywords, and controlled-vocabulary terms are present and from recognized vocabularies (e.g., MeSH, or another subject-specific ontology/thesaurus).
1Abstract and keywords are present, but not controlled-vocabulary terms (free-text keywords only).
0No abstract, or effectively no descriptive metadata is present.
Scores the subject and discovery terms only, not the vocabulary bindings already scored in 0.c (Interoperable) and 2.c (Standards).
Dataset abstract/description
This submission comprises affinity purification mass spectrometry (AP-MS) raw data generated from endogenously tagged MDA-MB-468 cells under Paclitaxel perturbation. The dataset focuses on drug-induced interactome rewiring of chromatin… [596 chars — expand]This submission comprises affinity purification mass spectrometry (AP-MS) raw data generated from endogenously tagged MDA-MB-468 cells under Paclitaxel perturbation. The dataset focuses on drug-induced interactome rewiring of chromatin modifiers relevant to the triple-negative breast cancer context. All AP-MS experiments were performed in four biological replicates. Each AP-MS batch includes an untagged parental control, 10 chromatin modifier tagged cell lines, and a positive control tagged line. DMSO-treated samples serve as vehicle control across all AP-MS experiments (Unpublished data).
596 chars
Keywords
  • proteomics
  • mass spectrometry
  • MassIVE
  • Endotagging
  • AP-MS
  • Triple Negative Breast Cancer
  • MDA-MB-468
  • Paclitaxel
  • CM4AI
  • Bridge2AI
Controlled-vocabulary terms present
✗ No 0 MeSH terms of 0 DefinedTerms
Controlled-vocabulary terms
none
Automated estimate 1
  • abstract and keywords present, no controlled-vocabulary terms
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

2.b Statistics

Provide domain-appropriate statistical characterizations of key features of the dataset to inform analysis planning. Ensure missing values are encoded consistently.
Questions (3)
  • Are descriptive statistics of key variables provided?
  • Are the chosen statistics domain-appropriate?
  • Are missing values encoded consistently?
2Missing values are encoded consistently, and descriptive statistics of key variables are provided. Where applicable, the dataset provides global record and variable counts, detailed per-variable descriptive statistics (central tendency, spread, min/max, and valid counts for numeric data; frequencies for categorical data), and strictly utilizes a single, documented convention for representing missing values (not mixed; e.g., blanks/NA/-999/0).
1Missing values are encoded consistently, but no descriptive statistics are present.
0Missing values are not consistently encoded.
Score on missing-value encoding consistency, expressed in terms appropriate to the modality (e.g., absent channels, dropped leads, corrupt or omitted slices, or gaps in a time series). For non-tabular data, per-variable descriptive statistics are not applicable (N/A). This criterion is N/A when the data has no representable missingness.
Entities carrying a summary-statistics link (hasSummaryStatistics)
2
Entities located
  • Dataset: summary.tsv -> ark:59853/dataset-msv000101915-183e739aaf87928c/summary-stats
  • Dataset: APMS_Paclitaxel_Sample_Annotation.tsv -> ark:59853/dataset-msv000101915-a8731f7640588a0b/summary-stats
Example summary-statistics reference
summary.tsv
{
  "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c",
  "name": "summary.tsv",
  "@type": [
    "prov:Entity",
    "https://w3id.org/EVI#Dataset"
  ],
  "author": "Richa Tiwari",
  "description": "TSV from MassIVE MSV000101915 at ccms_metadata/summary.tsv.",
  "version": "1.0",
  "associatedPublication": "",
  "contentUrl": "ftp://massive-ftp.ucsd.edu/v13/MSV000101915/ccms_metadata/summary.tsv",
  "isPartOf": [],
  "usedByComputation": [],
  "fairscapeVersion": "1.1.3",
  "prov:wasGeneratedBy": [
    {
      "@id": "ark:59853/experiment-msv000101915-unannotated"
    }
  ],
  "prov:wasDerivedFrom": [],
  "prov:wasAttributedTo": [
    "Richa Tiwari"
  ],
  "additionalType": "Dataset",
  "datePublished": "2026-05-22T13:23:35.699025+00:00",
  "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "Endotagging",
    "AP-MS",
    "…(+6 more)"
  ],
  "format": "tsv",
  "generatedBy": [
    {
      "@id": "ark:59853/experiment-msv000101915-unannotated"
    }
  ],
  "derivedFrom": [],
  "evi:Schema": {
    "@id": "ark:59853/schema-summary-tsv-schema-20260522-092748"
  },
  "contentSize": "577 B",
  "sha256": "5ad464dcab0ec7567d361e07440ac4c9e27fc7e609d0e306532a13feb0af39ff",
  "rowCount": 1,
  "columnCount": 9,
  "sampleSize": 1,
  "hasSummaryStatistics": {
    "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c/summary-stats"
  },
  "derivedTo": [
    {
      "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c/summary-stats"
    }
  ]
}
Missing-data statement (rai:dataCollectionMissingData)
Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps… [265 chars — expand]Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.
Tabular formats in the crate
tsv: 2 for non-tabular data, per-variable statistics are N/A but missing-value consistency is still scored in modality terms (absent channels, dropped leads, corrupt slices, time-series gaps); the criterion is N/A only if the data has no representable missingness at all
Automated estimate 2
  • 2 entities carry a summary-statistics link
  • missing-data statement present (its consistency and domain-appropriateness are asserted, not verified against the files)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

2.c Standards

Gate: must score above 0
Provide a machine-readable data dictionary or schema for each dataset, linked to the dataset metadata, and referencing any relevant standards.
Questions (3)
  • Is each file governed by a formal open specification with a built-in or paired external schema that names and types its data elements as required for downstream usage? (e.g., RDF/JSON-LD, OME-TIFF, OME-NGFF, NWB, WFDB, DICOM, NIfTI, Parquet, CSVW, Frictionless Table Schema, JSON Schema, XSD.)
  • Is at least one data element bound to a standard vocabulary term through (a) a dataset-level JSON-LD sidecar, or (b) through the format's inherent mechanism? Examples of inherent mechanisms include vocabulary IRIs in the data itself (where the data is RDF/JSON-LD), DICOM Code Sequences, CSVW propertyUrl or aboutUrl, Frictionless rdfType, Define-XML NCI Thesaurus codes, OME annotations, NWB ontology links, etc.
  • Is the standard a true schema/model that dictates structure and content, or merely a generic file container?
2For every format class present in the deposit, a formal schema and at least one populated standard vocabulary binding — as appropriate to the datatype — have been provided.
1For every format class present in the deposit, a formal schema is available, but no standard vocabulary binding has been provided.
0No formal schema has been provided. E.g., HDF5 with no NWB or other schema layer, such as pickle, .mat, npy, or undocumented vendor binaries.
A standard vocabulary is one registered in NCBO BioPortal, EBI Ontology Lookup Service, or OBO Foundry. A standard vocabulary binding must be populated in the deposited instance (capability alone scores 1). Scoring may be done per format class, and a sidecar covering a class covers every file in it. The formats named here are examples, not an enumeration.
Gating threshold (the "Standards" gate) — must score above 0. N/A is not permitted.
Machine-readable schema entities (EVI:Schema, JSON Schema dialect)
2
Non-image datasets linked to a schema
0.4% (2 of 557) 0 image datasets excluded — image files don't take a data dictionary; a schema covering a format class covers every file in that class
Standard vocabulary bindings populated
✓ Yes Cellosaurus (117), UniProt (57), OBO Foundry PURL (2)
File formats present (format-class view)
  • raw: 549
  • tsv: 2
  • application/json: 2
  • pdf: 1
  • txt: 1
  • fasta: 1
  • xml: 1
Example schema entity
summary.tsv schema
{
  "@id": "ark:59853/schema-summary-tsv-schema-20260522-092748",
  "@context": {
    "@vocab": "https://schema.org/",
    "EVI": "https://w3id.org/EVI#"
  },
  "@type": "EVI:Schema",
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "name": "summary.tsv schema",
  "description": "Inferred schema for summary.tsv, sampled from MassIVE MSV000101915 via FTPS on 2026-05-22.",
  "type": "object",
  "separator": "\t",
  "header": true,
  "required": [
    "Metadata_file",
    "Uploaded_file",
    "File_descriptor",
    "Conditions",
    "Conditions_url_params",
    "…(+4 more)"
  ],
  "properties": {
    "Metadata_file": "{…}",
    "Uploaded_file": "{…}",
    "File_descriptor": "{…}",
    "Conditions": "{…}",
    "Conditions_url_params": "{…}",
    "Bio_replicates": "{…}",
    "Bio_replicates_url_params": "{…}",
    "Tech_replicates": "{…}",
    "Tech_replicates_url_params": "{…}"
  },
  "additionalProperties": true,
  "conformsTo": [
    {
      "@id": "https://json-schema.org/draft/2020-12/schema"
    }
  ]
}
Automated estimate 2
  • 2 schema entities
  • standard vocabulary bindings populated: Cellosaurus, UniProt, OBO Foundry PURL
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

2.d Potential Sources of Bias

Describe known sources of bias in the data and assumptions made in collecting, processing, or interpreting the data. Include any known explanations regarding missing values, including reasons for missingness, and the degree to which data represents a state vs. a control (e.g., disease state vs. healthy state).
Questions (5)
  • Does the metadata describe known biases and assumptions?
  • Does the metadata describe, where applicable, reasons for missingness?
  • Does the metadata describe the state-vs-control balance?
  • Are demographic gaps and limitations in cohort representativeness (e.g., missing geographic or socioeconomic populations) explicitly disclosed within the dataset metadata or linked documentation?
  • If the data is derived from clinical settings, does the documentation describe how admission patterns or site selection might introduce systematic biases?
2Substantive and specific documentation details known biases and assumptions throughout collection, processing, and interpretation, including specific justifications for missing data and case-vs-control representation where applicable.
1Bias/assumptions are addressed but only in general terms; or applicable missingness/state-control elements were omitted.
0No description of bias or assumptions is given, or only a non-substantive or generic placeholder is offered (e.g., "no known bias" with no basis).
"Substantive" means that it engages the actual data (specific bias sources, assumptions, or reasons are named); it is not a boilerplate disclaimer. "Reasons for missingness" does not apply (N/A) if the dataset has no missing values. "State-vs-control" is not applicable (N/A) if the dataset has no such structure.
Bias description (rai:dataBiases)
Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.
Bias description present and substantive-length
✓ Yes
Missing-data reasons (rai:dataCollectionMissingData)
Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps… [265 chars — expand]Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.
Missingness explanation present and substantive-length
✓ Yes
Completeness statement
These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of… [317 chars — expand]These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.
Limitations (rai:dataLimitations)
This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.
Demographic / cohort-representativeness language found
✓ Yes represent
Case-vs-control language found
✗ No
Clinical admission-pattern / site-selection language found
✗ No only applicable to clinically-derived data
No automated estimate — this criterion needs human judgment.
Your score:

2.e Data quality

Have quality control procedures been applied? If so, provide a link to a description.
Questions (3)
  • Has data-type specific quality control (QC) been applied?
  • Is QC indicated? Does the link resolve? Is the description real and not just a placeholder?
  • Is a link provided to the specific protocol or software that was used?
2QC was applied and appropriate for the domain.
1QC was applied and documented, but it is incomplete for the domain.
0No QC was applied, or the description link doesn't resolve, or the description is inadequate.
Adequacy (1 vs 2) is a domain-expert determination.
Collection / QC description (rai:dataCollection)
Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi:… [347 chars — expand]Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized.
Missing-data handling
Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps… [265 chars — expand]Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.
QC language found in metadata
✗ No
Links to QC protocol/software found in the collection description
QC description link resolves: https://doi.org/10.1101/2024.05.21.589311
— not checked HTTP 429 — inconclusive
No automated estimate — this criterion needs human judgment.
Your score:

3. Pre-model Explainability

3.a Data documentation template

Machine-readable metadata and a linked human-readable document should support a domain-appropriate extension of the original Gebru Datasheets concept for pre-model explainability, as detailed in the present article. Healthsheets information may be a helpful supplement.
Questions (3)
  • Is there datasheet-style documentation of the dataset as machine-readable metadata?
  • Is there a linked human-readable document, such as Datasheets for Datasets (Gebru et al.)?
  • Does the datasheet substantively address its dimensions (composition, collection, processing, limitations/uses)?
2Both machine-readable metadata and a linked, human-readable datasheet-style document that substantively covers dataset composition, collection, processing, and known limitations/uses are provided.
1One is present but not the other, or one or both are present but are thin or partial.
0Neither is provided.
Scoring rule — score only whether the datasheet exists and covers its dimensions, and do not re-score bias documentation (2.d Potential Sources of Bias) or use-case guidance (3.b Fit for purpose) here.
Human-readable datasheet(s)
none
Machine-readable datasheet sections populated
7 of 7
Missing sections
none
Section — Collection
Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi:… [347 chars — expand]Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized.
Section — Use cases
AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or… [438 chars — expand]AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).
Section — Limitations
This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.
Section — Biases
Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.
Section — Release/maintenance plan
Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.
Section — Conditions of access
Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.
Automated estimate 1
  • no human-readable datasheet
  • 7 of 7 machine-readable sections populated
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

3.b Fit for purpose

Identify appropriate and inappropriate use cases for the data set. Link to any previously published analyses using this data.
Questions (2)
  • Does the dataset identify the appropriate and inappropriate use cases, and are there any links to previously published analyses using it?
  • Were the use cases provided in a structured manner (i.e., Croissant RAI or equivalent)?
2Appropriate and inappropriate use cases are identified, ideally in structured machine-readable form (e.g., Croissant RAI recommended/prohibited-use fields), and links to prior analyses are provided where such analyses exist.
1Use cases are identified but only partially (e.g., only appropriate uses are listed, or they are provided in prose with no structured form), or links to prior analyses are missing where such analyses exist.
0No use-case guidance is provided.
"Link to prior analyses" is not applicable (N/A) if the dataset is newly released with no published analyses; its absence doesn't count against the score.
Appropriate use cases (rai:dataUseCases)
AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or… [438 chars — expand]AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).
Limitations / inappropriate uses (rai:dataLimitations)
This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.
Prohibited uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.
Both appropriate and inappropriate uses stated
✓ Yes
Prior publications listed
1
Prior publications
Automated estimate 2
  • appropriate and inappropriate uses both stated
  • 1 prior publications linked
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

3.c Verifiable

Provide a mechanism for ensuring the integrity of each raw or processed dataset, such as a cryptographic hash.
Questions (1)
  • Is there an integrity mechanism (e.g., cryptographic hash) for each raw or processed dataset, recorded in the metadata that describes it?
2An integrity mechanism is provided in the metadata for each dataset (e.g., SHA-256 per file/dataset, or a Merkle tree/manifest of hashes covering all components), so metadata and data verifiably reference each other.
1An integrity mechanism is provided, but it is incomplete (i.e., some datasets/files are covered, others are not).
0No integrity mechanism is provided.
Hash coverage (datasets + software, embargoed excluded)
0.4% (2 of 557) counts md5/sha256 recorded on the metadata entities themselves, which is what the rubric requires (metadata and data verifiably reference each other)
Embargoed datasets excluded from the denominator
0
Example entity with a checksum
summary.tsv
{
  "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c",
  "name": "summary.tsv",
  "@type": [
    "prov:Entity",
    "https://w3id.org/EVI#Dataset"
  ],
  "author": "Richa Tiwari",
  "description": "TSV from MassIVE MSV000101915 at ccms_metadata/summary.tsv.",
  "version": "1.0",
  "associatedPublication": "",
  "contentUrl": "ftp://massive-ftp.ucsd.edu/v13/MSV000101915/ccms_metadata/summary.tsv",
  "isPartOf": [],
  "usedByComputation": [],
  "fairscapeVersion": "1.1.3",
  "prov:wasGeneratedBy": [
    {
      "@id": "ark:59853/experiment-msv000101915-unannotated"
    }
  ],
  "prov:wasDerivedFrom": [],
  "prov:wasAttributedTo": [
    "Richa Tiwari"
  ],
  "additionalType": "Dataset",
  "datePublished": "2026-05-22T13:23:35.699025+00:00",
  "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "Endotagging",
    "AP-MS",
    "…(+6 more)"
  ],
  "format": "tsv",
  "generatedBy": [
    {
      "@id": "ark:59853/experiment-msv000101915-unannotated"
    }
  ],
  "derivedFrom": [],
  "evi:Schema": {
    "@id": "ark:59853/schema-summary-tsv-schema-20260522-092748"
  },
  "contentSize": "577 B",
  "sha256": "5ad464dcab0ec7567d361e07440ac4c9e27fc7e609d0e306532a13feb0af39ff",
  "rowCount": 1,
  "columnCount": 9,
  "sampleSize": 1,
  "hasSummaryStatistics": {
    "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c/summary-stats"
  },
  "derivedTo": [
    {
      "@id": "ark:59853/dataset-msv000101915-183e739aaf87928c/summary-stats"
    }
  ]
}
Automated estimate 1
  • checksums on 2 of 557 entities
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

4. EthicsGating

4.a Ethically acquired

Gate: must score above 0
Describe ethical data acquisition consistent with accepted principles (e.g., Belmont Report, Menlo Report, CARE Principles) sufficient to evaluate for intended use, along with a management plan.
Questions (6)
  • Is ethical data acquisition consistent with accepted bioethics principles sufficient to evaluate for the intended use?
  • Are the following elements present: IRB/REC protocol, data governance committee chair, and ethics reviewer and institution?
  • For Indigenous-sourced data, is CARE principles conformance demonstrated, for example, by describing the composition of the governance committee with tribal representation, etc.?
  • Is acquisition legitimate — obtained without theft, coercion, or deception — with required permissions/consent secured?
  • For retrospective data, is the basis for secondary AI/ML use (e.g., a documented IRB waiver of authorization) explicitly stated in the metadata?
  • For prospective data, does the documented consent language explicitly cover downstream AI/ML training and potential commercialization?
2All of the following are present: acquisition is described with the source, including a permission or consent basis that specifically covers downstream AI/ML use (for prospective data), or explicitly details the specific basis for a waiver/exemption (for retrospective data); confirmation of legitimacy of data collection is included; and the DMP/DMSP is linked.
1Acquisition is described, but it is incomplete or generic.
0No acquisition description was offered.
Gating threshold — must score above 0. N/A is not permitted.
Collection description (rai:dataCollection)
Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi:… [347 chars — expand]Data collection processes are generally described in Clark T et al. (2024) "Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized.
Ethics reviewers (ethicalReview)
Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.
Human subjects research
None - data collected from commercially available cell lines
Human subjects exemption
Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.
Informed consent
not provided
At-risk populations
not provided
IRB / ethics-review references found
✗ No
Consent/ethics text mentions AI/ML or commercialization (prospective data: consent must explicitly cover downstream AI/ML use)
✓ Yes AI, Artificial Intelligence
Waiver / exemption / secondary-use language found (retrospective data: the basis for secondary AI/ML use must be explicit)
✓ Yes exempt
Management plan (rai:dataReleaseMaintenancePlan)
Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.
DMP link resolves: https://dataverse.lib.virginia.edu/
✓ Yes
No automated estimate — this criterion needs human judgment.
Your score:

4.b Ethically managed

Gate: must score above 0
Data management must align with ethical principles throughout the health AI lifecycle. Indicate privacy-protection processing, if any, sufficient to evaluate ethical status for intended use.
Questions (5)
  • Is data sensitivity properly described in alignment with disclosure risks?
  • Is data sensitivity aligned with intended use?
  • Is privacy protection described in a way aligned with the above?
  • Has a formal Privacy Impact Assessment (PIA; or equivalent risk assessment) been conducted for high-sensitivity data types (e.g., biometrics, longitudinal EHR)?
  • Is there a documented procedure or commitment to periodically reassess re-identification risks as new data linkage technologies emerge?
2DMP/DMSP content substantively describes management appropriate to that sensitivity. A PIA or equivalent risk assessment for high-sensitivity data exists. Alternatively, the data is documented as low-sensitivity with no privacy processing required, and that determination is justified. A documented plan for periodic reassessment of re-identification risks is detailed for sensitive data.
1Privacy protection is inadequate, and/or formal PIAs or reassessment plans are absent for sensitive data.
0There is no management description where sensitivity requires it.
Gating threshold — must score above 0. N/A is not permitted.
Confidentiality level
Unrestricted
Personal/sensitive information statement
not provided
Data governance committee
Jilian Parker
Ethics review
Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.
Management plan content (judge adequacy for the declared sensitivity)
Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.
Privacy Impact Assessment / risk-assessment language found (required for high-sensitivity data)
✗ No
Periodic re-identification-risk reassessment language found (required for sensitive data)
✗ No
No automated estimate — this criterion needs human judgment.
Your score:

4.c Ethically Disseminated

Gate: must score above 0
Specify a licensing agreement and/or data use agreement (DUA) on as open terms as ethical and sustainability considerations permit. Specify contact information for a data access committee.
Questions (3)
  • Assuming a license and/or DUA exists, are the permitted and prohibited uses ethically justified (i.e., "as open as possible, as closed as necessary" based on privacy/population risks)?
  • Where the modality warrants it, do the DUA(s) include explicit ethical prohibitions relevant to the specific data modality (e.g., prohibiting biometric synthesis, identity cloning, or discriminatory use) that go beyond minimum legal requirements?
  • If the dataset requires controlled access, is active information for a Data Access Committee (DAC) provided?
2The dataset is ethically cleared for fully open access under a license that clearly defines permitted and prohibited ethical boundaries. Or, for controlled-access datasets, the qualitative terms of the DUA do not unjustifiably restrict use, explicit modality-specific ethical prohibitions are clearly defined where required, and active contact information for a Data Access Committee (DAC) is present. Modality-specific prohibitions are required where the data itself can be used to synthesize, impersonate, or re-identify an individual beyond what a general re-identification prohibition addresses; for example, raw voice or speech, personal genomes, face or full-head imaging, dense longitudinal geolocation.
1The terms of access are more restrictive than ethically warranted; use boundaries are vague; access is controlled, and the DAC contact is missing. Alternatively, the DUA is well-scoped with an active DAC contact but lacks modality-specific ethical prohibitions where the modality warrants them.
0Access is controlled without any defined ethical framework or oversight mechanisms, or the dataset is openly released with no defined ethical terms.
Gating threshold — must score above 0. N/A is not permitted.
Conditions of access / DUA
Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.
Permitted uses stated
✓ Yes AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable …[truncated, 438 chars total]
Prohibited uses stated
✓ Yes These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.
High-risk modality signals in the metadata (voice, genomes, face imaging, geolocation, …)
✗ No where such a modality is present, a 2 requires explicit modality-specific ethical prohibitions (e.g., no biometric synthesis, identity cloning, discriminatory use)
Access-committee / dataset contact
tideker@health.ucsd.edu required (and active) where access is controlled
No automated estimate — this criterion needs human judgment.
Your score:

4.d Secure

Gate: must score above 0
Specify security requirements for storing and accessing this data, e.g., "public", "controlled access only", etc.
Questions (1)
  • Are security requirements for storing and accessing the dataset specified, following standards such as HL7's Confidentiality security labels (http://terminology.hl7.org/ValueSet/v3-Confidentiality)?
2Security level metadata is provided, indicating whether or not any level of protection is required to safeguard sensitive information contained in the dataset. These security requirements are specified using a standard, such as the HL7 security labels.
1Security level is specified in prose, without reference to a standard vocabulary.
0No security level metadata is provided, or the declared security level is not enforced, meaning controlled-access data is retrievable without authorization or authorized users are blocked.
Gating threshold — must score above 0. N/A is not permitted.
Confidentiality level
Unrestricted
Value is an HL7 v3-Confidentiality code
De-identification statement
not provided
Personal/sensitive information statement
not provided
Automated estimate 2
  • confidentiality level is an HL7 v3-Confidentiality code ('U')
  • whether the declared level is actually ENFORCED (the 0-rule) is not verified — score 0 if controlled data is retrievable without authorization or authorized users are blocked
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

5. Sustainability

5.a Persistent

Ensure unprocessed data is preserved in an archive that complies with privacy laws and retention guidelines, enabling future reprocessing and publishing of revised data.
Questions (1)
  • Is all unprocessed/raw data preserved in an archive that supports retention and future reprocessing?
2Raw data is preserved in a sustainable repository, with a persistent identifier and a retention commitment.
1Raw data is preserved but only in non-sustainable storage (no PID, or mutable hosting with no retention commitment).
0Raw data is not preserved (only processed/derived data are available).
Identifier
not provided
Identifier is a PID
✓ Yes scheme: ARK
Recognized archives detected (publisher + content hosts)
MassIVE (proteomics) a 2 requires a sustainable repository with a PID and a retention commitment
Unmanaged storage detected (excluded from 'sustainable' by the rubric glossary)
none
Datasets with a contentUrl (data parked somewhere)
99.6% (555 of 557) 555 of these are remote archive links; the rest live inside the crate deposit
Automated estimate 2
  • PID present (scheme: ARK)
  • recognized archive(s): MassIVE (proteomics)
  • an explicit retention commitment is assumed for recognized repositories, not verified
  • note the criterion covers RAW/unprocessed data — confirm the raw data itself is what's archived
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

5.b Domain-appropriate

Ensure single-domain raw or processed data are deposited in a FAIR, domain-appropriate specialist repository if available.
Questions (2)
  • For single-domain data, is it deposited in the repository the domain community treats as appropriate, where one is available for this data?
  • For multi-domain data, is it deposited in a widely recognized generalist repository (e.g., NIH GREI participants)?
2A specialist domain repository was used where one exists and accepts this data type, or no specialist repository exists, and a recognized generalist repository was used.
1Data is deposited in a generalist repository even though a specialist, domain-appropriate repository that accepts this type of data exists and is accepted by the community.
0Data was not deposited in any domain-appropriate repository.
A domain-appropriate repository is a sustainable repository recognized by the relevant research community as appropriate for a specific data type.
Publisher repository
MassIVE
Specialist (domain) hosts
ftp://massive-ftp.ucsd.edu — MassIVE (proteomics): 555 files
Generalist repository hosts
none
Unrecognized hosts
none
Dataset files on recognized repository hosts
100.0% (555 of 555)
Domain hint — collection types
Perturb-seq; IF imaging; SEC-MS
Domain hint — keywords
  • proteomics
  • mass spectrometry
  • MassIVE
  • Endotagging
  • AP-MS
  • Triple Negative Breast Cancer
  • MDA-MB-468
  • Paclitaxel
  • CM4AI
  • Bridge2AI
Automated estimate 2
  • data on specialist (domain) hosts: ftp://massive-ftp.ucsd.edu — MassIVE (proteomics)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

5.c Well-governed

Implement a data governance model before data capture, including a repository that facilitates data stewardship in the future and governance that accounts for maintenance, terms, policy changes, and fairness.
Questions (2)
  • Is a formal Data Management Plan (DMP) or long-term governance explicitly linked (i.e., via rai:dataReleaseMaintenancePlan)?
  • Does the linked governance plan actively account for future repository stewardship, data maintenance, and policy changes?
2A comprehensive DMP/governance plan is linked and explicitly details long-term repository stewardship, maintenance, and policy handling.
1A DMP/governance plan is linked, but it lacks detail regarding future maintenance, policy changes, or long-term stewardship.
0No formal DMP or governance model is specified.
Governance / maintenance plan (rai:dataReleaseMaintenancePlan)
Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.
DMP link extracted from the plan
DMP link resolves
✓ Yes
Governance plan present
✓ Yes
Stewardship / maintenance / policy language in the plan (the 1-vs-2 question)
✓ Yes long term, preservation
Data governance committee
Jilian Parker
Responsible party
Jilian Parker; Trey Ideker; tideker@health.ucsd.edu
Context — terms of access (conditionsOfAccess; no longer scored under 5.c)
Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.
No automated estimate — this criterion needs human judgment.
Your score:

5.d Associated

Document project-level connections between data components and elements in a machine-readable manner.
Questions (4)
  • Will all data users be able to link to all data required for analysis, as an associated set that persists?
  • Are all declared components retrievable?
  • Are the associations durable (archived with the data)?
  • Is this information provided as machine-readable or prose?
2All necessary components are accessible and associated in persistent machine-readable form (single-modality: sole component accessible and self-contained; multimodal: components linked machine-readably in an archived record, e.g. RO-Crate, Frictionless Data).
1All components are accessible, and associations are durable, but they are only documented in human-readable prose (persists with the deposit, not machine-actionable).
0Necessary components are not all accessible, or associations are only in a non-durable location (project page/README not archived with the data), or associations are absent.
Entities carrying at least one machine-readable provenance link
92.9% (793 of 854)
Sub-crates linked from the parent and present
0 of 0
hasPart references on the root
562
Evidence graphs (machine-derived association views)
none
Automated estimate 2
  • components associated machine-readably in the archived RO-Crate (hasPart + provenance links)
  • accessibility of every component not verified — downgrade if pieces are missing
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

6. Computability

6.a Standardized

≤ 2.c
Datasets adhere to established, documented standards, including metadata standards, and their adherence can be validated deterministically.
Questions (2)
  • Can the adherence to standards, as described in 2.c, be validated deterministically?
  • Does a programmatic validator (e.g., script, CLI tool, or automated pipeline) exist to deterministically verify this adherence without human intervention?
2Given a score of 2 in criterion 2.c: deterministic programmatic validation tools (e.g., OMOP DQD, pyshacl, JSON Schema validators, etc.) exist for the declared standard, and the dataset architecture allows these tools to execute successfully.
1Given a score of 1 in criterion 2.c: a formal schema is present, enabling structural validation. However, because no vocabulary bindings are populated, semantic conformance cannot be deterministically validated.
0No declared or documented standard is used. The data structures are arbitrary, proprietary, or undocumented, making programmatic validation impossible.
Scoring dependency rule — 6.a ≤ 2.c. The score for 6.a (Standardized validation) may never exceed the score for 2.c (Standards definition). Because 6.a evaluates the deterministic validation of the standards defined in 2.c, validation depth is limited by the completeness of the underlying schema and semantic bindings. A dataset cannot be fully validated (6.a = 2) if the standard itself is incomplete or missing (2.c < 2).
conformsTo declarations (metadata descriptor + root)
Recognized standards in @context / conformsTo
  • RO-Crate
  • schema.org
Deterministic validator known for the declared standards
✓ Yes JSON Schema — any JSON Schema validator (e.g. python-jsonschema); RO-Crate profile — rocrate-validator / ro-crate-py
Machine-readable schema entities
2
Populated standard vocabulary bindings (2.c evidence — without them 6.a caps at 1)
✓ Yes Cellosaurus (117), UniProt (57), OBO Foundry PURL (2)
File formats present
  • raw: 549
  • tsv: 2
  • application/json: 2
  • pdf: 1
  • txt: 1
  • fasta: 1
  • xml: 1
Automated estimate 2
  • declared standards: RO-Crate, schema.org
  • deterministic validator known: JSON Schema — any JSON Schema validator (e.g. python-jsonschema); RO-Crate profile — rocrate-validator / ro-crate-py
  • populated vocabulary bindings present (final score is capped at 2.c's score — rule 6.a ≤ 2.c)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

6.b Computationally accessible

Provide a mechanism to access data either through established exchange protocols or a well-documented API.
Questions (3)
  • Is there a programmatic mechanism to access the data at scale (e.g., cloud object storage, standard API)?
  • Does the access mechanism utilize standard interoperability protocols (e.g., GA4GH DRS, HL7 FHIR APIs, other REST API endpoints, etc.)?
  • Are the endpoints, payloads, and authentication methods formally documented (e.g., via an OpenAPI/Swagger specification)?
2Programmatic access is provided via formally documented APIs or standardized interoperability exchange protocols (e.g., cloud-native S3/GCS Buckets, GA4GH DRS URIs, HL7 FHIR Bulk Data Access, etc.). For access-controlled data, the documentation clearly details a straightforward, executable process for obtaining access.
1Programmatic access exists, but it introduces friction for automated pipelines. The access mechanism may rely on poorly documented or non-machine-readable API documentation (e.g., a plain-text README), require custom web scraping, or use non-standard authentication that requires human-in-the-loop intervention.
0Programmatic access is absent. The data is either entirely inaccessible, restricted to manual downloads (e.g., clicking a link or submitting a web form), or requires direct human intervention to complete the transfer.
Datasets with a distribution link (contentUrl)
99.6% (555 of 557)
Datasets with a REMOTE distribution link (http/ftp/s3-style)
99.6% (555 of 557) the rest are file:// paths inside the crate or embargoed placeholders
Protocols used by distribution links
ftp: 555 files file = paths inside the crate; none = placeholder values such as 'Embargoed'
Distribution hosts
ftp://massive-ftp.ucsd.edu: 555 files
Access instructions (conditionsOfAccess)
Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.
No automated estimate — this criterion needs human judgment.
Your score:

6.c Portable

Maximize portability across computational resources where possible. If data use requires specific resources, provide machine-readable documentation defining these resources.
Questions (3)
  • Can the data be processed or utilized without being locked into a specific, proprietary, or localized compute environment?
  • Are software dependencies (libraries, OS versions) and hardware requirements (GPUs, memory) explicitly documented?
  • Is this documentation provided in a machine-executable format (e.g., containers, workflow languages) that automates environment setup?
2The data requires no specialized compute environment, or all required resources and software dependencies are defined using standard, machine-executable deployment formats. These formats include containerization (e.g., Dockerfiles, Singularity images), environment managers (e.g., Conda environment.yml), or standard workflow languages (e.g., CWL, WDL, Nextflow), which allow the execution environment to be programmatically spun up across various infrastructures (e.g., cloud, HPC, local).
1Hardware constraints and software dependencies are documented in plain text (e.g., a README.md listing required Python versions or RAM limits), but lack machine-executable formats. Consequently, a user must manually provision the hardware and install dependencies to interact with the data.
0Compute requirements and software dependencies are undocumented. Alternatively, the data is tightly coupled to a proprietary system, an unshareable internal cluster, or hardcoded local file paths, making it virtually impossible for external users to port to their own environments.
Files in open/common formats
8 tsv: 2, application/json: 2, pdf: 1, txt: 1, fasta: 1, xml: 1
Files in proprietary vendor formats
549 raw — Thermo .raw (mass spec, vendor): 549
Computations with a container declared (usedContainer)
— no denominator
Example computation (environment description)
none found
Software descriptions
none
No automated estimate — this criterion needs human judgment.
Your score:

6.d Contextualized

Include any considerations regarding splits of the data, including any information withheld at any point of data collection and processing. Provide examples of data components to aid understanding of their general structure and content.
Questions (3)
  • Are machine-learning data splits (train/test/validation) explicitly defined, or is the rationale for a contiguous dataset documented?
  • Is any withheld information (such as de-identified/redacted fields, censored variables, or blinded evaluation holdouts) clearly described?
  • Are representative examples or synthetic datasets provided that strictly mirror the structure of the full dataset?
2The dataset composition is fully transparent. Any predefined splits (train/test/validation) are explicitly defined, including stratification strategies or random seeds, or the dataset is explicitly documented as not requiring splits. Redacted, censored, or blinded data — including information withheld during collection or processing — is clearly detailed so users understand exactly what is missing. Critically, machine-readable example data (e.g., a synthetic set or representative subset) is provided, enabling users to write and test code against the exact schema without needing access to the full cohort.
1Context is provided but incomplete. For example, splits are mentioned, but the underlying methodology (e.g., stratification) is missing; or withheld data is noted vaguely (e.g., "some patient data removed") without specifying the exact fields affected. Alternatively, structural examples are provided in a non-actionable format (e.g., a PDF screenshot of a table rather than a machine-readable JSON/CSV snippet).
0No context is provided regarding how the data might be split for training and evaluation. No documentation details what information was withheld or redacted during processing, and no structural examples or synthetic datasets are provided to help users understand the data format.
"Splits" is not applicable (N/A) where the dataset has no defined splits; the withholding and its example components still apply. Score on the applicable elements.
Datasets with split-like names (train/test/validation/holdout)
✗ No
Split-named datasets
none
Sampling strategies (d4d:samplingStrategies)
not provided
Withheld / missing information
Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps… [265 chars — expand]Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.
Preprocessing protocol (rai:dataPreprocessingProtocol)
not provided
Example / synthetic dataset provided
none found
Structural example — schema entity (documents exact record structure)
summary.tsv schema
{
  "@id": "ark:59853/schema-summary-tsv-schema-20260522-092748",
  "@context": {
    "@vocab": "https://schema.org/",
    "EVI": "https://w3id.org/EVI#"
  },
  "@type": "EVI:Schema",
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "name": "summary.tsv schema",
  "description": "Inferred schema for summary.tsv, sampled from MassIVE MSV000101915 via FTPS on 2026-05-22.",
  "type": "object",
  "separator": "\t",
  "header": true,
  "required": [
    "Metadata_file",
    "Uploaded_file",
    "File_descriptor",
    "Conditions",
    "Conditions_url_params",
    "…(+4 more)"
  ],
  "properties": {
    "Metadata_file": "{…}",
    "Uploaded_file": "{…}",
    "File_descriptor": "{…}",
    "Conditions": "{…}",
    "Conditions_url_params": "{…}",
    "Bio_replicates": "{…}",
    "Bio_replicates_url_params": "{…}",
    "Tech_replicates": "{…}",
    "Tech_replicates_url_params": "{…}"
  },
  "additionalProperties": true,
  "conformsTo": [
    {
      "@id": "https://json-schema.org/draft/2020-12/schema"
    }
  ]
}
Automated estimate 1
  • some context present but incomplete (example data or withheld-information description missing)
⚠ Estimated from the extracted evidence only — it may be wrong. Review the evidence above before accepting.
Your score:

Score summary

SectionGraded Points% N/AGate
0. FAIRness Gating
1. Provenance Gating
2. Characterization Gating
3. Pre-model Explainability
4. Ethics Gating
5. Sustainability
6. Computability
Total
Points and percentages cover graded criteria only; N/A answers are excluded from the denominator (N/A is not permitted on a gating criterion). The overall AI-readiness score is the unweighted average of the section percentages. Gate thresholds: 0.a must score 2; every other FAIRness, Provenance, and Ethics criterion, and 2.c (Standards), must score above 0 — a gate failure marks the section and the overall result "Gating FAIL" (the score is still computed). Dependency rules 1.b ≤ 1.a and 6.a ≤ 2.c are checked as you grade. Sections not yet graded plot at 0 on the chart.
your scores   unverified automated estimates