Datasheet

Variant calling on three sequenced samples

ark:59853/rocrate-variant-calling-on-three-sequenced-sampl-68f3c95

Version 1.0 License ↗ Released 2026-09-11T13:35:02.543766-04:00

snakemake variant calling bwa bcftools genomics

Datasheet Summary

Reads from three samples aligned to the reference with bwa-mem, sorted and indexed with samtools, and jointly called with bcftools; a Snakemake run captured as an EVI RO-Crate.

Dataset statistics
24.7 MBTotal size
22Datasets
13Computations
8Software

Also 1 schemas

Formats bam fastq binary yaml fasta amb ann bwt ns-proxy-autoconfig sa +3 more

Evidence graph

AI-Readiness review details ↓ · review page ↗
Substantive: 3 Partial: 10 Human review: 4 24/28 estimated
FAIRness
4/6
Provenance
5/8
Characterization
2/8
Pre-model Explainability
1/6
Ethics
0/6
Sustainability
3/8
Computability
1/6

SubstantivePartialHuman reviewNo points

16 of 48 points over the 24 criteria with a mechanical estimate; 4 await human review and are not counted either way.

Release Overview

RO-Crate ID
ark:59853/rocrate-variant-calling-on-three-sequenced-sampl-68f3c95
DOI
Not specified
Release Date
2026-09-11T13:35:02.543766-04:00
Version
1.0
Description
Reads from three samples aligned to the reference with bwa-mem, sorted and indexed with samtools, and jointly called with bcftools; a Snakemake run captured as an EVI RO-Crate.
Authors
Example Researcher
Keywords
snakemake, variant calling, bwa, bcftools, genomics

Human Subjects & Regulatory

Human subjects researchNot specified
De-identified samplesNot specified
FDA regulatedNot specified
IRB protocol IDNot specified
Institutional review boardNot specified
Human subjects exemptionsNot specified

AI-Ready Review

24/28 criteria with a mechanical estimate · 3 substantive, 10 partial, 11 absent, 4 awaiting human review

Rubric for Review of AI-readiness Evaluation Criteria, v1.8 (2026-09-10). Estimates are a mechanical application of the scoring rules to the extracted evidence; criteria calling for human judgment are left unscored rather than assumed. Generated without network checks. The review page lists every piece of evidence.

0. FAIRness · gating 4/6 pts
0.a Findable Partial
  • PID present (scheme: ARK)
  • PID present but publisher is not a recognized sustainable repository

7 evidence items collected.

0.b Accessible Human review

4 evidence items collected.

0.c Interoperable Partial
  • metadata is JSON-LD but references no standard vocabulary
  • 1 machine-readable schema entities

8 evidence items collected.

0.d Reusable Substantive
  • machine-readable license linked in the metadata
  • no AI/ML prohibition language found in license or use terms

5 evidence items collected.

1. Provenance · gating 5/8 pts
1.a Transparent Partial
  • 12 datasets carry provenance links
  • ground-truth elements missing: samples, instruments, experiments

8 evidence items collected.

1.b Traceable Substantive
  • 13 machine-readable transformation steps
  • every computation links its software
  • completeness of the provenance record (or disclosure of known gaps) is asserted, not verified
  • final score is capped at 1.a's score (rule 1.b ≤ 1.a)

6 evidence items collected.

1.c Interpretable Partial
  • 1 on mutable code hosting only
  • 7 with no link

9 evidence items collected.

1.d Key actors identified Partial
  • 0 of 1 authors carry a PID
  • 1 named in free text only

8 evidence items collected.

2. Characterization 2/8 pts
2.a Semantics Partial
  • abstract and keywords present, no controlled-vocabulary terms

4 evidence items collected.

2.b Statistics Human review

5 evidence items collected.

2.c Standards Partial
  • 1 schema entities but no standard-vocabulary binding found

5 evidence items collected.

2.d Potential Sources of Bias Absent
  • no bias, missingness, limitations, or completeness description anywhere in the metadata

9 evidence items collected.

2.e Data quality Absent
  • no QC description and no quality-control language anywhere in the metadata

4 evidence items collected.

3. Pre-model Explainability 1/6 pts
3.a Data documentation template Partial
  • human-readable datasheet linked
  • 1 of 7 machine-readable sections populated

4 evidence items collected.

3.b Fit for purpose Absent
  • no use-case guidance in the metadata

6 evidence items collected.

3.c Verifiable Absent
  • no checksums found

3 evidence items collected.

4. Ethics · gating 0/6 pts
4.a Ethically acquired Absent
  • no acquisition, consent, or ethics-review description anywhere in the metadata

10 evidence items collected.

4.b Ethically managed Absent
  • no sensitivity classification, sensitivity statement, or management description

7 evidence items collected.

4.c Ethically Disseminated Human review

6 evidence items collected.

4.d Secure Absent
  • no security-level metadata

4 evidence items collected.

5. Sustainability 3/8 pts
5.a Persistent Partial
  • PID present (scheme: ARK)
  • 19 datasets have a contentUrl but no recognized archive detected

5 evidence items collected.

5.b Domain-appropriate Absent
  • no recognized repository host detected

7 evidence items collected.

5.c Well-governed Absent
  • no DMP / governance plan in the metadata

7 evidence items collected.

5.d Associated Substantive
  • components associated machine-readably in the archived RO-Crate (hasPart + provenance links)
  • accessibility of every component not verified — downgrade if pieces are missing

4 evidence items collected.

6. Computability 1/6 pts
6.a Standardized Partial
  • formal schema/standard declared, so structural validation is possible
  • no populated standard-vocabulary bindings — semantic conformance cannot be deterministically validated (the 1-rule)

7 evidence items collected.

6.b Computationally accessible Absent
  • no dataset has a remote distribution link — no programmatic access mechanism visible

5 evidence items collected.

6.c Portable Human review

5 evidence items collected.

6.d Contextualized Absent
  • no splits, withheld-information statement, or example data anywhere in the metadata

7 evidence items collected.

Composition

Variant calling on three sequenced samples 1.8 GB 22 files13 computationsbamfastqbinary
RO-Crate ID
ark:59853/rocrate-variant-calling-on-three-sequenced-sampl-68f3c95
Description
Reads from three samples aligned to the reference with bwa-mem, sorted and indexed with samtools, and jointly called with bcftools; a Snakemake run captured as an EVI RO-Crate.
Authors
Example Researcher
Date
2026-09-11
Version
1.0
Size
1.8 GB
Keywords
snakemake, variant calling, bwa, bcftools, genomics
Provenance graph

Content summary

Files 22
Formats bam (6), fastq (3), binary (3), yaml (1), fasta (1), amb (1), ann (1), bwt (1), ns-proxy-autoconfig (1), sa (1), vcf (1), tsv (1), svg (1)
Access Available (19), No link (3)
With provenance 12 of 22
Software & instruments 8
Software 8 (snakemake (6), python (2))
Instruments 0
Inputs 10
Datasets 10
fastq (3), ann (1), amb (1), ns-proxy-autoconfig (1), fasta (1), sa (1), bwt (1), yaml (1)
Other components
Experiments 0
Computations 13
amb + ann + bwt + fasta + fastq + ns-proxy-autoconfig + sa → bam ×3
bam → binary ×3
bam → bam ×3
amb + ann + bwt + fasta + fastq + ns-proxy-autoconfig + sa → svg + tsv
bam + binary + fasta → vcf
vcf → svg
vcf → tsv
Inputs, outputs and commands are listed in the full dataset details.
Schemas 1
Other 0

View full dataset details

Distribution Information

License
Release Date
2026-09-11T13:35:02.543766-04:00
Version
1.0