ClinVar 2026-07: 951 records inside our regions changed classification since 2026-05. What changed
Sample report Support Log in
aimostı Check your file

Learn Guide

VCF, gVCF, BAM and CRAM: what each genome file can tell you

A sequenced genome usually arrives as a folder of files with three- and four-letter names. They are not copies of one another. Each step from the sequencer to the variant list keeps less than the step before, and whatever a file has dropped, no later analysis can get back from it.

Key takeaways

  • FASTQ holds the raw reads, a BAM or CRAM holds the same reads placed on the reference genome, and a VCF holds only the positions where the genome differs from that reference.
  • A plain VCF is silent everywhere else, so it cannot show whether a position matched the reference or was never read. Two real 30x plain VCFs we measured confirmed 0.16 and 0.17 percent of the bases our report examines[18].
  • A gVCF adds reference blocks, which say where the reads matched and how well. A complete one confirmed 99.0 percent of the same bases; another gVCF confirmed only 14.2 percent.
  • A BAM or CRAM keeps the evidence itself: read depth at every position, whole-gene deletions, difficult genes such as CYP2D6, and the option of calling the variants again with newer software.
  • A CRAM stores reads as differences from a reference and is 50 to 70 percent smaller than a BAM, but it cannot be decoded without the exact reference it was written against[11].

Sequencing a genome to 30x produces hundreds of millions of short reads: one public 30x run of the reference sample NA12878 holds 758 million of them, 113.7 billion bases in all[1]. Software places each read on a reference genome, and a second program compares the stacked reads with that reference and writes down where they disagree. Each stage has its own file format: FASTQ for the reads as they leave the sequencer[2], SAM, BAM or CRAM for the reads once placed[3], and VCF for the differences[4]. A gVCF is a VCF that also records where the reads matched.

Four files, three steps

The files form a ladder. FASTQ sits at the top, with every read and a quality score for every letter. A BAM or CRAM holds the same reads, each given a position on the reference genome. A variant caller then works through the stacked reads position by position and writes a VCF, which keeps only the places where the genome differs from the reference. A typical genome differs from the reference at 4.1 to 5.0 million sites[5], so a VCF of a few million lines stands for a genome of about 3.1 billion bases per set of chromosomes[6].

A human male karyotype: the chromosomes of one cell photographed under a microscope, each stained in light and dark bands, arranged in pairs from the largest to the smallest.
Figure 1. The genome every file in this guide describes, seen as chromosomes: a human male karyotype, stained so that each chromosome shows its own pattern of bands. A position in a VCF, BAM or CRAM is a coordinate on one of these, counted in bases along the reference sequence of that chromosome.Courtesy: National Human Genome Research Institute
Figure 2. Each step keeps less than the one before. Sizes are for one public 30x genome, sample NA12878 of the 1000 Genomes Project[1, 7]; the VCF figure is the range for a typical genome[5].

Four terms this guide relies on

Read
A short stretch of DNA as the sequencer read it, with a quality score for each letter. The reads of the NA12878 run average 150 letters[1].
Reference genome
A standard human genome sequence, such as GRCh38, that reads are placed against and variants are described relative to.
Depth
How many reads cover a given base, also called coverage. A 30x genome averages about thirty reads per base[8]; the gVCF records further down show positions with twelve.
Variant call
The caller's conclusion at one position: which two letters the person carries there, and how confident the caller is.

FASTQ: the raw reads

A FASTQ file is a list of reads with no position attached. Each read takes four lines: a name, the letters, a separator line, and a string of quality characters, one per letter[2]. The quality character encodes a Phred score, the sequencer's own estimate of the chance that the letter is wrong. A score of 30 means one chance in a thousand, and in the 30x 1000 Genomes data at least 91 percent of the bases scored 30 or higher[10].

FASTQ

@EXAMPLE:1:FC01:1:1101:1000:1000 1:N:0
CTTGCAGGTGTACCATTCAGGACTTCAAGGGCTCTAGGAAGCTCAT
+
FFFFFFFF:FFFFFFFFF,FFFFFFFF:FFFFFFFFFFFFF:FFFF
An invented read in the FASTQ layout. On recent Illumina instruments the quality line uses only a handful of characters, because the scores are binned into four values[11].

What FASTQ cannot say is where on the genome a read came from. That is decided in the next step, alignment, which places the reads against one particular reference assembly[3]. Every other file in the folder is built from the FASTQ, and it is also the largest.

BAM and CRAM: the reads, placed on the genome

Alignment gives every read a chromosome, a position and a record of how it lines up with the reference. SAM is the text form of the result and BAM the compressed binary form, holding exactly the same information[3]. Of the files a customer is likely to hold, this is the richest, because it still contains the evidence: at any position you can count the reads that cover it and see what each one says.

That evidence answers questions a variant list cannot. Read depth shows whether a gene was covered at all, so a report can say a region was examined instead of assuming it. A whole-gene deletion appears as a stretch where depth falls by half or to zero; in a VCF the same stretch looks exactly like one where nothing was found. Genes with near-identical neighbours are the hardest case. The pharmacogene CYP2D6 comes in whole-gene deletions, duplications and hybrids with its neighbouring pseudogene CYP2D7, and PharmVar's review of the gene notes that a sample carrying the deletion on both chromosomes may be called as two ordinary copies from sequencing files processed without a structural-variant caller[12].

Reads can also be read again with software that did not exist when the genome was sequenced. DeepVariant, for example, calls variants with a neural network trained on images of stacked reads[13]. A caller like that needs the aligned reads; a VCF made by an older caller cannot be upgraded in place.

Figure 3. A CRAM stores a read as its position plus its differences from the reference, so decoding needs the very reference it was written against. The file carries a checksum of each reference sequence so that a reader can confirm it has the right one[14, 15].

A CRAM holds the same alignments as a BAM in less space. Its founding idea was to store each read as its differences from the reference instead of as letters[16]. The format now also compresses names, positions and quality scores separately, column by column, and its author notes that this, more than the reference trick, is often the largest saving[11]. On Illumina data, CRAM 3.1 files are 50 to 70 percent smaller than the equivalent BAM[11]. The 30x CRAM of NA12878 is 15.8 GB[1].

The reference is effectively part of the file. A CRAM decoded against the wrong one does not come out slightly wrong: readers are expected to check each reference sequence's MD5 checksum and report a mismatch[15]. This is why we identify the reference a CRAM names before a Deep Read is paid for.

VCF: only the differences

A VCF is a text file, usually compressed, with a block of header lines followed by one line per position[17]. Each line names a chromosome and position, the reference letter, the alternative letters and a genotype. The genotype is written as two numbers: 0/0 for two copies of the reference, 0/1 for one of each, 1/1 for two copies of the alternative, and ./. when no call could be made[17].

VCF

#CHROM  POS       ID  REF  ALT  QUAL    FILTER  FORMAT    SAMPLE
20      10001617  .   C    A    493.77  .       GT:DP:GQ  0/1:38:99
One record of a plain VCF, simplified from the example in GATK's documentation[9]. Nothing is said about any position between this record and the next.

The format was designed as a generic way to store variants with their annotations[4], and listing differences only is what keeps it small. The cost is that silence carries two meanings. When a plain VCF has no line at a position, the person may match the reference there, or the sequencing may not have read the position well enough to make a call. The file does not say which. We have written about three ways that silence was misread on real files.

4.1 to 5.0 million

sites at which a typical genome differs from the reference[5]

3.1 billion

bases in one set of the reference chromosomes[6]

0.17%

of the bases our report examines that a real 30x plain VCF positively confirmed[18]

For a common position the first reading is far more likely, so reports usually assume it, and on a plain VCF so do we, labelled as inferred. What a plain VCF cannot support is the stronger statement that a gene was examined and found clear.

gVCF: the differences, and where the reads looked

A gVCF fills the silence. Besides the variant records it writes reference blocks: records that cover a run of positions where the reads matched the reference, with an END field marking where the run stops[17]. GATK, whose HaplotypeCaller made the format common, puts the difference in one sentence.

The key difference between a regular VCF and a GVCF is that the GVCF has records for all sites, whether there is a variant call there or not.

GATK documentation, GVCF: Genomic Variant Call Format[9]

Each block carries a genotype quality and the depth of the reads behind it. GATK merges neighbouring positions into one block only when their genotype qualities fall in the same band, which keeps the file small without hiding a weak stretch inside a strong one[9]. The alternative allele on a block is a placeholder, <NON_REF> in GATK's files and <*> in the current specification, standing for any possible alternative[17].

Figure 4. One stretch of genome as aligned reads, as a gVCF and as a plain VCF. The gVCF says which positions matched and how many reads stood behind each block. The plain VCF says nothing about any position without a variant, including the stretch that no read covered.

gVCF

#CHROM  POS    REF  ALT        INFO       FORMAT              NA12878
chr1    1      N    <NON_REF>  END=10017  GT:DP:GQ:MIN_DP:PL  0/0:0:0:0:0,0,0
chr1    10018  C    <NON_REF>  END=10018  GT:DP:GQ:MIN_DP:PL  0/0:13:4:13:0,4,312
chr1    10019  T    <NON_REF>  END=10019  GT:DP:GQ:MIN_DP:PL  0/0:12:1:12:0,1,287
The first three records of the 1000 Genomes Project's 30x gVCF for NA12878, with the empty ID, QUAL and FILTER columns left out and the rest aligned for reading[7].

A gVCF can also be incomplete. One of the two gVCFs we measured had no records at all over most of our panel, because the pipeline that wrote it skipped stretches of the genome[18]. The format makes proof of coverage possible; a given file still has to contain it.

What four real 30x files could prove

We cut four real 30x genome files, public exports from the Personal Genome Project, to the 1,054 regions our report examines, 25.5 million bases in all, and counted the bases each file positively confirms[18].

Table 1. Bases each genome file confirms in the regions our report examines
FileFormatBases confirmedRegions at least 90% confirmedVariants listed
Sequencing.com 30x, 2024gVCF99.0%1,038 of 1,05441,952
A second 30x providergVCF14.2%0 of 1,05450,015
Nebula Genomics 30x, 2025Plain VCF0.17%0 of 1,05443,261
Dante Labs 30x, 2024Plain VCF0.16%0 of 1,05441,164

Source: Aimosti measurements of 6 October 2026 on public Personal Genome Project files. A base counts as confirmed when the file gives it a called genotype, variant or reference, at a read depth of 10 or more[18].

Figure 5. The same measurement as a chart. The two plain VCFs each listed more than 41,000 variants in these regions and confirmed almost nothing else.

The two plain VCFs held what a good 30x genome should: 43,261 and 41,164 variants in these regions. What they could not do was say anything about the positions in between. The complete gVCF held a similar number of variants, 41,952, and proved 1,038 of the 1,054 regions at least 90 percent covered. The measured data page has the full numbers, including ten chip exports.

Which file answers which question

Table 2. What each file can and cannot answer
QuestionVCFgVCFBAM or CRAM
Which variants were found?YesYesYes, once called
Was this position read, and how well?NoYes, from reference blocksYes, from read depth
Was a whole gene deleted or duplicated?Only if the provider ran a structural-variant callerIndirectly, as low depth in the blocksYes, from read depth
CYP2D6, including hybrid genesSmall variants onlySmall variants onlyYes, with a caller built for it
Call the variants again with newer softwareNoNoYes
Size for one 30x genome4.1 to 5.0 million variant sites5.9 GB (NA12878)15.8 GB as CRAM (NA12878)

Source: The specifications and studies cited in the sections above[1, 5, 7, 9, 12, 13, 17].

No one file is best for every question. For most of a report the variant file does the work, and the reads settle the questions about coverage and structure. Providers' downloads differ in which of these files they include; our pages on Nebula, Dante Labs, Sequencing.com and Nucleus go through each one.

How to tell which file you have

File names are a weak guide. We have seen a file named .bam that was a CRAM inside, and a gVCF is often named like any other VCF. The first lines of the file settle it.

What the first lines of a file give away

  1. The first bytes. A BAM, once decompressed, begins with the characters BAM, a CRAM begins with CRAM, and a VCF begins with a text line such as ##fileformat=VCFv4.2[14, 15, 17].
  2. Reference blocks. Records with <NON_REF> or <*> as the alternative allele and an END= field are reference blocks, and a VCF that has them is a gVCF. GATK's gVCFs also carry ##GVCFBlock lines in the header[9, 17].
  3. The assembly. The header's ##contig lines give each chromosome's length. Chromosome 1 is 248,956,422 bases long in GRCh38 and 249,250,621 in GRCh37[19, 20], and the two assemblies give the same DNA different coordinates[21].
  4. A CRAM's reference. The @SQ lines in a BAM or CRAM header can carry an M5 checksum for each reference sequence, which identifies the exact sequence the reads were placed on[14].
  5. Or let a tool read it. Our free file check reads these lines in your browser and names the format, the assembly and the kind of file without uploading anything.

What Aimosti would (and wouldn't) show you

The report reads a VCF or gVCF for its clinical, carrier, medication, trait and ancestry panels, and each card says how many of its positions were read from the file and how many were inferred as the reference base. With a gVCF it can say a gene was examined; with a plain VCF it labels those positions as inferred. A BAM or CRAM adds the Deep Read panel, which measures read depth across the curated panel and calls CYP2D6 and whole-gene deletions from the reads. FASTQ is not used.

What we won't claim

We won't call a gene examined on the strength of a plain VCF, count a reference block with no reads behind it as a confirmed base, or decode a CRAM against any reference other than the one it names. A missing record is reported as inferred, never as a clean result.

Bottom line. The most useful folder holds a gVCF rather than a plain VCF, with the BAM or CRAM beside it. Each step down from the reads is smaller and says less, and no later step can put back what an earlier one left out.

Questions people ask

Is a gVCF the same as a VCF?

A gVCF is a VCF, with the same layout and read by the same tools. It adds reference blocks, records that cover runs of positions where the reads matched the reference, so the file says where the sequencing looked as well as what it found[9].

Can a plain VCF be converted into a gVCF?

No. A reference block records the caller's confidence that a run of positions matched the reference, and that confidence comes from the reads[9]. A plain VCF no longer holds them. A gVCF can be made again from the BAM or CRAM.

Why is a CRAM so much smaller than a BAM of the same genome?

A CRAM stores each read as its differences from the reference and compresses each kind of data, names, positions and quality scores, separately. On Illumina data that makes CRAM 3.1 files 50 to 70 percent smaller than BAM[11]. Nothing is lost unless the writer chose a lossy mode, but the file can only be decoded with the reference it names.

What does 30x mean?

It is the average number of reads covering each base. The 1000 Genomes Project's 30x genomes averaged 34x per sample, ranging from 27x to 71x[10]. An average says nothing about any one gene; the files that record depth, a gVCF, BAM or CRAM, can.

Is a 23andMe or AncestryDNA download one of these files?

No. A chip export is a text file listing the positions one genotyping chip was built to measure, between 563,320 and 955,958 rows in the ten we measured[18]. Chips read common variants well and rare ones poorly: in a UK Biobank comparison, only 16 percent of chip calls for variants rarer than 1 in 100,000 were confirmed by sequencing[22]. What that means for a scary result in your own file.

Which file does Aimosti need?

A VCF or gVCF for the full report, and a gVCF lets the report say which genes were examined. A BAM or CRAM adds the Deep Read panel. FASTQ is not used. The free file check names what you have before you pay anything.

References

  1. Run ERR3239334: 30x whole-genome sequencing of NA12878, read count, base count and CRAM size. European Nucleotide Archive, 2019. 758,135,204 reads; 113,720,280,600 bases; CRAM 15,797,182,294 bytes.
  2. Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Research, 2010. doi:10.1093/nar/gkp1137
  3. Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics, 2009. doi:10.1093/bioinformatics/btp352
  4. Danecek P, Auton A, Abecasis G, et al. The variant call format and VCFtools. Bioinformatics, 2011. doi:10.1093/bioinformatics/btr330
  5. The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature, 2015. doi:10.1038/nature15393
  6. Genome assembly GRCh38.p14. NCBI Datasets, 2022. Total sequence length 3,099,441,038 bases.
  7. 1000 Genomes 30x on GRCh38: per-sample HaplotypeCaller gVCFs. International Genome Sample Resource, 2021. NA12878.haplotypeCalls.er.raw.vcf.gz, 5,948,334,564 bytes.
  8. Sims D, Sudbery I, Ilott NE, Heger A, Ponting CP. Sequencing depth and coverage: key considerations in genomic analyses. Nature Reviews Genetics, 2014. doi:10.1038/nrg3642
  9. GVCF: Genomic Variant Call Format. GATK documentation, Broad Institute.
  10. Byrska-Bishop M, Evani US, Zhao X, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell, 2022. doi:10.1016/j.cell.2022.08.004
  11. Bonfield JK. CRAM 3.1: advances in the CRAM file format. Bioinformatics, 2022. doi:10.1093/bioinformatics/btac010
  12. Nofziger C, Turner AJ, Sangkuhl K, et al. PharmVar GeneFocus: CYP2D6. Clinical Pharmacology and Therapeutics, 2020. doi:10.1002/cpt.1643
  13. Poplin R, Chang PC, Alexander D, et al. A universal SNP and small-indel variant caller using deep neural networks. Nature Biotechnology, 2018. doi:10.1038/nbt.4235
  14. Sequence Alignment/Map Format Specification. GA4GH / hts-specs, 2025.
  15. CRAM format specification (version 3.1). GA4GH / hts-specs, 2025.
  16. Hsi-Yang Fritz M, Leinonen R, Cochrane G, Birney E. Efficient storage of high throughput DNA sequencing data using reference-based compression. Genome Research, 2011. doi:10.1101/gr.114819.110
  17. The Variant Call Format Specification, VCFv4.3. GA4GH / hts-specs, 2025.
  18. What your DNA file can actually read: measured on 14 real files. Aimosti, 2026. Measured 6 October 2026; data under CC BY 4.0.
  19. Homo sapiens chromosome 1, GRCh38.p14 Primary Assembly (NC_000001.11). NCBI Nucleotide.
  20. Homo sapiens chromosome 1, GRCh37.p13 Primary Assembly (NC_000001.10). NCBI Nucleotide.
  21. Schneider VA, Graves-Lindsay T, Howe K, et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Research, 2017. doi:10.1101/gr.213611.116
  22. Weedon MN, Jackson L, Harrison JW, et al. Use of SNP chips to detect rare pathogenic variants: retrospective, population based diagnostic evaluation. BMJ, 2021. doi:10.1136/bmj.n214

Last reviewed . Every number on this page links to the source it comes from; if one of them has moved, tell us.

Free · in your browser

Find out what your own file can read.

The free check opens your DNA file on your device and shows which format you have and what a report could read from it. The file is never uploaded.

Check your fileSee a sample report