Learn Guide
VCF, gVCF, BAM and CRAM: what each genome file can tell you
A sequenced genome usually arrives as a folder of files with three- and four-letter names. They are not copies of one another. Each step from the sequencer to the variant list keeps less than the step before, and whatever a file has dropped, no later analysis can get back from it.
Key takeaways
- FASTQ holds the raw reads, a BAM or CRAM holds the same reads placed on the reference genome, and a VCF holds only the positions where the genome differs from that reference.
- A plain VCF is silent everywhere else, so it cannot show whether a position matched the reference or was never read. Two real 30x plain VCFs we measured confirmed 0.16 and 0.17 percent of the bases our report examines[18].
- A gVCF adds reference blocks, which say where the reads matched and how well. A complete one confirmed 99.0 percent of the same bases; another gVCF confirmed only 14.2 percent.
- A BAM or CRAM keeps the evidence itself: read depth at every position, whole-gene deletions, difficult genes such as CYP2D6, and the option of calling the variants again with newer software.
- A CRAM stores reads as differences from a reference and is 50 to 70 percent smaller than a BAM, but it cannot be decoded without the exact reference it was written against[11].
Sequencing a genome to 30x produces hundreds of millions of short reads: one public 30x run of the reference sample NA12878 holds 758 million of them, 113.7 billion bases in all[1]. Software places each read on a reference genome, and a second program compares the stacked reads with that reference and writes down where they disagree. Each stage has its own file format: FASTQ for the reads as they leave the sequencer[2], SAM, BAM or CRAM for the reads once placed[3], and VCF for the differences[4]. A gVCF is a VCF that also records where the reads matched.
Four files, three steps
The files form a ladder. FASTQ sits at the top, with every read and a quality score for every letter. A BAM or CRAM holds the same reads, each given a position on the reference genome. A variant caller then works through the stacked reads position by position and writes a VCF, which keeps only the places where the genome differs from the reference. A typical genome differs from the reference at 4.1 to 5.0 million sites[5], so a VCF of a few million lines stands for a genome of about 3.1 billion bases per set of chromosomes[6].
Four terms this guide relies on
- Read
- A short stretch of DNA as the sequencer read it, with a quality score for each letter. The reads of the NA12878 run average 150 letters[1].
- Reference genome
- A standard human genome sequence, such as GRCh38, that reads are placed against and variants are described relative to.
- Depth
- How many reads cover a given base, also called coverage. A 30x genome averages about thirty reads per base[8]; the gVCF records further down show positions with twelve.
- Variant call
- The caller's conclusion at one position: which two letters the person carries there, and how confident the caller is.
FASTQ: the raw reads
A FASTQ file is a list of reads with no position attached. Each read takes four lines: a name, the letters, a separator line, and a string of quality characters, one per letter[2]. The quality character encodes a Phred score, the sequencer's own estimate of the chance that the letter is wrong. A score of 30 means one chance in a thousand, and in the 30x 1000 Genomes data at least 91 percent of the bases scored 30 or higher[10].
FASTQ
@EXAMPLE:1:FC01:1:1101:1000:1000 1:N:0
CTTGCAGGTGTACCATTCAGGACTTCAAGGGCTCTAGGAAGCTCAT
+
FFFFFFFF:FFFFFFFFF,FFFFFFFF:FFFFFFFFFFFFF:FFFFWhat FASTQ cannot say is where on the genome a read came from. That is decided in the next step, alignment, which places the reads against one particular reference assembly[3]. Every other file in the folder is built from the FASTQ, and it is also the largest.
BAM and CRAM: the reads, placed on the genome
Alignment gives every read a chromosome, a position and a record of how it lines up with the reference. SAM is the text form of the result and BAM the compressed binary form, holding exactly the same information[3]. Of the files a customer is likely to hold, this is the richest, because it still contains the evidence: at any position you can count the reads that cover it and see what each one says.
That evidence answers questions a variant list cannot. Read depth shows whether a gene was covered at all, so a report can say a region was examined instead of assuming it. A whole-gene deletion appears as a stretch where depth falls by half or to zero; in a VCF the same stretch looks exactly like one where nothing was found. Genes with near-identical neighbours are the hardest case. The pharmacogene CYP2D6 comes in whole-gene deletions, duplications and hybrids with its neighbouring pseudogene CYP2D7, and PharmVar's review of the gene notes that a sample carrying the deletion on both chromosomes may be called as two ordinary copies from sequencing files processed without a structural-variant caller[12].
Reads can also be read again with software that did not exist when the genome was sequenced. DeepVariant, for example, calls variants with a neural network trained on images of stacked reads[13]. A caller like that needs the aligned reads; a VCF made by an older caller cannot be upgraded in place.
A CRAM holds the same alignments as a BAM in less space. Its founding idea was to store each read as its differences from the reference instead of as letters[16]. The format now also compresses names, positions and quality scores separately, column by column, and its author notes that this, more than the reference trick, is often the largest saving[11]. On Illumina data, CRAM 3.1 files are 50 to 70 percent smaller than the equivalent BAM[11]. The 30x CRAM of NA12878 is 15.8 GB[1].
The reference is effectively part of the file. A CRAM decoded against the wrong one does not come out slightly wrong: readers are expected to check each reference sequence's MD5 checksum and report a mismatch[15]. This is why we identify the reference a CRAM names before a Deep Read is paid for.
VCF: only the differences
A VCF is a text file, usually compressed, with a block of header lines followed by one line per position[17]. Each line names a chromosome and position, the reference letter, the alternative letters and a genotype. The genotype is written as two numbers: 0/0 for two copies of the reference, 0/1 for one of each, 1/1 for two copies of the alternative, and ./. when no call could be made[17].
VCF
#CHROM POS ID REF ALT QUAL FILTER FORMAT SAMPLE
20 10001617 . C A 493.77 . GT:DP:GQ 0/1:38:99The format was designed as a generic way to store variants with their annotations[4], and listing differences only is what keeps it small. The cost is that silence carries two meanings. When a plain VCF has no line at a position, the person may match the reference there, or the sequencing may not have read the position well enough to make a call. The file does not say which. We have written about three ways that silence was misread on real files.
4.1 to 5.0 million
sites at which a typical genome differs from the reference[5]
3.1 billion
bases in one set of the reference chromosomes[6]
0.17%
of the bases our report examines that a real 30x plain VCF positively confirmed[18]
For a common position the first reading is far more likely, so reports usually assume it, and on a plain VCF so do we, labelled as inferred. What a plain VCF cannot support is the stronger statement that a gene was examined and found clear.
gVCF: the differences, and where the reads looked
A gVCF fills the silence. Besides the variant records it writes reference blocks: records that cover a run of positions where the reads matched the reference, with an END field marking where the run stops[17]. GATK, whose HaplotypeCaller made the format common, puts the difference in one sentence.
The key difference between a regular VCF and a GVCF is that the GVCF has records for all sites, whether there is a variant call there or not.
Each block carries a genotype quality and the depth of the reads behind it. GATK merges neighbouring positions into one block only when their genotype qualities fall in the same band, which keeps the file small without hiding a weak stretch inside a strong one[9]. The alternative allele on a block is a placeholder, <NON_REF> in GATK's files and <*> in the current specification, standing for any possible alternative[17].
gVCF
#CHROM POS REF ALT INFO FORMAT NA12878
chr1 1 N <NON_REF> END=10017 GT:DP:GQ:MIN_DP:PL 0/0:0:0:0:0,0,0
chr1 10018 C <NON_REF> END=10018 GT:DP:GQ:MIN_DP:PL 0/0:13:4:13:0,4,312
chr1 10019 T <NON_REF> END=10019 GT:DP:GQ:MIN_DP:PL 0/0:12:1:12:0,1,287A gVCF can also be incomplete. One of the two gVCFs we measured had no records at all over most of our panel, because the pipeline that wrote it skipped stretches of the genome[18]. The format makes proof of coverage possible; a given file still has to contain it.
What four real 30x files could prove
We cut four real 30x genome files, public exports from the Personal Genome Project, to the 1,054 regions our report examines, 25.5 million bases in all, and counted the bases each file positively confirms[18].
| File | Format | Bases confirmed | Regions at least 90% confirmed | Variants listed |
|---|---|---|---|---|
| Sequencing.com 30x, 2024 | gVCF | 99.0% | 1,038 of 1,054 | 41,952 |
| A second 30x provider | gVCF | 14.2% | 0 of 1,054 | 50,015 |
| Nebula Genomics 30x, 2025 | Plain VCF | 0.17% | 0 of 1,054 | 43,261 |
| Dante Labs 30x, 2024 | Plain VCF | 0.16% | 0 of 1,054 | 41,164 |
Source: Aimosti measurements of 6 October 2026 on public Personal Genome Project files. A base counts as confirmed when the file gives it a called genotype, variant or reference, at a read depth of 10 or more[18].
The two plain VCFs held what a good 30x genome should: 43,261 and 41,164 variants in these regions. What they could not do was say anything about the positions in between. The complete gVCF held a similar number of variants, 41,952, and proved 1,038 of the 1,054 regions at least 90 percent covered. The measured data page has the full numbers, including ten chip exports.
Which file answers which question
| Question | VCF | gVCF | BAM or CRAM |
|---|---|---|---|
| Which variants were found? | Yes | Yes | Yes, once called |
| Was this position read, and how well? | No | Yes, from reference blocks | Yes, from read depth |
| Was a whole gene deleted or duplicated? | Only if the provider ran a structural-variant caller | Indirectly, as low depth in the blocks | Yes, from read depth |
| CYP2D6, including hybrid genes | Small variants only | Small variants only | Yes, with a caller built for it |
| Call the variants again with newer software | No | No | Yes |
| Size for one 30x genome | 4.1 to 5.0 million variant sites | 5.9 GB (NA12878) | 15.8 GB as CRAM (NA12878) |
Source: The specifications and studies cited in the sections above[1, 5, 7, 9, 12, 13, 17].
No one file is best for every question. For most of a report the variant file does the work, and the reads settle the questions about coverage and structure. Providers' downloads differ in which of these files they include; our pages on Nebula, Dante Labs, Sequencing.com and Nucleus go through each one.
How to tell which file you have
File names are a weak guide. We have seen a file named .bam that was a CRAM inside, and a gVCF is often named like any other VCF. The first lines of the file settle it.
What the first lines of a file give away
- The first bytes. A BAM, once decompressed, begins with the characters
BAM, a CRAM begins withCRAM, and a VCF begins with a text line such as##fileformat=VCFv4.2[14, 15, 17]. - Reference blocks. Records with
<NON_REF>or<*>as the alternative allele and anEND=field are reference blocks, and a VCF that has them is a gVCF. GATK's gVCFs also carry##GVCFBlocklines in the header[9, 17]. - The assembly. The header's
##contiglines give each chromosome's length. Chromosome 1 is 248,956,422 bases long in GRCh38 and 249,250,621 in GRCh37[19, 20], and the two assemblies give the same DNA different coordinates[21]. - A CRAM's reference. The
@SQlines in a BAM or CRAM header can carry anM5checksum for each reference sequence, which identifies the exact sequence the reads were placed on[14]. - Or let a tool read it. Our free file check reads these lines in your browser and names the format, the assembly and the kind of file without uploading anything.
What Aimosti would (and wouldn't) show you
The report reads a VCF or gVCF for its clinical, carrier, medication, trait and ancestry panels, and each card says how many of its positions were read from the file and how many were inferred as the reference base. With a gVCF it can say a gene was examined; with a plain VCF it labels those positions as inferred. A BAM or CRAM adds the Deep Read panel, which measures read depth across the curated panel and calls CYP2D6 and whole-gene deletions from the reads. FASTQ is not used.
What we won't claim
We won't call a gene examined on the strength of a plain VCF, count a reference block with no reads behind it as a confirmed base, or decode a CRAM against any reference other than the one it names. A missing record is reported as inferred, never as a clean result.
Bottom line. The most useful folder holds a gVCF rather than a plain VCF, with the BAM or CRAM beside it. Each step down from the reads is smaller and says less, and no later step can put back what an earlier one left out.
Questions people ask
Is a gVCF the same as a VCF?
A gVCF is a VCF, with the same layout and read by the same tools. It adds reference blocks, records that cover runs of positions where the reads matched the reference, so the file says where the sequencing looked as well as what it found[9].
Can a plain VCF be converted into a gVCF?
No. A reference block records the caller's confidence that a run of positions matched the reference, and that confidence comes from the reads[9]. A plain VCF no longer holds them. A gVCF can be made again from the BAM or CRAM.
Why is a CRAM so much smaller than a BAM of the same genome?
A CRAM stores each read as its differences from the reference and compresses each kind of data, names, positions and quality scores, separately. On Illumina data that makes CRAM 3.1 files 50 to 70 percent smaller than BAM[11]. Nothing is lost unless the writer chose a lossy mode, but the file can only be decoded with the reference it names.
What does 30x mean?
It is the average number of reads covering each base. The 1000 Genomes Project's 30x genomes averaged 34x per sample, ranging from 27x to 71x[10]. An average says nothing about any one gene; the files that record depth, a gVCF, BAM or CRAM, can.
Is a 23andMe or AncestryDNA download one of these files?
No. A chip export is a text file listing the positions one genotyping chip was built to measure, between 563,320 and 955,958 rows in the ten we measured[18]. Chips read common variants well and rare ones poorly: in a UK Biobank comparison, only 16 percent of chip calls for variants rarer than 1 in 100,000 were confirmed by sequencing[22]. What that means for a scary result in your own file.
Which file does Aimosti need?
A VCF or gVCF for the full report, and a gVCF lets the report say which genes were examined. A BAM or CRAM adds the Deep Read panel. FASTQ is not used. The free file check names what you have before you pay anything.
References
- Run ERR3239334: 30x whole-genome sequencing of NA12878, read count, base count and CRAM size. European Nucleotide Archive, 2019. 758,135,204 reads; 113,720,280,600 bases; CRAM 15,797,182,294 bytes.
- Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Research, 2010. doi:10.1093/nar/gkp1137
- Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics, 2009. doi:10.1093/bioinformatics/btp352
- Danecek P, Auton A, Abecasis G, et al. The variant call format and VCFtools. Bioinformatics, 2011. doi:10.1093/bioinformatics/btr330
- The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature, 2015. doi:10.1038/nature15393
- Genome assembly GRCh38.p14. NCBI Datasets, 2022. Total sequence length 3,099,441,038 bases.
- 1000 Genomes 30x on GRCh38: per-sample HaplotypeCaller gVCFs. International Genome Sample Resource, 2021. NA12878.haplotypeCalls.er.raw.vcf.gz, 5,948,334,564 bytes.
- Sims D, Sudbery I, Ilott NE, Heger A, Ponting CP. Sequencing depth and coverage: key considerations in genomic analyses. Nature Reviews Genetics, 2014. doi:10.1038/nrg3642
- GVCF: Genomic Variant Call Format. GATK documentation, Broad Institute.
- Byrska-Bishop M, Evani US, Zhao X, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell, 2022. doi:10.1016/j.cell.2022.08.004
- Bonfield JK. CRAM 3.1: advances in the CRAM file format. Bioinformatics, 2022. doi:10.1093/bioinformatics/btac010
- Nofziger C, Turner AJ, Sangkuhl K, et al. PharmVar GeneFocus: CYP2D6. Clinical Pharmacology and Therapeutics, 2020. doi:10.1002/cpt.1643
- Poplin R, Chang PC, Alexander D, et al. A universal SNP and small-indel variant caller using deep neural networks. Nature Biotechnology, 2018. doi:10.1038/nbt.4235
- Sequence Alignment/Map Format Specification. GA4GH / hts-specs, 2025.
- CRAM format specification (version 3.1). GA4GH / hts-specs, 2025.
- Hsi-Yang Fritz M, Leinonen R, Cochrane G, Birney E. Efficient storage of high throughput DNA sequencing data using reference-based compression. Genome Research, 2011. doi:10.1101/gr.114819.110
- The Variant Call Format Specification, VCFv4.3. GA4GH / hts-specs, 2025.
- What your DNA file can actually read: measured on 14 real files. Aimosti, 2026. Measured 6 October 2026; data under CC BY 4.0.
- Homo sapiens chromosome 1, GRCh38.p14 Primary Assembly (NC_000001.11). NCBI Nucleotide.
- Homo sapiens chromosome 1, GRCh37.p13 Primary Assembly (NC_000001.10). NCBI Nucleotide.
- Schneider VA, Graves-Lindsay T, Howe K, et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Research, 2017. doi:10.1101/gr.213611.116
- Weedon MN, Jackson L, Harrison JW, et al. Use of SNP chips to detect rare pathogenic variants: retrospective, population based diagnostic evaluation. BMJ, 2021. doi:10.1136/bmj.n214
Last reviewed . Every number on this page links to the source it comes from; if one of them has moved, tell us.
Free · in your browser
Find out what your own file can read.
The free check opens your DNA file on your device and shows which format you have and what a report could read from it. The file is never uploaded.