GC Content Calculator

Paste a DNA or RNA sequence, or open a FASTA or GenBank file, to get its GC percentage. Multi-FASTA input gets one row per sequence, longer sequences get a sliding-window GC plot, and primers get a Tm and 3′ clamp check.

Runs in your browser. Sequences are not uploaded.

1 FASTA sequence, 600 bases. Example: Synthetic sequence with a GC-rich block, to show the sliding-window plot. Paste your own sequence to replace it.

Examples

GC content

52.00%

312 of 600 bases are G or C.

G+C
31252.00%
A+T
28848.00%
Ambiguous
0
Length
600
A
154 25.7%
C
169 28.2%
G
143 23.8%
T
134 22.3%

GC profile

GC% in a 50 bp window, moved 5 bp at a time. Blue above 50% marks GC-rich stretches, orange below 50% marks AT-rich ones.

Show window values (111 windows)
StartEndGC %GC skew
15048.0%-0.250
65548.0%-0.250
116046.0%-0.304
166546.0%-0.217
217044.0%-0.364
267544.0%-0.364
318046.0%-0.217
368542.0%-0.048
419042.0%-0.048
469540.0%-0.100
5110042.0%-0.048
5610548.0%-0.083
6111046.0%-0.130
6611550.0%-0.040
7112056.0%0.000
7612556.0%0.000
8113054.0%-0.037
8613556.0%-0.071
9114056.0%-0.143
9614552.0%0.000
10115054.0%-0.185
10615554.0%-0.111
11116054.0%-0.111
11616552.0%-0.231
12117052.0%-0.231
12617554.0%-0.259
13118056.0%-0.214
13618556.0%-0.214
14119058.0%-0.172
14619562.0%-0.226
15120062.0%-0.226
15620562.0%-0.290
16121068.0%-0.294
16621572.0%-0.222
17122072.0%-0.222
17622576.0%-0.105
18123080.0%-0.200
18623582.0%-0.220
19124082.0%-0.220
19624582.0%-0.220
20125086.0%-0.116
20625582.0%-0.073
21126082.0%0.024
21626578.0%0.026
22127078.0%0.026
22627576.0%0.000
23128076.0%0.053
23628578.0%0.077
24129078.0%0.128
24629580.0%0.100
25130078.0%0.128
25630584.0%0.048
26131086.0%-0.070
26631586.0%-0.070
27132088.0%-0.045
27632588.0%-0.091
28133086.0%-0.070
28633584.0%-0.095
29134082.0%-0.122
29634584.0%-0.095
30135080.0%-0.200
30635574.0%-0.135
31136068.0%-0.118
31636568.0%-0.118
32137062.0%-0.161
32637556.0%-0.143
33138050.0%-0.200
33638546.0%-0.043
34139042.0%-0.143
34639536.0%0.000
35140038.0%0.158
35640538.0%0.158
36141036.0%0.222
36641536.0%0.222
37142034.0%0.176
37642534.0%0.176
38143034.0%0.294
38643532.0%0.250
39144032.0%0.375
39644530.0%0.200
40145026.0%0.077
40645530.0%-0.067
41146028.0%0.000
41646524.0%-0.167
42147026.0%-0.077
42647526.0%-0.231
43148026.0%-0.231
43648530.0%-0.200
44149030.0%-0.333
44649530.0%-0.333
45150032.0%-0.250
45650528.0%0.000
46151034.0%0.059
46651536.0%0.000
47152038.0%-0.053
47652540.0%-0.100
48153044.0%-0.182
48653538.0%-0.263
49154046.0%-0.217
49654546.0%-0.130
50155046.0%-0.130
50655546.0%-0.217
51156044.0%-0.182
51656542.0%-0.143
52157038.0%-0.053
52657536.0%0.000
53158036.0%0.111
53658538.0%0.053
54159030.0%0.200
54659534.0%0.294
55160032.0%0.250

How to calculate GC content

GC content is the percentage of bases in a DNA or RNA molecule that are guanine (G) or cytosine (C).

GC content (%) = (G + C) / (A + T + G + C) × 100

For RNA, use U in place of T.

  1. 1Count the G and C bases.
  2. 2Count every base whose identity is known: A, T (or U), G and C.
  3. 3Divide the first count by the second and multiply by 100.

ATGCGGCTTA has 3 G and 2 C, so G + C = 5 out of 10 bases and the GC content is 5 / 10 × 100 = 50%.

How this calculator counts

  • Line breaks, spaces, line numbers, gap characters (- and .) and FASTA headers are ignored. Lowercase bases count the same as uppercase ones, so soft-masked repeats are included.
  • S is always G or C, so it counts toward GC. W is always A or T, so it counts toward AT.
  • N and the other ambiguity codes (R, Y, K, M, B, D, H, V) are left out of both the count and the length. This is the default in Biopython's gc_fraction. When a sequence has ambiguous bases you can switch to weighted counting, where R counts as half a G or C and B as two thirds, or count them in the length.
  • Anything else, such as * or letters that only appear in protein sequences, is skipped and listed under the input box so it never changes the result silently.

GC content for PCR primers

Most primer design guidelines aim for 40 to 60% GC. In that range a primer binds firmly enough to be specific without being so stable that it anneals to partial matches.

  • Keep one to three G or C among the last five bases at the 3′ end. This GC clamp holds the end where the polymerase starts, while four or five G or C there make mispriming more likely.
  • Avoid runs of more than four of the same base, and dinucleotide repeats such as ATATATAT.
  • Match the two primers of a pair to within about 5 °C Tm. GC content drives Tm, so primers of similar length and GC% usually have similar Tm.

The calculator runs these checks when every sequence you paste is 8 to 60 bases long. Paste a forward and a reverse primer as two FASTA records to see the Tm difference.

Tm uses the basic formula, Tm = 64.9 + 41 × (G + C − 16.4) / N for primers of 14 or more bases and Tm = 2 × (A + T) + 4 × (G + C) for shorter ones, at 50 mM Na+. It is a quick estimate. Nearest-neighbor methods that account for salt and primer concentration are more accurate for final designs.

What counts as high or low GC content

There is no single cutoff. Sequences above about 60% GC are usually called GC-rich and those below about 40% AT-rich. Whole genomes span a wide range.

Approximate genome GC content of common organisms
OrganismGenome GC
Plasmodium falciparum (malaria parasite)about 19%
Arabidopsis thaliana (thale cress)about 36%
Saccharomyces cerevisiae (baker's yeast)about 38%
Homo sapiens (human)about 41%
Escherichia coli K-12 (lab strain)about 51%
Streptomyces coelicolor (soil bacterium)about 72%

G–C pairs form three hydrogen bonds, one more than A–T pairs, and GC-rich stretches also stack more strongly. That is why GC-rich DNA melts at a higher temperature. In the lab this matters in a few places.

  • PCR. Templates above about 60% GC fold into stable secondary structures and do not fully melt, so they amplify poorly. Additives such as DMSO or betaine and polymerases sold for GC-rich templates help.
  • Gene synthesis and cloning. Synthesis providers check overall and local GC content. A GC-rich or AT-rich block in the profile plot is worth smoothing with synonymous codons before you order.
  • Sequencing. Very GC-rich and very AT-rich fragments tend to get lower coverage in short-read libraries. FastQC plots per-sequence GC content for this reason.

Worked examples

Both primers are in the calculator's examples, so you can check every number.

A standard sequencing primer

M13 forward (−20): GTAAAACGACGGCCAGT

  1. 1Count the bases: 5 G, 4 C, 6 A and 2 T, 17 in total.
  2. 2G + C = 9, so GC content = 9 / 17 × 100 = 52.94%.
  3. 3The last five bases, CCAGT, include three G or C, which gives a GC clamp.
  4. 4Basic Tm = 64.9 + 41 × (9 − 16.4) / 17 = 47.1 °C.

52.94% GC, inside the 40 to 60% range recommended for primers.

A degenerate primer

16S rRNA primer 27F: AGAGTTTGATCMTGGCTCAG. The M at position 12 means A or C.

  1. 1Known bases: 6 G, 3 C, 4 A and 6 T, so G + C = 9 out of 19.
  2. 2Leaving M out (the default): 9 / 19 × 100 = 47.37%.
  3. 3Weighted, with M counted as 0.5: (9 + 0.5) / 20 × 100 = 47.50%.
  4. 4Counting M in the length: 9 / 20 × 100 = 45.00%.

The three conventions differ by up to 2.4 percentage points. Different handling of N and other ambiguity codes is the usual reason two calculators disagree.

Base counts from GC content

In double-stranded DNA every A pairs with a T and every G with a C, so A = T and G = C (Chargaff's rule). One percentage fixes all four bases: G = C = GC% / 2 and A = T = (100 − GC%) / 2.

For example, a 150,000-base stretch of double-stranded DNA that is 64% G+C is 36% A+T, so thymine makes up 18% of the bases: 0.18 × 150,000 = 27,000 thymines. The estimator below starts with these numbers. The rule does not apply to single-stranded DNA or RNA.

BaseShareCount
AAdenine18%27,000
TThymine18%27,000
GGuanine32%48,000
CCytosine32%48,000

GC content 64%, 150,000 nucleotides in total.

GC content FAQ

What is the formula for GC content?

GC content (%) = (G + C) / (A + T + G + C) × 100. Count the guanine and cytosine bases, divide by the number of bases, and multiply by 100. A 20-base primer with 11 G or C is 55% GC.

What is considered high GC content?

Above about 60% is usually called GC-rich, and below about 40% AT-rich. The human genome averages about 41%. GC-rich templates are harder to amplify by PCR and often need additives such as DMSO or betaine.

What GC content should a primer have?

Aim for 40 to 60%. Also check the 3′ end: one to three G or C among the last five bases helps the primer bind, and more than three raises the risk of mispriming. Keep the two primers of a pair within about 5 °C Tm of each other.

Does the calculator count N or other ambiguous bases?

Not by default. N and the other IUPAC ambiguity codes are left out of both the GC count and the length, the same as Biopython's gc_fraction. S counts as GC and W as AT. When ambiguous bases are present you can switch to weighted counting or count them in the length.

Why does another tool give a different GC%?

Usually because it handles N differently. Some tools divide by the full length including N, which lowers the result. Set ambiguous bases to "Count in length" to match them. Headers, spaces and line numbers do not change the result here.

Can I calculate GC content for RNA?

Yes. U counts in place of T, and the results show A+U instead of A+T when the sequence has U and no T.

Can I check several sequences at once?

Yes. Paste a multi-FASTA, FASTQ reads or a GenBank file, or open one. You get one row per sequence with its length, GC% and counts, plus the combined total, and you can download the table as CSV.

What does the GC profile show?

GC% measured in a window that slides along the sequence. Peaks above 50% are GC-rich stretches and dips below 50% are AT-rich ones. GC skew, (G − C) / (G + C), shows strand bias; in bacterial genomes it changes sign near the origin and terminus of replication.

Is my sequence uploaded anywhere?

No. Parsing and counting run in your browser, and files you open are read locally.