Goal
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts.
Worldwide context
Saved once here, used across the site.
Currency changes display only. Country selection guides tax input; no tax rate is guessed.
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts.
N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq².A clearer path to an answer
This page keeps the calculation transparent: define the goal, enter the matching values, inspect the method, and decide what the result means in your situation.
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts.
Homozygous dominant count (AA) · Heterozygous count (Aa) · Homozygous recessive count (aa)
N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq².
Calculate, review the assumptions below, then compare a related tool when the decision needs more context.
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts.
Open the Hardy–Weinberg Equilibrium Calculator pageMore science tools
Download PDFDownload Word (.doc)
Enter your values above and choose Calculate to see the result here.
Calculation map
N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq².
Bounded, transparent calculation
Your recent runs stay in this browser session only.
Formula: N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq².
This worksheet uses observed genotype counts for a diploid population with two alleles at one locus. It derives allele frequencies and the genotype counts expected under the Hardy–Weinberg model, then reports a Pearson chi-square diagnostic for comparison. The result is educational and exploratory, not a complete inference about a population.
Worked example: N = 100, p = 0.60, q = 0.40, expected counts = 36, 48, and 16, with χ² = 0.
The displayed limits are checked before the handler runs. Model-specific domain checks may also reject impossible or non-finite inputs.
Methodology: This calculator follows the WorldCalculate input, formula, precision, and boundary policy. Read the official methodology.
Calculator usage statistics
This section counts anonymous successful Calculate submissions, not unique visitors. Counts and top tools appear only when trusted aggregate data is available; country analysis is shown only under the same condition and reporting threshold.
Answer-first guide
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts. Start with one clearly defined goal, enter values in the units shown, and keep the result attached to the assumptions below.
This tool is useful when your question includes Hardy Weinberg calculator, population genetics calculator, allele frequency calculator. It returns the outputs declared in the calculator contract rather than a live quote, approval, diagnosis, or professional sign-off.
Homozygous dominant count (AA) · Heterozygous count (Aa) · Homozygous recessive count (aa). Keep the same time period, unit system, and currency wherever the form requires comparable values.
Run the worked example first, compare its output with the page's example, then change one input at a time. This makes an unexpected result easier to trace to a unit, boundary, or assumption.
Need a wider view? Browse Science Calculators or compare the related tools below. The WorldCalculate methodology explains how formulas, examples, limits, and revisions are reviewed.
N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq².
This worksheet uses observed genotype counts for a diploid population with two alleles at one locus. It derives allele frequencies and the genotype counts expected under the Hardy–Weinberg model, then reports a Pearson chi-square diagnostic for comparison. The result is educational and exploratory, not a complete inference about a population.
N = 100, p = 0.60, q = 0.40, expected counts = 36, 48, and 16, with χ² = 0.
Context and background
Science calculators define a system, choose an equation, apply units and constants, and show the substitution. Effects outside that model remain outside the result.
Introductory science problem solving builds from measured quantities and idealized relationships. Those models are valuable for learning and first-pass estimates, while experiments and engineering decisions need additional evidence.
Research and review
Researched by Hassan ALRowaie, Founder and editorial researcher at WorldCalculate.
This guide follows the live calculator's declared inputs, formula, worked example, assumptions, validation boundaries, and source-backed methodology. The review date describes editorial review of the calculator explanation; it is not a promise that external facts or rates remain current.
The Hardy–Weinberg model is useful because it separates three related ideas that are often mixed together: observed genotype counts, allele frequencies, and expected genotype proportions. This calculator takes counts for AA, Aa, and aa, derives p and q, reconstructs the p², 2pq, q² expectation, and reports a chi-square screening value. It is a learning and planning aid, not a declaration that a real population is in equilibrium or that a genetic association has been proved.
Enter the number of individuals observed in the homozygous dominant, heterozygous, and homozygous recessive groups. The page sums them into N, counts the two allele copies carried by each diploid individual, and estimates the frequency of the chosen allele as p. The other allele frequency is q = 1 − p. It then calculates the expected counts under the equilibrium proportions p², 2pq, and q².
The final chi-square value measures the squared, expected-count-scaled differences between observed and expected counts. A value of zero means the entered counts exactly match the model's expectation after the allele frequencies are estimated. A nonzero value does not tell you why the difference exists or whether it is statistically meaningful without a specified testing framework.
A genotype describes the pair of alleles carried at a locus, while an allele frequency describes the share of all allele copies that are a particular form. In a sample of N diploid individuals there are 2N allele copies. Each AA individual contributes two A copies, each Aa contributes one A and one a, and each aa contributes no A copies.
This distinction explains the numerator 2AA + Aa in the p formula. Counting people in the heterozygous group as if they carried two A copies would overstate p. Before entering numbers, confirm that the three categories are mutually exclusive and collectively exhaustive for the locus and sample being studied.
For counts AA, Aa, and aa, the calculator uses p = (2AA + Aa)/(2N) and q = (Aa + 2aa)/(2N). Since every allele copy is either A or a, the two frequencies add to one apart from ordinary numerical rounding. The page reports both values so a student can inspect the complement rather than trusting a hidden transformation.
For example, with 36 AA, 48 Aa, and 16 aa individuals, N is 100. There are 120 A copies among 200 total copies, so p = 0.60. The remaining 80 copies are a, giving q = 0.40. Those frequencies become the inputs for the expected genotype proportions.
The ideal model treats the two allele draws that form a genotype as independent with probabilities p and q. Two A draws have probability p², one A and one a can occur in two orders and therefore have probability 2pq, and two a draws have probability q². These proportions sum to one because they cover the three possible unordered genotypes.
Multiplying each proportion by N converts a proportion into an expected count. The expected numbers are not predictions about named individuals; they are model-based reference counts for the population table. A fractional expected count is normal and should not be rounded before the chi-square calculation.
With 36 AA, 48 Aa, and 16 aa, p = 0.6 and q = 0.4. The expected proportions are 0.36, 0.48, and 0.16. Multiplying by 100 gives 36, 48, and 16, exactly matching the observations. Each observed-minus-expected difference is zero, so the chi-square screening value is zero.
This result is a useful arithmetic check, not evidence that every assumption has been met. The frequencies were estimated from the same table, the sample could still be biased, and a perfect classroom example can hide the questions that matter in real data. Keep the worked steps visible when using it for teaching or review.
Suppose a sample contains 10 AA, 50 Aa, and 40 aa. N is 100 and p = (20 + 50)/200 = 0.35, so q = 0.65. The expected counts become 12.25 AA, 45.5 Aa, and 42.25 aa. The observed table differs from those expectations, and the chi-square value summarizes the magnitude after scaling by expected counts.
The difference may reflect chance, selection, nonrandom mating, population mixing, genotyping error, an incorrect locus model, or a sample that is not representative. The calculator does not label the cause. Use the output to ask the next scientific question rather than to attach a biological story automatically.
Hardy–Weinberg equilibrium is a model condition in which allele frequencies remain stable from one generation to the next under a set of idealized assumptions. It is commonly taught with random mating, a very large population, no selection, no mutation, no migration, and no genetic drift. The calculation here focuses on the genotype-frequency relationship, not on verifying every assumption.
A real population can depart from the model for many reasons. Even if observed counts are close to expectation, the result does not prove that mating is random or that no evolutionary force is present. It may simply mean that the sample and statistical power do not reveal a departure.
A sample made by combining distinct subpopulations can show an overall genotype pattern that differs from each subgroup. Conversely, a mixture can sometimes hide departures within groups. Record where, when, and how the sample was collected, and do not treat a convenient sample as a census of a species or community.
The unit of analysis matters. Individuals from different families, locations, age groups, or breeding lines may not share one allele pool. If the scientific question concerns a defined population, analyse that population according to the study design and report exclusions, missing genotypes, and quality-control rules.
The calculator expects a complete count in one of three genotype categories. Laboratory calls can be uncertain, missing, or subject to allele dropout and misclassification. Forcing an unknown observation into AA, Aa, or aa changes both the allele frequencies and the expected counts. Keep a separate record of missing or failed calls and follow the assay's quality rules.
If the locus has more than two alleles, copy-number variation, polyploidy, or an uncertain ploidy state, this three-category model is not enough. Do not collapse a complex locus into two groups merely to make the input fields fit. The correct model should reflect the biology and the measurement process.
The Pearson chi-square screen is the sum of (observed − expected)² divided by expected for the three genotype groups. Large contributions identify which groups differ most from the expectation, but this page reports only the total. A formal test requires a chosen null hypothesis, degrees of freedom, significance threshold, and attention to how allele frequencies were estimated.
Small expected counts can make an ordinary asymptotic chi-square approximation unreliable. In those cases, a specialist may use an exact test, simulation, or another method. The calculator intentionally does not print a universal pass-or-fail label because the correct inference depends on design and context.
Counts are whole individuals, not percentages. Enter the same locus and allele coding throughout the table. If the dominant or recessive terminology is not biologically appropriate, use neutral allele labels in the study record and interpret the fields as the first homozygote, heterozygote, and second homozygote categories.
Dominance is a phenotype relationship and does not determine how allele frequency is calculated. A recessive phenotype may conceal heterozygotes, which is why genotype data and phenotype data are not interchangeable. If only phenotypes are available, this calculator should not be used as though the missing genotypes were known.
For a classroom exercise, begin with a table whose expected counts are integers, then change one genotype count and observe p, q, the expected counts, and chi-square. Ask students to calculate the allele copies by hand before checking the page. This makes the two-copy denominator and the heterozygote contribution easier to remember.
A second exercise can compare a small sample with a large sample that has the same observed proportions. The proportions may be identical while the chi-square value changes because expected counts scale with N. That illustrates why effect size, sample size, and uncertainty should be discussed together.
A useful handoff includes the locus name, allele coding, sample definition, observed counts, missing-call policy, p and q estimates, expected counts, chi-square value, and the formal method planned for inference. Store the calculator output with the dataset version rather than copying only a rounded p value into a report.
If the result will inform a health, conservation, breeding, or population decision, involve a qualified geneticist or statistician. A simple equilibrium screen can be a good quality-control question, but it cannot establish disease risk, ancestry, fitness, or treatment response by itself.
The most common arithmetic error is entering percentages as counts or forgetting that a heterozygote contributes one copy of each allele. Another is entering phenotype categories that do not map one-to-one onto genotypes. A third is mixing samples, loci, or allele orientations. Check that AA + Aa + aa equals the intended analysed sample.
Also check the scale of the result. p and q must lie between zero and one and should add to one within rounding. Expected counts should add to N. If a value fails those simple checks, correct the data definition before interpreting the chi-square output.
The calculator does not identify whether a population is evolving, whether a variant is harmful, whether a sample is contaminated, or whether a study meets an ethics or reporting standard. It does not replace laboratory validation, a pedigree analysis, a population-structure model, or a pre-registered statistical plan.
Its value is transparency. A reader can see which counts produced the allele frequencies and expected values, reproduce the arithmetic, and decide whether a more appropriate method is needed. That is safer than presenting a single equilibrium label without the assumptions that created it.
Confirm the locus, allele names, ploidy, and genotype categories. Confirm that counts are from the same population and that missing calls are handled explicitly. Enter whole counts, inspect N, verify p + q, and compare expected counts with the observed table.
Then document the sampling design and choose the formal inferential method. Treat the chi-square value as a diagnostic prompt, not a verdict. Re-run the calculation if quality control changes any genotype count, and retain the original table so another reviewer can reproduce the decision.
A phenotype is an observed trait while an allele frequency is a count of allele copies. A dominant phenotype can contain both AA and Aa individuals, so phenotype counts alone usually cannot fill the three genotype fields. The page is designed for genotype data or for a teaching exercise where the genotypes are explicitly known.
If a study has only phenotypes, do not infer hidden heterozygotes by casually assigning all dominant-looking individuals to AA. That shortcut changes p and q and can create a false appearance of equilibrium. Use an appropriate genetic model or collect genotype information under a validated protocol.
Some datasets arrive as proportions rather than integer counts. Convert them back only when the original sample size and rounding rules are known. Entering rounded percentages multiplied by an assumed N can create totals that do not correspond to real individuals and can distort the chi-square screen.
For teaching, it is fine to create a synthetic table, but label it as synthetic. For research, preserve the raw count table and use the calculator as a check. A frequency may be useful for a graph while the count is needed for expected values and uncertainty.
Hardy–Weinberg teaching examples often list random mating as an assumption, while the calculator input comes from a sample. Random sampling and random mating are different concepts. A well-sampled population can still have nonrandom mating, and a randomly mating population can be sampled poorly.
Keep those questions separate in the interpretation. The arithmetic can show that counts fit p², 2pq, and q², but it cannot observe mate choice, geographic structure, or the route by which individuals entered the sample. Good notes state which assumptions were measured and which were only idealized.
The ideal equilibrium model assumes a very large population. In a small population, random genetic drift can change allele frequencies between generations even when no selection is present. A single sample that fits or misses the expected counts does not tell the full history of drift.
If the data come from a conservation, breeding, or isolated population, describe the census and effective population context. The page can calculate p and q for that population, but a specialist should choose the evolutionary model and uncertainty method for the real question.
A researcher may check equilibrium at many loci. Some chi-square values can look unusual by chance when many tests are performed. The calculator reports one locus at a time and does not adjust a collection of tests or choose a significance threshold.
Keep a pre-specified analysis plan when the result will support a study. Record the number of loci checked, quality filters, missing data, and correction method. A visually surprising single locus should be replicated and reviewed rather than promoted to a biological conclusion immediately.
The simple p², 2pq, q² model assumes a diploid locus with the same allele-copy structure for the analysed individuals. Sex-linked loci, haploid stages, polyploid organisms, copy-number variants, and mixed ploidy need different denominators or models.
If the biological system is not a standard autosomal diploid case, do not force it into the fields because the labels look familiar. Write the copy-number rule, stratify the sample if appropriate, and use a specialist method. A clean result from the wrong model is not a correct result.
Allele frequencies estimated from a finite sample have uncertainty. The calculator reports point estimates and expected counts but no confidence interval. A formal report may use a binomial or multinomial approach, resampling, or a model that accounts for the study design.
Include the sample size and raw counts beside p and q. A value of p = 0.60 from N = 100 does not carry the same information as p = 0.60 from a much larger or much more structured sample. The displayed decimals are not a substitute for uncertainty analysis.
Laboratories often use signal thresholds and quality scores to assign genotype categories. A change in threshold can move observations between AA, Aa, and aa without any biological change. When reviewing a result, record the assay, call rate, quality filters, and version of the calling pipeline.
The calculator accepts final counts, so it cannot inspect a genotype cluster or an ambiguous read. If uncertain calls are important, perform the quality review before entering the table. A clean mathematical output should not conceal measurement uncertainty at the source.
The simple model treats sampled individuals as observations from the population, but close relatives can carry correlated alleles. A family-heavy sample may not represent independent population draws. Record whether the design includes relatives, and use a method that accounts for relatedness when the scientific question requires it.
This does not make the calculator useless. It makes the interpretation narrower: it can summarize the entered counts, while the study design determines whether the summary is a valid equilibrium test. Keep the count table and sampling frame together.
Allele frequencies may differ between life stages if survival or reproduction is associated with the locus. A sample of adults is not automatically comparable with a sample of embryos, seedlings, or newborns. Name the life stage and sampling time in the analysis record.
A departure from expected genotype counts can be a clue about selection, but it can also reflect structure, sampling, or genotyping error. The calculator cannot distinguish those explanations. Use it to identify a pattern and then test a biologically appropriate hypothesis.
Store the three observed counts in a small table with column names, allele labels, locus identifier, population identifier, and missing-data note. Save the exact inputs used in the calculator. A later reviewer should be able to recompute N, p, q, expected counts, and chi-square without asking which row was intended.
When publishing, show enough digits to reproduce the rounded result and state the software or worksheet version. Do not replace the raw counts with only p and q; many different samples can share similar frequencies while carrying different evidence and uncertainty.
A useful follow-up is to compare the observed table with an alternative biological or sampling explanation. The calculator supplies the equilibrium expectation, while a researcher can test structure, selection, or genotyping error with a model suited to the data. Do not choose the explanation simply because it makes the chi-square value smaller.
Replication across an independent sample or time point is often more informative than a single surprising table. Preserve the same allele coding and quality rules so the results can be compared. If the samples differ in design, report that difference instead of forcing one pooled interpretation.
A plain-language summary can say that the page counted allele copies, estimated p and q, and compared the observed genotype groups with the ideal p², 2pq, q² pattern. It should also say that the comparison is a screen and that real populations may violate the assumptions.
Avoid phrases such as proves equilibrium or proves evolution. A careful summary tells the reader what was measured, what model was used, and what needs expert review next. That is useful for a student report as well as for a research handoff.
A clear result sentence names the sample size, allele frequencies, expected genotype counts, and chi-square diagnostic, then states that the model assumptions require review. Include the raw genotype counts so a reader can distinguish a measured table from a theoretical example.
This wording keeps the arithmetic useful without turning a screening comparison into a biological verdict. It also makes the result easier to translate, teach, audit, and compare with a later sample that uses the same definitions.
Estimate allele frequencies, expected two-allele genotype counts, and a chi-square screening value from observed genotype counts.
N = AA + Aa + aa; p = (2AA + Aa)/(2N); q = 1 − p; expected counts = Np², N(2pq), and Nq². This worksheet uses observed genotype counts for a diploid population with two alleles at one locus. It derives allele frequencies and the genotype counts expected under the Hardy–Weinberg model, then reports a Pearson chi-square diagnostic for comparison. The result is educational and exploratory, not a complete inference about a population.
Enter Homozygous dominant count (AA), Heterozygous count (Aa), Homozygous recessive count (aa), then choose Calculate.
There are two alleles at the locus and individuals are diploid. AA, Aa, and aa counts refer to the same sampled population and locus. The heterozygote label combines Aa and aA because genotype order is not distinguished. Allele frequencies are estimated from the observed counts. Expected counts use p², 2pq, and q² under the idealized equilibrium model. The chi-square output is a screening value and does not include a chosen critical value or p-value. Sampling design, population structure, selection, migration, mutation, and linkage are not modeled. Biological interpretation requires the locus definition, sampling context, and appropriate statistical review.
This calculator is part of the WorldCalculate library. Its formula, example, assumptions, input bounds, and output formatting follow the official methodology.
These WorldCalculate collections connect this tool with related questions while keeping each calculation separate and transparent.