Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The denominator tells you which probability you are calculating: use the grand total for a joint or marginal probability, and use the size of the stated subgroup for a conditional probability. A two-way table makes the distinction visible: a cell is a joint probability, a row or column total gives a marginal probability, and a cell divided by its row or column total gives a conditional probability.

Start with one table and three readings

Imagine recording two characteristics for each of 200 observations. The table below uses the counts from MacEwan University’s Introduction to Applied Statistics. B and S are simply labels for the row and column categories in that example.

Count S Not S Row total
B 10 30 40
Not B 20 140 160
Column total 30 170 200

Each count describes the same set of observations. What changes is the question you ask and, with it, the reference population in the denominator.

Joint probability: a cell means both

A joint probability asks whether two events occur together. To find the probability of B and S, select their intersecting cell and divide by all 200 observations: P(B ∩ S) = 10/200 = 0.05. In words, 5% of the full set is both B and S.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The joint probability can be written P(A ∩ B) or P(A, B). The Delft MUDE textbook describes a joint probability as the probability that two variables take particular values and calculates it from the cell count over the grand total: Contingency tables.

Marginal probability: a margin means one characteristic, regardless of the other

A marginal probability asks about one variable while ignoring the other. Add the relevant row or column, then divide by the grand total. For B, add the B row: P(B) = 40/200 = 0.20. For S, add the S column: P(S) = 30/200 = 0.15.

The totals are called marginal because they sit at the table’s margins. The same result comes from adding joint probabilities across every possible value of the variable you are ignoring. For example, P(B) = P(B ∩ S) + P(B ∩ not S). This is also called summing out the other variable; see the Delft MUDE explanation and the ProbabilityCourse discussion of joint and marginal distributions.

Conditional probability: a slice means within a subgroup

A conditional probability narrows the reference population to cases where a condition is already true. For P(S|B), read “the probability of S among observations that are B.” Keep only the B row: 10 of its 40 observations are S, so P(S|B) = 10/40 = 0.25.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In general, P(A|B) = P(A ∩ B)/P(B), provided P(B) is greater than zero. The denominator is the probability of the condition, B, because those are the cases being considered. The corresponding counts formula is the joint cell divided by the conditioning row or column total. The Delft MUDE textbook describes the conditional probability as the joint probability divided by the probability of the conditioning event.

How to choose the denominator

Before calculating, ask: “What is my reference population?” The wording of the question often tells you.

  • “Out of all observations” points to the grand total. A cell over the grand total is joint; a row or column total over the grand total is marginal.
  • “Among,” “given,” or “of those who…” points to a subgroup total. Divide the selected cell by the row or column total for that group.
  • A changed condition means a changed denominator. “Among B” uses the B total; “among S” uses the S total.

For the example, P(S|B) = 10/40 = 0.25, while P(B|S) = 10/30 ≈ 0.333. The first considers only the 40 B observations; the second considers only the 30 S observations. Colorado State University’s conditional probability module likewise emphasizes that the denominator determines what population a percentage describes.

Why P(A|B) is not usually P(B|A)

The expression after the vertical bar names the condition—the group you have narrowed to. In P(A|B), B defines the reference group; in P(B|A), A does. The overlap A ∩ B is the same joint event in either direction, but the group totals generally differ. In the table, the overlap count is 10 in both calculations, while the denominators are 40 and 30.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That is why a conditional probability is directional. Knowing the share of B cases that are also S does not, by itself, tell you the share of S cases that are B. They would be equal only in special cases, such as when the relevant base rates and joint probability happen to make the ratios equal.

Connect the three with the product rule

Rearranging the conditional probability formula gives the product rule:

P(A ∩ B) = P(A|B)P(B) = P(B|A)P(A).

Each expression calculates the same joint probability: the probability of landing in the conditioning group multiplied by the probability of the other event within that group. With the example, 0.25 × 0.20 = 0.05, and approximately 0.333 × 0.15 = 0.05. The small displayed rounding in 0.333 aside, both routes lead to the joint probability 0.05.

Bayes’ theorem changes which conditional you are asking for

Bayes’ theorem follows from setting the two product-rule expressions equal and solving for the reverse conditional:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

P(A|B) = P(B|A)P(A)/P(B), where P(B) is greater than zero.

It is useful when you know the probability of B among A cases, P(B|A), but want the probability of A among B cases, P(A|B). The calculation also needs the base rate P(A) and the overall probability of the evidence, P(B). Penn State’s STAT 414 lesson on Bayes’ theorem explains its use when the reverse conditional is known.

In the table, P(B|S) can be found from P(S|B)P(B)/P(S): (0.25 × 0.20)/0.15 ≈ 0.333. Bayes does not make the two conditional probabilities interchangeable; it converts between them using the relevant base rates.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check a conditional distribution within its slice

Once you fix a condition, the probabilities of all possible outcomes for the remaining variable should add to 1. Within the B row, P(S|B) = 10/40 = 0.25 and P(not S|B) = 30/40 = 0.75; together they equal 1. If they do not, check that you included every possible outcome and used the same subgroup denominator for each.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A quick practice routine

  1. Name the event or group in the question. Mark the relevant cell, row, or column in the table.
  2. Identify the scope. Two characteristics together means joint; one characteristic regardless of the other means marginal; one characteristic within a stated group means conditional.
  3. Say the denominator aloud. Use “all observations” for joint or marginal relative frequencies, and “all observations in this condition” for conditional probability.
  4. Calculate and sanity-check. A probability must be between 0 and 1; conditional outcomes covering the full slice must sum to 1.

Try this with the table: what is the probability of not S among B cases? The B row is the reference group, so the answer is 30/40 = 0.75. The key is not memorizing a denominator—it is identifying who is still in the population after the condition is applied.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.