Previous Knowledge Required: Binomial Variables (S6)
Motivation
Suppose a standard die is rolled 30 times and we write down the number of times the die lands on the number 6. Let
be the random variable counting the number of times the die lands on 6. We know that
. You might think that
but we do not know if the die is fair.
You would assume it is. This is what your null hypothesis (denoted
) will be: a generally-accepted result or assumption. In this case
is that the die is fair, so
. In that case, we would expect to see 5 times the number 6 (the expectation of
).
Of course, we do not always observe exactly 5 times the die landing on 6. It is going to vary or fluctuate. We know that some outcomes are more likely than others, for example, the die landing on 6 five times is more likely than it landing on 6 all 30 times. The question is, at what point is a result statistically unlikely, and at what point are our basic assumptions wrong? That is the goal of the null hypothesis significance testing (NHST).
Basic Definitions
Using our opening problem, let us say that we threw very few sixes, then we would suspect that
. If we observed a lot of sixes, we would suspect that
. Whichever the observed result is, it will lead to our alternative hypothesis (denoted
). When we do hypothesis testing, we are dealing with one of those two situations typically (having
either greater or lower than a particular number). This is known as a one-tailed test. If we think that
we are conducting a left-side/lower tail test. When we suspect
we are conducting a right-side/upper tail test.
Say we observed the die land on 6 only once. This is our observed value. If we graph the distribution of
, we see that 1 is in the lower tail. We need a measure of the amount of improbability of this particular outcome. So what is the probability of getting less than, or equal to, the observed value? That is, what is:
![]()
That particular value in most cases is
![]()
![]()
We can either reject the null hypothesis or fail to reject the null hypothesis. An important distinction here is that we do not accept the null hypothesis, we simply disprove it or do not have enough evidence to disprove it.
If our alternative hypothesis is simply that the probability is different from the suspected value, then this requires a two-tailed test. If we take a significance level
, we test at half that significance level on each side of the tail.
How to Conduct a NHST
When we undertake a hypothesis test, we take the following steps:
- Establish a null and alternative hypothesis, with relevant probabilities (typically stated in the
question). - Assign probabilities to our null and alternative hypotheses.
- Write out our binomial distribution, assuming the null hypothesis.
- Compare probability against significance level.
- Draw a Conclusion – Reject or fail to reject the null hypothesis.
Example: A standard six-sided die used in a board game is thought to be biased because it does not roll a 1 as frequently as the other five values. In
rolls, 1 appears only three times. At a
significance level, test whether this die is biased. Let
be the number of times a six-sided die was rolled, giving a result of one. Let
be the probability of rolling a 1 in a roll of the die.
Another way to approach this problem is by understanding the critical values. These are the threshold values below/above which we can certainly reject the null hypothesis. In the context of our previous problem, if we rolled a 1 a total of 20 times, it would be obvious that the die is not fair and we would reject the null hypothesis. What about if we roll a 1 only 6 times? What about 9? When do we draw that line. That is the critical value.
Example: Is a normal six-sided die fair when
six is thrown in
throws? Conduct a hypothesis test at a
significance level.
In simpler terms, we observed a number of sixes that is relatively close to the expected number. Therefore, there is evidence to suggest the die is fair, or there is not enough evidence to suggest it is biased. We could also reason this differently. Let us find the number of times that would lead us to suspect the die is biased. That is, we want an
and
such that:
![]()
Example: A teacher believes that
of students watch football on a Saturday afternoon. The teacher asks
students, and
of them watch football on the weekend. State at a
level of significance whether or not the percentage of pupils that watch football on a Saturday afternoon is different from
.
Type I and Type II Errors
These errors are statistical concepts. A type I error is a false positive, i.e.: when you incorrectly reject a true null hypothesis. A type II error is a false negative, i.e.: when you fail to reject a false null hypothesis. In short, we have:
| Reality: | Reality: | |
| We accept | Correct Decision | Type II error |
| We reject | Type I error | Correct Decision |
The process of NHST is the following, we observe (say) some
. We reject the null hypothesis because the probability is lower than our significance level. To calculate the type 1 error, we calculate:
![]()
To calculate a type II error, we find the probability of the complementary event assuming
Example: A quality control engineer believes that
of electronic components produced by a factory require software firmware updates before shipping. They test a random batch of
components and find that
of them actually need the update. Test at a
significance level whether the true proportion is different from
, finding the critical values explicitly. Then, evaluate the probabilities of committing Type I and Type II errors for this quality tracking process.
Exercises
For each of the following independent scenarios, identify the population parameter of interest (
), state the null hypothesis (
) and the alternative hypothesis (
) using proper statistical notation, and clarify whether a one-tailed or a two-tailed test is required.
- Scenario A: A pharmaceutical company claims that
of patients experience immediate relief from a standard cough syrup. A medical researcher suspects that a newly formulated batch has a lower success rate and tests it on a sample of patients. - Scenario B: Historical data shows that
of online shopping carts are abandoned before completion on a retail website. The platform introduces a simplified one-click checkout system and wants to test if the abandonment rate has changed. - Scenario C: A standard seed pack states that the germination rate of wild sunflowers is
. An organic gardener believes that using a specialized soil treatment will increase the proportion of seeds that successfully sprout.
A manufacturing plant produces precision steel bolts. Historically, the machine has a known defect rate of
. After a software optimization update, the floor manager suspects that the proportion of defective bolts has decreased. A random sample of
bolts is drawn from the production line, and exactly
bolt is found to be defective.
- 1. State the null hypothesis (
) and the alternative hypothesis (
) using the population parameter
. - 2. Define the random variable
and state its distribution under the assumption that the null hypothesis is true. - 3. Use a calculator to determine the probability of this observed sample result.
- 4. Conduct the hypothesis test at a
significance level (
) and state the conclusion. - 5. Conduct the hypothesis test at a
significance level (
) and draw the relevant conclusion. - 6. Comment on why changing the significance level alters the final contextual conclusion of the test.
A magician suspects that a coin used in a theatrical performance is biased in favour of heads. To investigate this, the coin is flipped
times.
- 1. State the null hypothesis (
) and the alternative hypothesis (
). - 2. Assuming
is true, define a relevant random variable
and state its distribution. - 3. Use a calculator to find the critical value and the corresponding rejection region for an upper-tail test at 5% significance level.
- 4. Suppose that exactly
heads were observed during the trial. Determine whether the magician can reject the null hypothesis at this significance level. - 5. If instead exactly
heads had been observed, determine whether the magician could reject the null hypothesis. Justify your answer using your established critical values.
A delivery courier claims that
of express parcels arrive before the scheduled time window. A logistics manager suspects the real success rate is lower. They track a random sample of
express shipments to conduct a lower-tail test at a
significance level (
). Under the null hypothesis, the number of on-time arrivals follows the distribution
.
- 1. State the null hypothesis (
) and alternative hypothesis (
). - 2. Use a calculator to show that the critical region (rejection region) for this lower-tail test is given by
. - 3. Calculate the exact probability of committing a Type I error.
- 4. Suppose that the true underlying on-time delivery rate has actually degraded to
(
). Use a calculator to determine the probability of a Type II error.