T-Test Calculator
Run a one-sample, independent two-sample (Welch or pooled) or paired t-test from raw data or summary statistics, with the p-value, confidence interval and every step.
Null hypothesis
H₀: μ = 14
Alternative hypothesis
H₁: μ ≠ 14
- t-statistic
- 1.2247
- Degrees of freedom
- 6
- p-value
- 0.2666
- Two-tailed
- Significance level
- α = 0.05
- Sample mean
- 15
- μ₀ = 14
- Standard error
- 0.8165
- 95% confidence interval
- 13.0021 to 16.9979
- For the population mean; t* = 2.4469
- Effect size (Cohen’s d)
- 0.4629
- (x̄ − μ₀) / s
Estimate, interval and test agree: the 95% confidence interval contains the null value 14, which matches the decision at α = 0.05. Statistical significance is not practical importance: a tiny effect can be significant in a large sample, so judge the size of the difference (the interval and Cohen’s d) in context. Cohen’s d measures magnitude, not significance; 0.2, 0.5 and 0.8 are often quoted as small, medium and large, but those are rough conventions.
Visualization
Sample mean with its 95% confidence interval
Data summary
| Statistic | Sample |
|---|---|
| n | 7 |
| Mean | 15 |
| Median | 15 |
| SD | 2.1602 |
| SE | 0.8165 |
| Minimum | 12 |
| Maximum | 18 |
Step-by-step calculation
Step 1: Sample statistics
n = 7, x̄ = 15, s = 2.1602
Step 2: Standard error
SE = s / √n = 2.1602 / √7 = 0.8165
Step 3: Degrees of freedom
df = n − 1 = 7 − 1 = 6
Step 4: t-statistic
t = (estimate − null value) / SE = (15 − 14) / 0.8165
t = 1.2247
Step 5: p-value
From the t-distribution with df = 6 (two-tailed: both tails).
p = 0.2666
Step 6: Decision
p = 0.2666 ≥ α = 0.05: fail to reject H₀
Running many t-tests on the same data raises the chance that at least one is significant by chance alone. A significant result shows evidence of a difference in means, not what caused it.
T-test formulas
One-sample, df = n − 1t = (x̄ − μ₀) / (s / √n)
Welch two-sample, Welch–Satterthwaite dft = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂)
Pooled two-sample, df = n₁ + n₂ − 2t = (x̄₁ − x̄₂) / √(sₚ²(1/n₁ + 1/n₂))
Paired, df = n − 1t = d̄ / (s_d / √n)
x̄ = sample mean, μ₀ = hypothesized mean, s = sample standard deviation, n = sample size, sₚ² = [(n₁ − 1)s₁² + (n₂ − 1)s₂²] / (n₁ + n₂ − 2) is the pooled variance, d̄ and s_d = mean and standard deviation of the paired differences.
Quick guide: choosing the test
- One-sample: one sample, compared with a reference or target value (average delivery time vs a 30-minute target).
- Independent two-sample: two separate groups with different subjects (two classes taught by different methods). Use Welch’s version unless you have good reason to assume equal variances.
- Paired: each value in one set has a natural partner in the other (the same people before and after, or matched pairs).
| T-test | Z-test |
|---|---|
| Uses the sample SD when the population SD is unknown | Uses a known population SD, or a large-sample normal approximation for some tests |
| Student t-distribution | Standard normal distribution |
| The usual test for comparing means | Used where its assumptions apply, such as some proportion tests |
Comparing more than two group means? Use the ANOVA Calculator rather than many t-tests. To estimate a population mean range from sample statistics, see the Confidence Interval Calculator; to standardize an individual value rather than compare means, the Z-Score Calculator; and to plan how many observations you need, the Sample Size Calculator. Get the mean and SD of raw data with the Standard Deviation Calculator.
A t-test checks whether a mean, or a difference between means, is far enough from a hypothesized value to be unlikely by chance. This t-test calculator runs the three common versions: the one-sample t-test (one mean against a reference value), the independent two-sample t-test (two separate groups, with Welch’s or the pooled equal-variance method) and the paired t-test (two measurements on the same subjects).
Enter raw data or summary statistics, choose a two-tailed, left-tailed or right-tailed alternative and a significance level, and the calculator returns the t-statistic, degrees of freedom, p-value and decision, the hypotheses tested, a confidence interval, Cohen's d, a data summary, a chart and every calculation step.
Worked Calculation Examples
| Scenario | Result | Calculation Step |
|---|---|---|
| One-sample: delivery times vs a 30-minute target (hypothetical) | t ≈ 3.07, df = 11, p ≈ 0.0053 | x̄ ≈ 32.17, s ≈ 2.44, n = 12, SE ≈ 0.705. t = (32.17 − 30) / 0.705 ≈ 3.07 with df = 11, one-tailed p ≈ 0.0053 < 0.05, so reject H₀: the data provide evidence that the mean delivery time exceeds 30 minutes. The 95% one-sided lower bound is about 30.9 minutes, and Cohen’s d ≈ 0.89. |
| Independent two-sample: two teaching methods (hypothetical summary statistics) | Welch t ≈ 1.70, df ≈ 50.7, p ≈ 0.096 | Mean difference 72.4 − 68.2 = 4.2 points, SE = √(8.1²/30 + 10.5²/28) ≈ 2.475, t ≈ 1.70 with Welch df ≈ 50.7, two-tailed p ≈ 0.096 > 0.05: fail to reject H₀. The 95% confidence interval for the difference, −0.77 to 9.17 points, includes 0 but also differences of several points, so the study is inconclusive rather than showing no difference. |
| Paired: scores before and after training (hypothetical) | t ≈ 4.49, df = 7, p ≈ 0.0028 | Differences (After − Before): 4, 3, 3, −1, 6, 6, 3, 4. d̄ = 3.5, s_d ≈ 2.20, SE ≈ 0.779, t ≈ 4.49 with df = 7, two-tailed p ≈ 0.0028, so reject H₀: scores tend to be higher after training. The 95% CI for the mean improvement is 1.66 to 5.34 points. Without a control group, this does not show the training caused the improvement. |
What Is a T-Test?
A t-test is a hypothesis test for means. It divides the difference between an estimate and the value stated in the null hypothesis by the standard error of that estimate, giving the t-statistic, and then asks how unusual that value would be under Student's t-distribution. Unlike a z-test, it uses the sample standard deviation, and the t-distribution's heavier tails allow for the extra uncertainty of estimating the spread from the data.
Which T-Test Should You Use?
Use a one-sample t-test when you have one sample and want to compare its mean with a reference or target value, such as whether average delivery time exceeds a 30-minute target.
Use an independent two-sample t-test when you have two separate groups made up of different subjects, such as students taught by two different methods. Welch's version is the safer default because it does not assume equal variances; use the pooled version only when equal population variances are a reasonable assumption. Deciding by first running a variance test is not recommended.
Use a paired t-test when every value in one set has a natural partner in the other: the same people measured before and after, the same product tested under two conditions, or matched pairs. The test works on the differences within each pair, which removes the variation between subjects. Treating paired data as independent groups throws that advantage away.
How to Calculate a T-Test
One-sample: find the sample mean x̄ and standard deviation s, compute the standard error SE = s / √n, then t = (x̄ − μ₀) / SE with df = n − 1. For the sample 12, 15, 14, 18, 16, 13, 17 and μ₀ = 14: x̄ = 15, s ≈ 2.160, SE ≈ 0.816, t ≈ 1.225, df = 6, two-tailed p ≈ 0.267.
Independent (Welch): compute each group's mean and variance, the difference x̄₁ − x̄₂, SE = √(s₁²/n₁ + s₂²/n₂), t = difference / SE and the Welch–Satterthwaite degrees of freedom. For 10, 12, 15, 14, 13 against 18, 17, 20, 19, 21: difference −6.2, SE ≈ 1.114, t ≈ −5.568, df ≈ 7.71, p ≈ 0.0006.
Paired: compute each difference d = After − Before, then run a one-sample t-test on the differences against 0. For Before 10, 12, 15, 13, 16 and After 12, 15, 16, 15, 18: differences 2, 3, 1, 2, 2, d̄ = 2, s_d ≈ 0.707, SE ≈ 0.316, t ≈ 6.325, df = 4, p ≈ 0.0032.
How to Interpret T-Test Results
Compare the p-value with the significance level α chosen in advance. If p < α, reject the null hypothesis: the result is statistically significant at that level. If p ≥ α, fail to reject it: the data do not provide enough evidence of a difference. Failing to reject is not the same as showing there is no difference, and nothing is ever proved true or false.
A two-tailed test looks for a difference in either direction; a one-tailed test looks in one direction only and must be chosen before seeing the data. The confidence interval gives the range of plausible values for the mean or difference; when a two-sided interval excludes the null value, the two-tailed test rejects at the same α.
T-Test Assumptions
The outcome is quantitative (measured on a numerical scale).
Observations are independent: within each group for the one-sample and independent tests, and between pairs for the paired test. The two groups in an independent test must be unrelated.
For small samples, the data (or, for a paired test, the differences) should be roughly normal, without severe outliers. The t-test is fairly robust to mild non-normality, and larger samples make it more so, but strong skew and outliers can still distort it.
The pooled two-sample test assumes equal population variances; Welch's test does not.
The calculator cannot check how your data were collected. It flags values beyond Tukey's fences (1.5 × IQR past the quartiles) for review, but never removes them.
Statistical Significance vs Practical Significance
A small p-value says the data are hard to explain by chance alone; it does not say the difference is large or important. With a big enough sample, a difference of a fraction of a point can be significant. Look at the estimated difference and its confidence interval in the units you care about, at an effect size such as Cohen's d, and at the context: what size of difference would actually matter? Cohen's d values of 0.2, 0.5 and 0.8 are often called small, medium and large, but these are rough conventions, not rules.
If you run many t-tests on the same data, the chance that at least one is significant by chance alone increases. For more than two groups, use ANOVA rather than many pairwise t-tests.
T-Test vs Z-Test and ANOVA
A z-test uses the standard normal distribution and a known population standard deviation, or a large-sample normal approximation in tests such as those for proportions. A t-test uses the sample standard deviation and the t-distribution. It is not simply a matter of sample size: for means with an unknown standard deviation the t-test is correct at any sample size, and it becomes almost identical to the z-test as the degrees of freedom grow.
A t-test compares one mean with a value or two means with each other. ANOVA compares the means of three or more groups in a single test, which controls the false-positive rate better than running every pairwise t-test.
How to Use the T-Test Calculator
- Choose the test: one-sample (one mean against a reference value), independent two-sample (two separate groups) or paired (two measurements on the same subjects or matched pairs).
- Enter raw data, one list per group, or switch to summary statistics and enter the means, standard deviations and sample sizes.
- Choose the alternative hypothesis (two-tailed, left-tailed or right-tailed) and the significance level α; for two groups, choose Welch’s test (recommended) or the pooled equal-variance test.
- Read the t-statistic, degrees of freedom, p-value and decision, with the hypotheses that were tested, the confidence interval for the mean or difference, and Cohen’s d.
- Check the data summary and chart for unusual values, and follow the step-by-step calculation.
Frequently Asked Questions
What is a t-test?
A hypothesis test for means: it compares a sample mean with a value, or two means with each other, using the t-statistic and Student’s t-distribution.
When should I use a t-test?
When the outcome is numerical and you want to compare one mean with a reference value, two independent group means, or paired measurements, and the population standard deviation is unknown.
What is a one-sample t-test?
It tests whether a population mean differs from a hypothesized value μ₀: t = (x̄ − μ₀) / (s/√n) with df = n − 1.
What is an independent two-sample t-test?
It compares the means of two separate groups of different subjects. Welch’s version, the recommended default, does not assume equal variances; the pooled version does.
What is a paired t-test?
It compares two measurements on the same subjects or matched pairs by testing whether the mean of the within-pair differences is 0: t = d̄ / (s_d/√n), df = n − 1.
What is Welch’s t-test?
A two-sample t-test that uses each group’s own variance, SE = √(s₁²/n₁ + s₂²/n₂), and the Welch–Satterthwaite degrees of freedom, which are often not a whole number. It does not assume equal variances.
What is the t-statistic?
The estimate minus the hypothesized value, divided by its standard error. It measures how many standard errors the observed mean or difference is from the null value.
How do you calculate a t-test p-value?
From the t-distribution with the test’s degrees of freedom: for a two-tailed test, p = 2 × P(T > |t|); for a right-tailed test, P(T > t); for a left-tailed test, P(T < t).
What are the degrees of freedom in a t-test?
n − 1 for one-sample and paired tests, n₁ + n₂ − 2 for the pooled two-sample test, and the Welch–Satterthwaite value for Welch’s test. They set the shape of the t-distribution.
What does p < 0.05 mean in a t-test?
At the 0.05 significance level you reject the null hypothesis: a difference this large would be unlikely if the null hypothesis were true. It does not say the difference is large or important.
What is the difference between a t-test and a z-test?
A t-test uses the sample standard deviation and the t-distribution; a z-test uses a known population standard deviation or a large-sample normal approximation and the normal distribution.
Does a significant t-test prove causation?
No. It shows evidence of a difference in means. Whether one thing caused it depends on the study design, for example random assignment to groups.
What is the difference between statistical and practical significance?
Statistical significance means the difference is unlikely to be due to chance; practical significance means it is large enough to matter. Judge the latter from the effect size, the confidence interval and the context.
Related Calculators
ANOVA Calculator
Run a one-way ANOVA (classical or Welch) on two or more groups: F-statistic, p-value, ANOVA table, effect sizes, charts and every calculation step.
Confidence Interval Calculator
Calculate a confidence interval for a population mean from the sample mean, standard deviation and sample size, using t or z, with the margin of error and every step.
Z-Score Calculator
Calculate a z-score from a value, mean and standard deviation, with the steps, what it means, a bell-curve chart and the normal-distribution percentile.
Sample Size Calculator
Find the minimum sample size for a survey from the confidence level, margin of error, expected proportion and optional population size, with every step shown.
Standard Deviation Calculator
Calculate sample and population standard deviation and variance of a dataset, with the mean, a deviation table and step-by-step working.
Chi-Square Calculator
Run a chi-square goodness-of-fit test or test of independence: χ², degrees of freedom, p-value, expected counts, contributions and a clear decision.
Last updated: September 27, 2026.