T-Test Calculator

Run a one-sample, independent two-sample (Welch or pooled) or paired t-test from raw data or summary statistics, with the p-value, confidence interval and every step.

Test type

7 values

The reference value you are testing against

Alternative hypothesis

Tests whether the mean (or difference) differs in either direction.

Significance level (α)
One-sample t-test (two-tailed)t = 1.2247df = 6, p = 0.2666, α = 0.05Fail to reject the null hypothesisAt the 0.05 significance level, the data do not provide enough evidence that the population mean differs from 14. This does not show the null hypothesis is true.

Null hypothesis

H₀: μ = 14

Alternative hypothesis

H₁: μ ≠ 14

t-statistic
1.2247
Degrees of freedom
6
p-value
0.2666
Two-tailed
Significance level
α = 0.05
Sample mean
15
μ₀ = 14
Standard error
0.8165
95% confidence interval
13.0021 to 16.9979
For the population mean; t* = 2.4469
Effect size (Cohen’s d)
0.4629
(x̄ − μ₀) / s

Estimate, interval and test agree: the 95% confidence interval contains the null value 14, which matches the decision at α = 0.05. Statistical significance is not practical importance: a tiny effect can be significant in a large sample, so judge the size of the difference (the interval and Cohen’s d) in context. Cohen’s d measures magnitude, not significance; 0.2, 0.5 and 0.8 are often quoted as small, medium and large, but those are rough conventions.

Visualization

12131415161718μ₀ = 14SampleSample: 12Sample: 15Sample: 14Sample: 18Sample: 16Sample: 13Sample: 17mean 15

Sample mean with its 95% confidence interval

141516null value 14Mean

Data summary

StatisticSample
n7
Mean15
Median15
SD2.1602
SE0.8165
Minimum12
Maximum18

Step-by-step calculation

  1. Step 1: Sample statistics

    n = 7, x̄ = 15, s = 2.1602

  2. Step 2: Standard error

    SE = s / √n = 2.1602 / √7 = 0.8165

  3. Step 3: Degrees of freedom

    df = n − 1 = 7 − 1 = 6

  4. Step 4: t-statistic

    t = (estimate − null value) / SE = (15 − 14) / 0.8165

    t = 1.2247

  5. Step 5: p-value

    From the t-distribution with df = 6 (two-tailed: both tails).

    p = 0.2666

  6. Step 6: Decision

    p = 0.2666 ≥ α = 0.05: fail to reject H₀

Running many t-tests on the same data raises the chance that at least one is significant by chance alone. A significant result shows evidence of a difference in means, not what caused it.

T-test formulas

One-sample, df = n − 1t = (x̄ − μ₀) / (s / √n)

Welch two-sample, Welch–Satterthwaite dft = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂)

Pooled two-sample, df = n₁ + n₂ − 2t = (x̄₁ − x̄₂) / √(sₚ²(1/n₁ + 1/n₂))

Paired, df = n − 1t = d̄ / (s_d / √n)

x̄ = sample mean, μ₀ = hypothesized mean, s = sample standard deviation, n = sample size, sₚ² = [(n₁ − 1)s₁² + (n₂ − 1)s₂²] / (n₁ + n₂ − 2) is the pooled variance, d̄ and s_d = mean and standard deviation of the paired differences.

Quick guide: choosing the test

  • One-sample: one sample, compared with a reference or target value (average delivery time vs a 30-minute target).
  • Independent two-sample: two separate groups with different subjects (two classes taught by different methods). Use Welch’s version unless you have good reason to assume equal variances.
  • Paired: each value in one set has a natural partner in the other (the same people before and after, or matched pairs).
T-test vs z-test
T-testZ-test
Uses the sample SD when the population SD is unknownUses a known population SD, or a large-sample normal approximation for some tests
Student t-distributionStandard normal distribution
The usual test for comparing meansUsed where its assumptions apply, such as some proportion tests

Comparing more than two group means? Use the ANOVA Calculator rather than many t-tests. To estimate a population mean range from sample statistics, see the Confidence Interval Calculator; to standardize an individual value rather than compare means, the Z-Score Calculator; and to plan how many observations you need, the Sample Size Calculator. Get the mean and SD of raw data with the Standard Deviation Calculator.

A t-test checks whether a mean, or a difference between means, is far enough from a hypothesized value to be unlikely by chance. This t-test calculator runs the three common versions: the one-sample t-test (one mean against a reference value), the independent two-sample t-test (two separate groups, with Welch’s or the pooled equal-variance method) and the paired t-test (two measurements on the same subjects).

Enter raw data or summary statistics, choose a two-tailed, left-tailed or right-tailed alternative and a significance level, and the calculator returns the t-statistic, degrees of freedom, p-value and decision, the hypotheses tested, a confidence interval, Cohen's d, a data summary, a chart and every calculation step.

Worked Calculation Examples

ScenarioResultCalculation Step
One-sample: delivery times vs a 30-minute target (hypothetical)t ≈ 3.07, df = 11, p ≈ 0.0053x̄ ≈ 32.17, s ≈ 2.44, n = 12, SE ≈ 0.705. t = (32.17 − 30) / 0.705 ≈ 3.07 with df = 11, one-tailed p ≈ 0.0053 < 0.05, so reject H₀: the data provide evidence that the mean delivery time exceeds 30 minutes. The 95% one-sided lower bound is about 30.9 minutes, and Cohen’s d ≈ 0.89.
Independent two-sample: two teaching methods (hypothetical summary statistics)Welch t ≈ 1.70, df ≈ 50.7, p ≈ 0.096Mean difference 72.4 − 68.2 = 4.2 points, SE = √(8.1²/30 + 10.5²/28) ≈ 2.475, t ≈ 1.70 with Welch df ≈ 50.7, two-tailed p ≈ 0.096 > 0.05: fail to reject H₀. The 95% confidence interval for the difference, −0.77 to 9.17 points, includes 0 but also differences of several points, so the study is inconclusive rather than showing no difference.
Paired: scores before and after training (hypothetical)t ≈ 4.49, df = 7, p ≈ 0.0028Differences (After − Before): 4, 3, 3, −1, 6, 6, 3, 4. d̄ = 3.5, s_d ≈ 2.20, SE ≈ 0.779, t ≈ 4.49 with df = 7, two-tailed p ≈ 0.0028, so reject H₀: scores tend to be higher after training. The 95% CI for the mean improvement is 1.66 to 5.34 points. Without a control group, this does not show the training caused the improvement.

What Is a T-Test?

A t-test is a hypothesis test for means. It divides the difference between an estimate and the value stated in the null hypothesis by the standard error of that estimate, giving the t-statistic, and then asks how unusual that value would be under Student's t-distribution. Unlike a z-test, it uses the sample standard deviation, and the t-distribution's heavier tails allow for the extra uncertainty of estimating the spread from the data.

Which T-Test Should You Use?

Use a one-sample t-test when you have one sample and want to compare its mean with a reference or target value, such as whether average delivery time exceeds a 30-minute target.

Use an independent two-sample t-test when you have two separate groups made up of different subjects, such as students taught by two different methods. Welch's version is the safer default because it does not assume equal variances; use the pooled version only when equal population variances are a reasonable assumption. Deciding by first running a variance test is not recommended.

Use a paired t-test when every value in one set has a natural partner in the other: the same people measured before and after, the same product tested under two conditions, or matched pairs. The test works on the differences within each pair, which removes the variation between subjects. Treating paired data as independent groups throws that advantage away.

How to Calculate a T-Test

One-sample: find the sample mean x̄ and standard deviation s, compute the standard error SE = s / √n, then t = (x̄ − μ₀) / SE with df = n − 1. For the sample 12, 15, 14, 18, 16, 13, 17 and μ₀ = 14: x̄ = 15, s ≈ 2.160, SE ≈ 0.816, t ≈ 1.225, df = 6, two-tailed p ≈ 0.267.

Independent (Welch): compute each group's mean and variance, the difference x̄₁ − x̄₂, SE = √(s₁²/n₁ + s₂²/n₂), t = difference / SE and the Welch–Satterthwaite degrees of freedom. For 10, 12, 15, 14, 13 against 18, 17, 20, 19, 21: difference −6.2, SE ≈ 1.114, t ≈ −5.568, df ≈ 7.71, p ≈ 0.0006.

Paired: compute each difference d = After − Before, then run a one-sample t-test on the differences against 0. For Before 10, 12, 15, 13, 16 and After 12, 15, 16, 15, 18: differences 2, 3, 1, 2, 2, d̄ = 2, s_d ≈ 0.707, SE ≈ 0.316, t ≈ 6.325, df = 4, p ≈ 0.0032.

How to Interpret T-Test Results

Compare the p-value with the significance level α chosen in advance. If p < α, reject the null hypothesis: the result is statistically significant at that level. If p ≥ α, fail to reject it: the data do not provide enough evidence of a difference. Failing to reject is not the same as showing there is no difference, and nothing is ever proved true or false.

A two-tailed test looks for a difference in either direction; a one-tailed test looks in one direction only and must be chosen before seeing the data. The confidence interval gives the range of plausible values for the mean or difference; when a two-sided interval excludes the null value, the two-tailed test rejects at the same α.

T-Test Assumptions

The outcome is quantitative (measured on a numerical scale).

Observations are independent: within each group for the one-sample and independent tests, and between pairs for the paired test. The two groups in an independent test must be unrelated.

For small samples, the data (or, for a paired test, the differences) should be roughly normal, without severe outliers. The t-test is fairly robust to mild non-normality, and larger samples make it more so, but strong skew and outliers can still distort it.

The pooled two-sample test assumes equal population variances; Welch's test does not.

The calculator cannot check how your data were collected. It flags values beyond Tukey's fences (1.5 × IQR past the quartiles) for review, but never removes them.

Statistical Significance vs Practical Significance

A small p-value says the data are hard to explain by chance alone; it does not say the difference is large or important. With a big enough sample, a difference of a fraction of a point can be significant. Look at the estimated difference and its confidence interval in the units you care about, at an effect size such as Cohen's d, and at the context: what size of difference would actually matter? Cohen's d values of 0.2, 0.5 and 0.8 are often called small, medium and large, but these are rough conventions, not rules.

If you run many t-tests on the same data, the chance that at least one is significant by chance alone increases. For more than two groups, use ANOVA rather than many pairwise t-tests.

T-Test vs Z-Test and ANOVA

A z-test uses the standard normal distribution and a known population standard deviation, or a large-sample normal approximation in tests such as those for proportions. A t-test uses the sample standard deviation and the t-distribution. It is not simply a matter of sample size: for means with an unknown standard deviation the t-test is correct at any sample size, and it becomes almost identical to the z-test as the degrees of freedom grow.

A t-test compares one mean with a value or two means with each other. ANOVA compares the means of three or more groups in a single test, which controls the false-positive rate better than running every pairwise t-test.

How to Use the T-Test Calculator

  1. Choose the test: one-sample (one mean against a reference value), independent two-sample (two separate groups) or paired (two measurements on the same subjects or matched pairs).
  2. Enter raw data, one list per group, or switch to summary statistics and enter the means, standard deviations and sample sizes.
  3. Choose the alternative hypothesis (two-tailed, left-tailed or right-tailed) and the significance level α; for two groups, choose Welch’s test (recommended) or the pooled equal-variance test.
  4. Read the t-statistic, degrees of freedom, p-value and decision, with the hypotheses that were tested, the confidence interval for the mean or difference, and Cohen’s d.
  5. Check the data summary and chart for unusual values, and follow the step-by-step calculation.

Frequently Asked Questions

What is a t-test?

A hypothesis test for means: it compares a sample mean with a value, or two means with each other, using the t-statistic and Student’s t-distribution.

When should I use a t-test?

When the outcome is numerical and you want to compare one mean with a reference value, two independent group means, or paired measurements, and the population standard deviation is unknown.

What is a one-sample t-test?

It tests whether a population mean differs from a hypothesized value μ₀: t = (x̄ − μ₀) / (s/√n) with df = n − 1.

What is an independent two-sample t-test?

It compares the means of two separate groups of different subjects. Welch’s version, the recommended default, does not assume equal variances; the pooled version does.

What is a paired t-test?

It compares two measurements on the same subjects or matched pairs by testing whether the mean of the within-pair differences is 0: t = d̄ / (s_d/√n), df = n − 1.

What is Welch’s t-test?

A two-sample t-test that uses each group’s own variance, SE = √(s₁²/n₁ + s₂²/n₂), and the Welch–Satterthwaite degrees of freedom, which are often not a whole number. It does not assume equal variances.

What is the t-statistic?

The estimate minus the hypothesized value, divided by its standard error. It measures how many standard errors the observed mean or difference is from the null value.

How do you calculate a t-test p-value?

From the t-distribution with the test’s degrees of freedom: for a two-tailed test, p = 2 × P(T > |t|); for a right-tailed test, P(T > t); for a left-tailed test, P(T < t).

What are the degrees of freedom in a t-test?

n − 1 for one-sample and paired tests, n₁ + n₂ − 2 for the pooled two-sample test, and the Welch–Satterthwaite value for Welch’s test. They set the shape of the t-distribution.

What does p < 0.05 mean in a t-test?

At the 0.05 significance level you reject the null hypothesis: a difference this large would be unlikely if the null hypothesis were true. It does not say the difference is large or important.

What is the difference between a t-test and a z-test?

A t-test uses the sample standard deviation and the t-distribution; a z-test uses a known population standard deviation or a large-sample normal approximation and the normal distribution.

Does a significant t-test prove causation?

No. It shows evidence of a difference in means. Whether one thing caused it depends on the study design, for example random assignment to groups.

What is the difference between statistical and practical significance?

Statistical significance means the difference is unlikely to be due to chance; practical significance means it is large enough to matter. Judge the latter from the effect size, the confidence interval and the context.

Last updated: September 27, 2026.