Method · 5 min read · 2026-10-06

How many requests before a success rate means something

The arithmetic behind sample sizes: how wide the uncertainty on a success rate is at 30, 100 and 400 requests, and how to compute it yourself.

A success rate is a proportion measured on a sample, so it carries uncertainty. Two providers at 82% and 85% on thirty requests each have not been ranked; the difference is well inside the noise. This note shows the arithmetic so you can size a test before you run it.

The simple approximation

For a proportion p measured on n independent requests, the standard error is the square root of p(1-p)/n, and a 95% interval is about 1.96 times that either side of p. It is an approximation that works best when p is not near 0 or 1 and n is not tiny.

python
from math import sqrt

def half_width(p: float, n: int) -> float:
    return 1.96 * sqrt(p * (1 - p) / n)

for n in (30, 100, 400):
    print(n, round(half_width(0.8, n) * 100, 1))

At a measured 80%, that gives roughly 14 percentage points either side at 30 requests, 8 at 100 and 4 at 400. To tell two providers apart by a few points you need hundreds of requests per target for each, and more as the gap shrinks.

What the approximation hides

  • Requests are not independent. Retries on one address, or bursts in one minute, correlate. Spread the test over time and addresses.
  • Targets differ. Pool results across targets only if you will pool them in production; otherwise report each separately.
  • A rate near 100% or 0% needs a better interval, such as the Wilson interval, because the normal approximation can run past the ends.

A practical rule

Decide the smallest difference that would change your decision, work out the sample size that resolves it, and fix that number before you start. Then publish it beside the result, as the method page requires.

Test it on your own targets.

Create an account, run the measurement harness against your real workload, and read the numbers before you commit to anything.