Part 5

16 min read6 headingsSplit lesson page

Lesson overview | Previous part | Next part

Online Experimentation and AB Testing: Part 5: Sequential Testing

5. Sequential Testing

Sequential Testing is the part of online experimentation and ab testing that turns the approved TOC into a concrete learning path. The subsections below keep the focus on Chapter 17's canonical job: measurement, reliability, uncertainty, and decision support for AI systems.

5.1 Peeking risk

Peeking risk is part of the canonical scope of online experimentation and ab testing. In this chapter, the object under study is not merely a dataset or a model, but the full online randomized experiment: the items, prompts, outputs, graders, uncertainty statements, and decision rules that turn model behavior into evidence.

The basic mathematical pattern is an empirical estimator. For a model or system $m$ evaluated on items $z_1,\ldots,z_n$ , the local estimate is written

\widehat{\operatorname{ATE}} = \frac{1}{n}\sum_{i=1}^n Y_i(1)-Y_i(0).

The formula is intentionally simple. The difficulty lies in deciding what counts as an item, which loss or score is meaningful, whether the items are independent, and whether the estimate answers the real product or research question. For peeking risk, those choices determine whether the reported number is evidence or decoration.

A useful invariant is that every evaluation claim should be reproducible as a tuple $(m,\mathcal{T},\pi,g,\rho)$ , where $m$ is the system, $\mathcal{T}$ is the task sample, $\pi$ is the prompt or intervention policy, $g$ is the grader, and $\rho$ is the aggregation rule. If any part of this tuple is missing, the number cannot be audited.

Component	What to record	Why it matters
Item definition	IDs, source, split, and allowed transformations	Prevents accidental drift in peeking risk
Scoring rule	Exact formula for Y_i(1)-Y_i(0)	Makes comparisons repeatable
Aggregation	Mean, weighted mean, worst group, or pairwise model	Determines the scientific claim
Uncertainty	Standard error, interval, or posterior summary	Separates signal from sampling noise
Audit trail	Code version and random seeds	Makes failures debuggable

Examples of correct use:

Report peeking risk with item count, prompt protocol, grader version, and a confidence interval.
Use paired comparisons when two models answer the same evaluation items.
Inspect at least one meaningful slice before concluding that the aggregate result is reliable.
Store raw outputs so future graders can be replayed without querying the model again.
Document whether the metric is measuring capability, reliability, user value, or risk.

Non-examples:

A leaderboard point estimate without sample size.
A benchmark score produced with an undocumented prompt template.
A model-graded result without judge identity, rubric, or agreement check.
A robustness claim measured only on the easiest in-distribution examples.
An online win declared before the randomization and logging checks pass.

Worked evaluation pattern for peeking risk:

Define the evaluation population in words before writing code.
Choose the smallest metric set that answers the decision question.
Compute the point estimate and an uncertainty statement together.
Run a slice or paired analysis to check whether the aggregate hides structure.
Archive raw outputs, scores, and seeds before changing the prompt or grader.

For AI systems, peeking risk is especially delicate because the same model can be used with many prompts, decoding policies, tools, retrieval contexts, and safety filters. The measured quantity is therefore a property of the system configuration, not just the base weights.

AI connection	Evaluation consequence
Prompting	Treat prompt templates as part of the protocol, not as invisible setup
Decoding	Temperature and sampling change both mean score and variance
Retrieval	Retrieved context creates an extra source of failure and leakage
Tool use	Tool errors need separate attribution from model reasoning errors
Safety layer	Guardrail behavior can improve risk metrics while changing capability metrics

Implementation checklist:

Use deterministic seeds for synthetic or sampled evaluation subsets.
Print metric denominators, not only percentages.
Keep missing, invalid, timeout, and refusal outcomes explicit.
Prefer typed result records over loose CSV columns.
Separate raw model outputs from normalized grader inputs.
Track the smallest reproducible command that generated the result.
Record whether the estimate is item-weighted, token-weighted, user-weighted, or domain-weighted.
Write the decision rule before seeing the final score whenever the result will guide a release.

The mathematical habit to build is skepticism with structure. A score is not ignored because it is noisy; it is interpreted through the design that produced it. Peeking risk is one place where that habit becomes concrete.

5.2 Alpha spending

Alpha spending is part of the canonical scope of online experimentation and ab testing. In this chapter, the object under study is not merely a dataset or a model, but the full online randomized experiment: the items, prompts, outputs, graders, uncertainty statements, and decision rules that turn model behavior into evidence.

The basic mathematical pattern is an empirical estimator. For a model or system $m$ evaluated on items $z_1,\ldots,z_n$ , the local estimate is written

\widehat{\operatorname{ATE}} = \frac{1}{n}\sum_{i=1}^n Y_i(1)-Y_i(0).

The formula is intentionally simple. The difficulty lies in deciding what counts as an item, which loss or score is meaningful, whether the items are independent, and whether the estimate answers the real product or research question. For alpha spending, those choices determine whether the reported number is evidence or decoration.

Component	What to record	Why it matters
Item definition	IDs, source, split, and allowed transformations	Prevents accidental drift in alpha spending
Scoring rule	Exact formula for Y_i(1)-Y_i(0)	Makes comparisons repeatable
Aggregation	Mean, weighted mean, worst group, or pairwise model	Determines the scientific claim
Uncertainty	Standard error, interval, or posterior summary	Separates signal from sampling noise
Audit trail	Code version and random seeds	Makes failures debuggable

Examples of correct use:

Report alpha spending with item count, prompt protocol, grader version, and a confidence interval.
Use paired comparisons when two models answer the same evaluation items.
Inspect at least one meaningful slice before concluding that the aggregate result is reliable.
Store raw outputs so future graders can be replayed without querying the model again.
Document whether the metric is measuring capability, reliability, user value, or risk.

Non-examples:

A leaderboard point estimate without sample size.
A benchmark score produced with an undocumented prompt template.
A model-graded result without judge identity, rubric, or agreement check.
A robustness claim measured only on the easiest in-distribution examples.
An online win declared before the randomization and logging checks pass.

Worked evaluation pattern for alpha spending:

Define the evaluation population in words before writing code.
Choose the smallest metric set that answers the decision question.
Compute the point estimate and an uncertainty statement together.
Run a slice or paired analysis to check whether the aggregate hides structure.
Archive raw outputs, scores, and seeds before changing the prompt or grader.

For AI systems, alpha spending is especially delicate because the same model can be used with many prompts, decoding policies, tools, retrieval contexts, and safety filters. The measured quantity is therefore a property of the system configuration, not just the base weights.

AI connection	Evaluation consequence
Prompting	Treat prompt templates as part of the protocol, not as invisible setup
Decoding	Temperature and sampling change both mean score and variance
Retrieval	Retrieved context creates an extra source of failure and leakage
Tool use	Tool errors need separate attribution from model reasoning errors
Safety layer	Guardrail behavior can improve risk metrics while changing capability metrics

Implementation checklist:

Use deterministic seeds for synthetic or sampled evaluation subsets.
Print metric denominators, not only percentages.
Keep missing, invalid, timeout, and refusal outcomes explicit.
Prefer typed result records over loose CSV columns.
Separate raw model outputs from normalized grader inputs.
Track the smallest reproducible command that generated the result.
Record whether the estimate is item-weighted, token-weighted, user-weighted, or domain-weighted.
Write the decision rule before seeing the final score whenever the result will guide a release.

The mathematical habit to build is skepticism with structure. A score is not ignored because it is noisy; it is interpreted through the design that produced it. Alpha spending is one place where that habit becomes concrete.

5.3 Always-valid p-values

Always-valid p-values is part of the canonical scope of online experimentation and ab testing. In this chapter, the object under study is not merely a dataset or a model, but the full online randomized experiment: the items, prompts, outputs, graders, uncertainty statements, and decision rules that turn model behavior into evidence.

The basic mathematical pattern is an empirical estimator. For a model or system $m$ evaluated on items $z_1,\ldots,z_n$ , the local estimate is written

\widehat{\operatorname{ATE}} = \frac{1}{n}\sum_{i=1}^n Y_i(1)-Y_i(0).

The formula is intentionally simple. The difficulty lies in deciding what counts as an item, which loss or score is meaningful, whether the items are independent, and whether the estimate answers the real product or research question. For always-valid p-values, those choices determine whether the reported number is evidence or decoration.

Component	What to record	Why it matters
Item definition	IDs, source, split, and allowed transformations	Prevents accidental drift in always-valid p-values
Scoring rule	Exact formula for Y_i(1)-Y_i(0)	Makes comparisons repeatable
Aggregation	Mean, weighted mean, worst group, or pairwise model	Determines the scientific claim
Uncertainty	Standard error, interval, or posterior summary	Separates signal from sampling noise
Audit trail	Code version and random seeds	Makes failures debuggable

Examples of correct use:

Report always-valid p-values with item count, prompt protocol, grader version, and a confidence interval.
Use paired comparisons when two models answer the same evaluation items.
Inspect at least one meaningful slice before concluding that the aggregate result is reliable.
Store raw outputs so future graders can be replayed without querying the model again.
Document whether the metric is measuring capability, reliability, user value, or risk.

Non-examples:

A leaderboard point estimate without sample size.
A benchmark score produced with an undocumented prompt template.
A model-graded result without judge identity, rubric, or agreement check.
A robustness claim measured only on the easiest in-distribution examples.
An online win declared before the randomization and logging checks pass.

Worked evaluation pattern for always-valid p-values:

Define the evaluation population in words before writing code.
Choose the smallest metric set that answers the decision question.
Compute the point estimate and an uncertainty statement together.
Run a slice or paired analysis to check whether the aggregate hides structure.
Archive raw outputs, scores, and seeds before changing the prompt or grader.

For AI systems, always-valid p-values is especially delicate because the same model can be used with many prompts, decoding policies, tools, retrieval contexts, and safety filters. The measured quantity is therefore a property of the system configuration, not just the base weights.

AI connection	Evaluation consequence
Prompting	Treat prompt templates as part of the protocol, not as invisible setup
Decoding	Temperature and sampling change both mean score and variance
Retrieval	Retrieved context creates an extra source of failure and leakage
Tool use	Tool errors need separate attribution from model reasoning errors
Safety layer	Guardrail behavior can improve risk metrics while changing capability metrics

Implementation checklist:

Use deterministic seeds for synthetic or sampled evaluation subsets.
Print metric denominators, not only percentages.
Keep missing, invalid, timeout, and refusal outcomes explicit.
Prefer typed result records over loose CSV columns.
Separate raw model outputs from normalized grader inputs.
Track the smallest reproducible command that generated the result.
Record whether the estimate is item-weighted, token-weighted, user-weighted, or domain-weighted.
Write the decision rule before seeing the final score whenever the result will guide a release.

The mathematical habit to build is skepticism with structure. A score is not ignored because it is noisy; it is interpreted through the design that produced it. Always-valid p-values is one place where that habit becomes concrete.

5.4 Stopping rules

Stopping rules is part of the canonical scope of online experimentation and ab testing. In this chapter, the object under study is not merely a dataset or a model, but the full online randomized experiment: the items, prompts, outputs, graders, uncertainty statements, and decision rules that turn model behavior into evidence.

The basic mathematical pattern is an empirical estimator. For a model or system $m$ evaluated on items $z_1,\ldots,z_n$ , the local estimate is written

\widehat{\operatorname{ATE}} = \frac{1}{n}\sum_{i=1}^n Y_i(1)-Y_i(0).

The formula is intentionally simple. The difficulty lies in deciding what counts as an item, which loss or score is meaningful, whether the items are independent, and whether the estimate answers the real product or research question. For stopping rules, those choices determine whether the reported number is evidence or decoration.

Component	What to record	Why it matters
Item definition	IDs, source, split, and allowed transformations	Prevents accidental drift in stopping rules
Scoring rule	Exact formula for Y_i(1)-Y_i(0)	Makes comparisons repeatable
Aggregation	Mean, weighted mean, worst group, or pairwise model	Determines the scientific claim
Uncertainty	Standard error, interval, or posterior summary	Separates signal from sampling noise
Audit trail	Code version and random seeds	Makes failures debuggable

Examples of correct use:

Report stopping rules with item count, prompt protocol, grader version, and a confidence interval.
Use paired comparisons when two models answer the same evaluation items.
Inspect at least one meaningful slice before concluding that the aggregate result is reliable.
Store raw outputs so future graders can be replayed without querying the model again.
Document whether the metric is measuring capability, reliability, user value, or risk.

Non-examples:

A leaderboard point estimate without sample size.
A benchmark score produced with an undocumented prompt template.
A model-graded result without judge identity, rubric, or agreement check.
A robustness claim measured only on the easiest in-distribution examples.
An online win declared before the randomization and logging checks pass.

Worked evaluation pattern for stopping rules:

Define the evaluation population in words before writing code.
Choose the smallest metric set that answers the decision question.
Compute the point estimate and an uncertainty statement together.
Run a slice or paired analysis to check whether the aggregate hides structure.
Archive raw outputs, scores, and seeds before changing the prompt or grader.

For AI systems, stopping rules is especially delicate because the same model can be used with many prompts, decoding policies, tools, retrieval contexts, and safety filters. The measured quantity is therefore a property of the system configuration, not just the base weights.

AI connection	Evaluation consequence
Prompting	Treat prompt templates as part of the protocol, not as invisible setup
Decoding	Temperature and sampling change both mean score and variance
Retrieval	Retrieved context creates an extra source of failure and leakage
Tool use	Tool errors need separate attribution from model reasoning errors
Safety layer	Guardrail behavior can improve risk metrics while changing capability metrics

Implementation checklist:

Use deterministic seeds for synthetic or sampled evaluation subsets.
Print metric denominators, not only percentages.
Keep missing, invalid, timeout, and refusal outcomes explicit.
Prefer typed result records over loose CSV columns.
Separate raw model outputs from normalized grader inputs.
Track the smallest reproducible command that generated the result.
Record whether the estimate is item-weighted, token-weighted, user-weighted, or domain-weighted.
Write the decision rule before seeing the final score whenever the result will guide a release.

The mathematical habit to build is skepticism with structure. A score is not ignored because it is noisy; it is interpreted through the design that produced it. Stopping rules is one place where that habit becomes concrete.

5.5 Multiple online tests

Multiple online tests is part of the canonical scope of online experimentation and ab testing. In this chapter, the object under study is not merely a dataset or a model, but the full online randomized experiment: the items, prompts, outputs, graders, uncertainty statements, and decision rules that turn model behavior into evidence.

The basic mathematical pattern is an empirical estimator. For a model or system $m$ evaluated on items $z_1,\ldots,z_n$ , the local estimate is written

\widehat{\operatorname{ATE}} = \frac{1}{n}\sum_{i=1}^n Y_i(1)-Y_i(0).

The formula is intentionally simple. The difficulty lies in deciding what counts as an item, which loss or score is meaningful, whether the items are independent, and whether the estimate answers the real product or research question. For multiple online tests, those choices determine whether the reported number is evidence or decoration.

Component	What to record	Why it matters
Item definition	IDs, source, split, and allowed transformations	Prevents accidental drift in multiple online tests
Scoring rule	Exact formula for Y_i(1)-Y_i(0)	Makes comparisons repeatable
Aggregation	Mean, weighted mean, worst group, or pairwise model	Determines the scientific claim
Uncertainty	Standard error, interval, or posterior summary	Separates signal from sampling noise
Audit trail	Code version and random seeds	Makes failures debuggable

Examples of correct use:

Report multiple online tests with item count, prompt protocol, grader version, and a confidence interval.
Use paired comparisons when two models answer the same evaluation items.
Inspect at least one meaningful slice before concluding that the aggregate result is reliable.
Store raw outputs so future graders can be replayed without querying the model again.
Document whether the metric is measuring capability, reliability, user value, or risk.

Non-examples:

A leaderboard point estimate without sample size.
A benchmark score produced with an undocumented prompt template.
A model-graded result without judge identity, rubric, or agreement check.
A robustness claim measured only on the easiest in-distribution examples.
An online win declared before the randomization and logging checks pass.

Worked evaluation pattern for multiple online tests:

Define the evaluation population in words before writing code.
Choose the smallest metric set that answers the decision question.
Compute the point estimate and an uncertainty statement together.
Run a slice or paired analysis to check whether the aggregate hides structure.
Archive raw outputs, scores, and seeds before changing the prompt or grader.

For AI systems, multiple online tests is especially delicate because the same model can be used with many prompts, decoding policies, tools, retrieval contexts, and safety filters. The measured quantity is therefore a property of the system configuration, not just the base weights.

AI connection	Evaluation consequence
Prompting	Treat prompt templates as part of the protocol, not as invisible setup
Decoding	Temperature and sampling change both mean score and variance
Retrieval	Retrieved context creates an extra source of failure and leakage
Tool use	Tool errors need separate attribution from model reasoning errors
Safety layer	Guardrail behavior can improve risk metrics while changing capability metrics

Implementation checklist:

Use deterministic seeds for synthetic or sampled evaluation subsets.
Print metric denominators, not only percentages.
Keep missing, invalid, timeout, and refusal outcomes explicit.
Prefer typed result records over loose CSV columns.
Separate raw model outputs from normalized grader inputs.
Track the smallest reproducible command that generated the result.
Record whether the estimate is item-weighted, token-weighted, user-weighted, or domain-weighted.
Write the decision rule before seeing the final score whenever the result will guide a release.

The mathematical habit to build is skepticism with structure. A score is not ignored because it is noisy; it is interpreted through the design that produced it. Multiple online tests is one place where that habit becomes concrete.

Online Experimentation and AB Testing: Part 5 - Sequential Testing

Online Experimentation and AB Testing: Part 5: Sequential Testing

5. Sequential Testing

5.1 Peeking risk

5.2 Alpha spending

5.3 Always-valid p-values

5.4 Stopping rules

5.5 Multiple online tests

Test this lesson

Which module does this lesson belong to?

Which section is covered in this lesson content?

Which term is most central to this lesson?

What is the best way to use this lesson for real learning?