1.1 Why a Single-Arm Phase II Trial?
Cancer clinical trials investigate the efficacy and toxicity of experimental cancer therapies through four phases of clinical trials, designated phase I, II, III and IV. A phase III cancer clinical trial is usually a large-scale randomized study to compare the efficacy of a new or experimental treatment to that of the best current standard treatment or placebo, while a phase II cancer clinical trial is often conducted as a single-arm study to determine whether a new treatment has sufficient antitumor activities to warrant further investigation in a large-scale randomized phase III trial. However, a single-arm trial compares to a reference group from historical data and doesn't have a concurrent control group. Thus, the results from a single-arm phase II trial could be biased. However, a single-arm trial requires relatively small sample size and all patients are treated with the new treatment. Furthermore for patients with a rare life-threatening disease, the use of a control arm may be unethical or unfeasible. Therefore, single-arm phase II trials are still frequently used in cancer studies to provide some preliminary efficacy and toxicity assessment for new or experimental treatments.
1.2 Primary Endpoint of a Single-Arm Phase II Trial
The antitumor activity for evaluating cytotoxic compounds might be quantified by tumor response which is often categorized to be a binary endpoint as responder if the patient achieved a complete response (CR) or partial response (PR) or non-responder if the patient had stable disease (SD) or progressive disease (PD), where the response evaluation criteria is often based on response evaluation criteria in solid tumor (RECIST). For the trial design with tumor response rate as the primary endpoint, investigators identify the response rate (p0) (CR or PR) of the standard treatment from historical or literature data as the null hypothesis, e.g., . Therefore, for single-arm phase II trial design, p0 is assumed known and is not subject to variation even though it may be estimated from historical or literature data or a reference level selected by the investigators. This is a fundamental difference between a single-arm phase II trial and a randomized phase II trial.
1.3 Hypothesis of a Single-Arm Phase II Trial
For a single-arm phase II trial with tumor response rate as the primary endpoint, the research hypothesis is if the new or experimental treatment can improve the tumor response rate for a target patient group. Therefore the research hypothesis is typically set to be as a one-sided hypothesis
where p is the true response rate of the new treatment which is unknown. If the true response rate of the new treatment is less than or equal to p0, the new treatment is not promising for further investigation and if the true response rate of the new treatment is greater than p0, the new treatment is promising for further investigation in the future in a phase III trial.
1.4 Test Statistic and Study Design
To calculate the required sample size, investigators have to choose a response rate of the new treatment (p1), e.g., , where is a minimum clinical meaningful effect size to detect for the sample size calculation. For the study design, typically, the reference response rate p0 is chosen to be a value of the probability of tumor response to the standard treatment for the same disease group, and p1 or effect size is selected to identify a minimum clinically meaningful improvement over the historical value. To design the study, we have to select a test statistic to make the inference to the research hypothesis, e.g., reject null hypothesis H0. A nature test statistic is which is an estimated probability of response rate for the experimental treatment, where X is the number of subjects who experienced response (CR or PR) and n is the total number of subjects in the study. However the distribution of is often skewed when sample size is small. To make the distribution more normal, we take an arcsin square root transformation to be . Using delta method, it can be shown that
then, the standardized test statistic Z is given by
which is approximately standard normal distributed under the null hypothesis. A large observed value of Z indicates a large treatment effect, therefore, given a type I error rate α, we reject null hypothesis H0 if , where and is the cumulative distribution function (CDF) of standard normal.
To design a study, we need to calculate the sample size (n) or number of subjects required to have sufficient statistical power () to detect the treatment effect. The statistical power is the probability to reject null hypothesis H0 when the alternative hypothesis H1 is true. Under the alternative hypothesis , the test Z is approximately normal with mean and unit variance. Thus, the study power satisfies the following equation:
By inverting the CDF from the above equation, we have
Solving n, we obtain the following sample size formula for the Z test