Zixuan Zhao

Research

My research projects

I mainly work on three questions.

Discretize the covariate into strata, or balance directly?

Two baseline covariates, an unadjusted treatment-effect estimate, 1,000 simulated trials per design.

1 Covariate-adaptive randomization

Discretization in covariate-adaptive randomization: gains and losses

Zixuan Zhao, Feifang Hu

Major revision, Journal of the American Statistical Association

Trials try to spread patient characteristics evenly across treatment arms. Age is continuous, but in practice trials discretize it into bins (e.g., 'smaller than 50', '50 to 65', 'larger than 65'). This paper asks what the discretization costs and what it buys. Discretization turns out to be the robust default. It protects when your assumption about how age affects the outcome is wrong. But if you genuinely know that relationship, say linear, balancing on the continuous age directly gives more precision from the same number of patients.

Full abstract

Covariate-adaptive randomization (CAR) is widely implemented in clinical trials to balance prognostic covariates across treatment arms. Continuous covariates are often discretized into strata in practice, yet its consequences are not clearly understood. This paper provides a comprehensive overview on the impact of discretization, on both the CAR design process and the inferential results thereafter. We establish the asymptotic properties of both imbalance measures and treatment effect estimators under discretized and non-discretized settings. Practical recommendations are given on when and how discretization should be employed. We show that discretization in design is generally recommended, as it enhances robustness against model misspecification. However, if the true model is known, the most efficient strategy is to balance covariates according to that model in the design. The theoretical results are corroborated by extensive simulation studies and an empirical application to a diabetes trial dataset. Together, the results clarify the gains and losses of discretization in CAR and pave the way for learning impact of discretization to other designs and beyond.

Getting patient data back out of a picture

A trial is simulated, its survival curve is "published", and only that curve plus the numbers-at-risk table is handed to the reconstruction. Everything runs live.

2 Synthetic patient data

SynthIPD: training-free synthetic individual patient data generation

Zixuan Zhao*, Zexin Ren*, Guannan Zhai, Feifang Hu, Will Ma, En Xie, Qian Shi — *equal contribution

Major revision, Journal of the American Statistical Association

Trial results are published as summary statistics, and usually no individual patient data is disclosed. SynthIPD reads the published survival curve, recovering each patient's event from the image, then fills in patient characteristics so the reconstructed dataset reproduces the trial's report faithfully. Existing generative modeling approaches need a large real dataset to learn from first, which is not always available. This one needs none.

Full abstract

Individual patient data (IPD) are essential for statistical analysis in clinical research, yet access is often limited due to privacy concerns, high data-sharing costs, and proprietary restrictions. Conventional approaches to synthetic data generation, such as generative adversarial networks (GANs), require a large piece of IPD as training set. We propose a training-free, three-step method to create synthetic IPD for survival data, requiring no IPD for training. In summary, it digitizes the reported Kaplan–Meier (KM) plots and generates covariates that match the reconstructed data. Compared with existing IPD reconstruction methods, our approach is the first to exploit Scalable Vector Graphics (SVG) for high-accuracy digitization, and the first to provide covariate information. We demonstrate the method's potential through two detailed case studies and complementary simulation studies.

Bootstrap covariance estimation under minimization

The heat map shows the bootstrap estimate minus its target; the second panel shows test size at zero treatment effect and power otherwise.

3 Covariate-adaptive randomization

Consistent covariances estimation for stratum imbalances under minimization method for covariate-adaptive randomization

Zixuan Zhao, Yanglei Song, Wenyu Jiang, Dongsheng Tu

Scandinavian Journal of Statistics, 2024, 51(3), 861–890

Minimization is one of the most widely used ways to assign patients to arms, but it introduces conservative and sometimes inflated statistical tests afterwards. Repairing this needs a no-closed form quantity whose existence was proved only recently and which no one knew how to compute. We estimate it by bootstrap, prove the estimate is consistent, and show the corrected tests hit their nominal error rate.

Full abstract

Pocock and Simon's minimization method is a popular approach for covariate-adaptive randomization in clinical trials. Valid statistical inference with data collected under the minimization method requires the knowledge of the limiting covariance matrix of within-stratum imbalances, whose existence is only recently established. In this work, we propose a bootstrap-based estimator for this limit and establish its consistency, in particular, by Le Cam's third lemma. As an application, we consider in simulation studies adjustments to existing robust tests for treatment effects with survival data by the proposed estimator. It shows that the adjusted tests achieve a size close to the nominal level, and unlike other designs, the robust tests without adjustment may have an asymptotic size inflation issue under the minimization method.

What several disagreeing agents buy you

Results as reported in the paper: 30 Monte Carlo runs on the NSCLC (n=2236) benchmark and the reproduction of a published network meta-analysis. Hover any point.

4 AI in clinical trials: evidence synthesis

Systematic literature reviews with two multi-agentic systems and human-in-the-loop

Zexin Ren, Zixuan Zhao, Qiyun Li, Yawen Wu, Lanjing Wang, Renjie Luo, Yi Xu, Qing Guo, Jin Shi, En Xie, Feifang Hu, Qian Shi

Submitted · arXiv preprint

To design a clinical trial, you have to establish a benchmark estimate for the current standard of care. Conventional meta-analyses require up to a year of heavy manual work. We built two systems of cooperating AI agents: one screens papers using several agents given deliberately different perspectives that cross-check each other over multiple rounds; the other extracts the numbers through a correction loop. Human stay in the loop by design. The accuracy metrics are evaluated for both systems. As a real application, we reproduced a published meta-analysis. The system recovered every trial the original authors found plus eligible ones they had missed.

Full abstract

Systematic literature review of clinical trials drives regulatory decision-making, but conventional screening and extraction are time-consuming, labor-intensive, and vulnerable to study selection bias. We propose two fit-to-purpose multi-agentic systems (MAS) for systematic literature review, with human-in-the-loop. The screening MAS uses multiple LLM agents with heterogeneous personas and multi-round cross-review, and uniformly improves accuracy over a single-LLM baseline. The extraction MAS combines standardization, an iterative correction loop, and retrieval-based context control to ensure accuracy and scalability. Both MAS are specifically designed to support Human-In-The-Loop which is essential for clinical decisions. The novelty of the proposed approach lies in the system architecture rather than in any single foundation tools: the system can naturally benefit from future improvements in the underlying tools, for instance, stronger LLM agents, retrieval engines, image recognition methods, etc. As a real-world application, a published network meta-analysis is reproduced by the MAS. The result recovers all trials from the original study and identifies additional eligible trials missed by manual review, leading to updated clinical conclusions.

Does a fast endpoint predict a slow one?

The paper's own data. Each bubble is one two-arm comparison; hover for the trial. Reported R² values and the copula odds ratios are transcribed from the paper's tables.

NDTE NDTI RRMM

5 AI in clinical trials: surrogate endpoints

Leveraging AI to evaluate minimal residual disease endpoint surrogacy in multiple myeloma

Zexin Ren, Zixuan Zhao, Andrew J. Cowan, Will Ma, En Xie, Qian Shi

Cancer Research Communications, 2026, 6(5), 1206–1215

Learning whether a myeloma drug extends survival takes years. Regulators now accept a faster signal (minimal residual disease, MRD), but the evidence behind that substitution needs curation. We used an AI pipeline to find and extract those trials automatically, evaluated the trial-level associations, and used synthetic patient data to investigate individual-level associations.

Full abstract

Minimal residual disease (MRD) has been endorsed by the FDA Oncology Drugs Advisory Committee as an endpoint for accelerated approval in multiple myeloma based on individual patient data collected from randomized trials. However, emerging data from recent trials were not included. A novel artificial intelligence (AI)–assisted framework is proposed, which automates information identification and extraction, providing up-to-date analyses that confirm moderate trial-level and strong individual patient–level associations between MRD rates at suspected complete response (MRD-CR) and survival endpoints in multiple myeloma. Specifically, this study utilized an AI-assisted framework that identifies relevant studies and filters critical information to analyze published data via 2 independent objectives. First, we examined the trial-level association using the coefficients of determination (R²) and its 95% confidence intervals (CI) based on published statistics of treatment effects on MRD and various endpoints. Next, we generated synthetic individual patient data with covariates through AI-curated tools to estimate the individual-level association. The AI tool searched for eligible randomized clinical trials (RCT). A total of 20 two-arm comparisons from 19 RCTs were analyzed. Trial-level analysis showed an R² of 0.71 (95% CI, 0.52–0.89) pooling disease subpopulations. Furthermore, AI techniques were applied to create synthetic individual data, combining information extracted from Kaplan–Meier curves and subgroup analyses from published literatures. Using generated synthetic data, we estimated the individual-level correlation between MRD-CR rates and progression-free survival outcomes with a bivariate copula model and calculated a global odds ratio of 7.28 (95% CI, 5.60–8.95).

6 Efficient treatment effect estimation for covariate-adaptive randomization with right-censored data Zixuan Zhao, Yanglei Song, Wenyu Jiang · Working paper