Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
Publish this paper in The AIPR Journal
Are you an author? Turn this AI review into a permanent, citable journal entry with a cover, open comments, and Scholar metadata.
AIPR assessment
This is a hard, competitive problem space: continual learning under adaptive optimizers is a crowded area, and the paper tries to extract a non-obvious failure mode rather than just improve a benchmark by brute force. The strengths compound well, because a simple repair, broad ablations, low overhead, and cross-family consistency all point in the same direction. The weaknesses also compound in a milder way, because the key mechanism is partly diagnostic, the adaptive schedule is heuristic, and t
Abstract
Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition has a hidden failure mode. In a high-overlap, non-adaptive 8-domain continual LM, all shared-routing projection baselines collapse close to vanilla forgetting (12.5--12.8 vs. 13.2). A 0.5% replay buffer is the strongest shared alternative but still reaches 11.6, while fixed-strength decoupling falls below vanilla at 14.1. Only adaptive decoupled routing remains stable at 9.4, improving over vanilla by 3.8 units. On a 16-domain stream, its gain over the strongest shared-routing projection baseline grows to 4.5--4.8 units. The failure is largely invisible on clean benchmarks. We explain this effect through Adam's second-moment pathway: in the tested regime, projection induces a 1/(1-alpha) inflation of the old-direction effective learning rate, matching measurements within 8% across eight alpha values. The same conflict appears with penalty methods, replay mixing, and at 7B scale under LoRA. Our fix routes the modified gradient only to the first moment while preserving magnitude-faithful second-moment statistics, with overlap-aware adaptive strength. This simple change is the only tested configuration that consistently avoids collapse across methods, optimizers, and scale.
Score Breakdown
More from this week
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
- How Hard is it to Decide if a Fact is Relevant to a Query?
- Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain
- On first-order model checking parameterized by the number of variables
- Entrywise Low-Rank Approximation and Matrix \(p \rightarrow q\) Norms via Global Correlation Rounding