Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws

Gerard Ben Arous, Murat A Erdogdu, Nuri Mert Vural, Denny Wu

New York University· University of Toronto· Vector Institute· Flatiron Institute

scaling laws stochastic gradient descent shallow neural network multi-index model

Abstract

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as

y \propto \sum_{j = 1}^{r} λ_{j} σ (⟨ θ_{j}, x ⟩), x \sim N (0, I_{d})

, where

σ

is the 2nd Hermite polynomial, and

{θ_{j}}_{j = 1}^{r} \subset R^{d}

are orthonormal signal directions.

We consider the extensive-width regime

r ≍ d^{β}

for

β \in (0, 1)

, and assume a power-law decay on the (non-negative) second-layer coefficients

λ_{j} ≍ j^{- α}

for

α \geq 0

We provide a sharp analysis of the SGD dynamics in the feature learning regime, for both the population limit and the finite-sample (online) discretization, and derive scaling laws for the prediction risk that highlight the power-law dependencies on the optimization time, the sample size, and the model width.

Our analysis combines a precise characterization of the associated matrix Riccati differential equation with novel matrix monotonicity arguments to establish convergence guarantees for the infinite-dimensional effective dynamics.