Localized Dudley Entropy Integral
Published:
Chaining across all resolution scales at once: process complexity is the integral of root-entropy, localized to a ball.
Published:
Chaining across all resolution scales at once: process complexity is the integral of root-entropy, localized to a ball.
Published:
Any function of independent inputs that no single input can move very far concentrates like a Gaussian.
Published:
The expected max of n sub-Gaussian variables costs only √(2 log n) — no independence required.
Published:
Variance-aware concentration for bounded random variables: the Gaussian rate until the range term takes over.
Published:
Principal Component Analysis from a linear algebraic perspecive.
Published:
Composing with a κ-Lipschitz function costs at most a factor κ in Rademacher complexity.
Published:
The ghost sample trick: uniform deviations are controlled by the class’s ability to correlate with random signs.
Published:
Chaining across all resolution scales at once: process complexity is the integral of root-entropy, localized to a ball.
Published:
Annotated study notes on Ma Chapter 5: margin theory, Rademacher complexity of linear models and two-layer nets, and the covering-number route to deep nets.
Published:
Why ‘learning is optimization’: no free lunch, the three-term excess risk decomposition, and the coupling that makes deep learning theory interesting.
Published:
Principal Component Analysis from a linear algebraic perspecive.
Published:
Why ‘learning is optimization’: no free lunch, the three-term excess risk decomposition, and the coupling that makes deep learning theory interesting.
Published:
Annotated study notes on Ma Chapter 5: margin theory, Rademacher complexity of linear models and two-layer nets, and the covering-number route to deep nets.
Published:
Composing with a κ-Lipschitz function costs at most a factor κ in Rademacher complexity.
Published:
The ghost sample trick: uniform deviations are controlled by the class’s ability to correlate with random signs.
Published:
Any function of independent inputs that no single input can move very far concentrates like a Gaussian.
Published:
The expected max of n sub-Gaussian variables costs only √(2 log n) — no independence required.
Published:
Chaining across all resolution scales at once: process complexity is the integral of root-entropy, localized to a ball.
Published:
Variance-aware concentration for bounded random variables: the Gaussian rate until the range term takes over.
Published:
Annotated study notes on Ma Chapter 5: margin theory, Rademacher complexity of linear models and two-layer nets, and the covering-number route to deep nets.
Published:
Why ‘learning is optimization’: no free lunch, the three-term excess risk decomposition, and the coupling that makes deep learning theory interesting.