H²V — Scalable Neural Network Geometric Robustness Validation via Hölder Optimisation
In brief
H²V is a formal method for the deep validation of robustness properties of large machine-learning models over continuous regions of the input space, rather than at isolated test points. Based on optimisation, the method scales this analysis to models with up to 300 million tunable parameters, including large vision transformers and 3D video models. This opened the way to formally assessing robustness properties at the scale of contemporary machine-learning models used in physical AI systems.
Problem and contribution
Traditional formal neural-network verification methods, such as symbolic interval propagation, can establish that a model behaves correctly for every input in a specified region. However, their computational cost and the looseness of the resulting bounds limit their applicability to production-scale models and to establishing practically useful safety envelopes. H²V tackles the problem in a novel way through global optimisation. A Hilbert space-filling construction reduces a multidimensional robustness problem to a one-dimensional one, after which Hölder optimisation progressively constructs and refines lower bounds on the model's robustness margin. This avoids the layer-by-layer symbolic propagation used by conventional verifiers, thereby enabling analysis at substantially larger model scales. In this paper, the regions considered are generated by geometric perturbations such as rotation, translation and scaling, but the approach applies more generally to other low-dimensional transformations.

Why it matters
The key advantage of H²V lies in its scalability. H²V makes rigorous robustness analysis possible on networks up to two orders of magnitude larger than those typically accessible to standard verification techniques such as symbolic interval propagation and branch-and-bound. The paper demonstrates scalability up to 300M parameters and, for the first time, to large 3D video classifiers. H²V is formally sound when its convergence conditions are established; outside that soundness envelope, carefully chosen optimisation settings provide a high-confidence validation regime in which incorrect robustness conclusions are made very unlikely. Across the extensive soundness benchmarks reported in the paper, no incorrect H²V result was observed, even though errors have been reported in implementations of methods that are theoretically sound.
Technical direction and further work
The approach is particularly effective when the relevant perturbation space is low-dimensional even though the model and its raw input are very large. This is a natural setting for studying model robustness against environment variability. Our subsequent work extends this approach to vision-language and vision-language-action models, scaling region-level validation to models with tens of billions of parameters. Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models This opens a route towards providing safety envelopes for forthcoming physical-AI systems, where multimodal models increasingly sit in the loop between perception and action.
Citation and BibTeX
@inproceedings{h2v-2025,
title = {Scalable Neural Network Geometric Robustness Validation via Hölder Optimisation},
author = {Zhang, Yanghao and Kouvaros, Panagiotis and Lomuscio, Alessio},
booktitle = {Advances in Neural Information Processing Systems 38},
year = {2025},
url = {https://papers.nips.cc/paper_files/paper/2025/hash/1435ac9924d7621e44aa2407a1cfdec7-Abstract-Conference.html},
doi = {10.52202/085713-0458},
}