TTOpt: A Maximum Volume Quantized Tensor Train-based Optimization and its Application to Reinforcement Learning
Konstantin SozykinAndrei ChertkovRoman SchutskiAnh-Huy PhanAndrzej S. CichockiIvan V. Oseledets
Develops a gradient-free optimization algorithm based on quantized tensor train decomposition and the maximum matrix volume principle, enabling direct training of quantized neural network policies for reinforcement learning with substantially fewer function evaluations than existing methods.
Many modern artificial intelligence and control applications require optimizing complex, non-differentiable target functions where traditional gradient-based methods fail. While gradient-free approaches such as evolutionary algorithms provide a flexible alternative, they typically suffer from slow convergence and demand massive numbers of function evaluations. Furthermore, deploying neural network policies directly onto low-power edge hardware requires quantized, discrete parameters, a constraint that traditional continuous optimization methods struggle to handle efficiently.
The article introduces and evaluates the Tensor Train Optimizer (TTOpt), a novel gradient-free algorithm designed to solve multivariable optimization problems across discrete and continuous domains. The primary objective is to demonstrate that tensor network decomposition combined with submatrix volume maximization can locate optimal parameters substantially faster and with fewer function evaluations than existing gradient-free benchmarks, particularly in reinforcement learning settings.
To evaluate this approach, the authors reformulated optimization problems into discrete multidimensional tensors and used low-rank Tensor Train structures alongside a maximum matrix volume selection principle to sample only a tiny fraction of total possibilities. They benchmarked TTOpt against standard zeroth-order methods (such as genetic algorithms and evolutionary strategies) and classical gradient-based algorithms across ten standard mathematical benchmark functions and four continuous-control reinforcement learning environments, testing problems scaling up to 500 dimensions on standard consumer computing hardware.
The findings show that TTOpt consistently achieved superior precision on mathematical benchmarks, maintaining stable accuracy even as problem dimensionality scaled from 10 to 500 variables while running in 2.3 to 2.6 seconds. In reinforcement learning tasks, TTOpt matched or exceeded baseline performance across coarse and fine discrete grids, notably solving the InvertedPendulum-v2 control task with a perfect cumulative reward (1000.00 with zero variance) where evolutionary baselines showed high variance. Additionally, the algorithm exhibited significantly faster execution times and more consistent convergence curves across multiple runs compared to popular evolutionary strategies.
These results demonstrate that direct optimization of heavily quantized, discrete model weights is practical and highly effective for continuous control tasks. This capability enables organizations to train compact neural network policies directly for low-power edge devices, drastically lowering computational runtime, hardware power demands, and cloud training costs without sacrificing control accuracy or policy performance.
Organizations developing edge-deployed artificial intelligence or complex control systems should consider testing tensor-based optimization workflows as an alternative to standard evolutionary search. Further development should focus on testing TTOpt in larger-scale industrial reinforcement learning pipelines and multi-agent systems to validate scaling trade-offs before broad production deployment.
While the empirical results are strong, users should note that the multidimensional sweep formulation currently lacks formal theoretical guarantees for global convergence rates. Readers should treat convergence behavior as monotonic but empirical, paying close attention to initial rank selection and computational query budgets when deploying on uncharacterized objective functions.
- Paper: Tensor Decomposition for Signal Processing and Machine Learning, Nicholas D. Sidiropoulos et al. (2016). Provides a comprehensive foundation in multilinear algebra and low-rank tensor decompositions that underpin TTOpt's tensor-network formulation.
- Paper: Evolution Strategies as a Scalable Alternative to Reinforcement Learning, Tim Salimans et al. (2017). Introduces the use of gradient-free evolutionary strategies for training neural network policies in continuous control, which serves as a primary baseline and motivation for TTOpt.
- Paper: The CMA Evolution Strategy: A Tutorial, Nikolaus Hansen (2016). Details the principles of Covariance Matrix Adaptation Evolution Strategies (CMA-ES), a classic zeroth-order optimization benchmark evaluated against TTOpt.
- Paper: Benchmarking Deep Reinforcement Learning for Continuous Control, Yan Duan et al. (2016). Establishes the standard continuous control reinforcement learning benchmarks and baseline comparisons used to evaluate policy search algorithms like TTOpt.
- Paper: Fast Tensor Completion via Approximate Richardson Iteration, Mehrdad Ghadiri et al. (2025). Develops fast approximate methods for completing missing data in tensor train and related decomposition formats, offering complementary algorithmic acceleration for low-rank tensor operations.
- Paper: OMPQ: Orthogonal Mixed Precision Quantization, Yuexiao Ma et al. (2023). Explores layer-wise mixed-precision neural network quantization, extending the goal of deploying low-precision policies onto resource-constrained edge hardware.
- Paper: Adaptive Data-Free Quantization, Biao Qian et al. (2023). Investigates low-bit neural network quantization under data-free constraints, complementing TTOpt's study of heavily quantized discrete parameter spaces.
