NVIDIA TensorRT Model Optimizer is a comprehensive library of state-of-the-art post-training and training-in-the-loop model optimization techniques. It offers quantization techniques, enables ultra-low precision inference, and provides model compression with sparsity.
Table of contents
Quantization techniquesEnabling ultra-low precision inference for next-generation platformsModel compression with sparsityComposable model optimization APIsGet startedShare this post