TensorRT
Tag102 stories
TensorRT news and updates for the NVIDIA SDK that compiles and optimizes trained neural networks for inference on NVIDIA GPUs. Readers can learn about TensorRT-LLM, engine building, quantization, Triton deployment, RTX desktop support and reported latency results.
Accelerate Generative AI Inference Performance with NVIDIA TensorRT Model Optimizer, Now Publicly AvailableHow TensorRT Accelerates AI on RTX PCsCity of Raleigh Taps NVIDIA Metropolis to Improve TrafficNVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training QuantizationDetecting Real-Time Waste Contamination Using Edge Computing and Video AnalyticsDeploying LLMs Into Production Using TensorRT LLMEmulating the Attention Mechanism in Transformer Models with a Fully Convolutional Networkcollabora/WhisperFusion: WhisperFusion builds upon the capabilities of WhisperLive and WhisperSpeech to provide a seamless conversations with an AI.How Amazon and NVIDIA Help Sellers Create Better Product Listings With AIColossal-AI Team Open-Sources SwiftInfer: A TensorRT-Based Implementation of the StreamingLLM Algorithm