• All tags
  • tensorrt

TensorRT

Tagยท102 stories

TensorRT news and updates for the NVIDIA SDK that compiles and optimizes trained neural networks for inference on NVIDIA GPUs. Readers can learn about TensorRT-LLM, engine building, quantization, Triton deployment, RTX desktop support and reported latency results.

Accelerate Generative AI Inference Performance with NVIDIA TensorRT Model Optimizer, Now Publicly AvailableHow TensorRT Accelerates AI on RTX PCsCity of Raleigh Taps NVIDIA Metropolis to Improve TrafficNVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training QuantizationDetecting Real-Time Waste Contamination Using Edge Computing and Video AnalyticsDeploying LLMs Into Production Using TensorRT LLMEmulating the Attention Mechanism in Transformer Models with a Fully Convolutional Networkcollabora/WhisperFusion: WhisperFusion builds upon the capabilities of WhisperLive and WhisperSpeech to provide a seamless conversations with an AI.How Amazon and NVIDIA Help Sellers Create Better Product Listings With AIColossal-AI Team Open-Sources SwiftInfer: A TensorRT-Based Implementation of the StreamingLLM Algorithm

Recommended TensorRT stories

Who to follow for TensorRT

Top sources covering TensorRT

Most upvoted TensorRT posts

Best discussed TensorRT posts

All posts about TensorRT