AI Inference
Tag1.8K stories
AI Inference news and updates covering the stage where a trained model serves predictions, as distinct from training. Readers can learn about serving frameworks, batching and KV caching, quantization, latency and throughput tuning, accelerator selection, and the cost of running models in production.
Accelerate GenAI App Development with New Updates to Databricks Model ServingProduction-Quality RAG Applications with DatabricksServing and Deploying Machine Learning Models with BentoML: Germany Car Price Prediction Case StudyServing LLMs on an RTX4090 with SequoiaA Hitchhiker’s Guide to Speculative DecodingHuawei AI Introduces ‘Kangaroo’: A Novel Self-Speculative Decoding Framework Tailored for Accelerating the Inference of Large Language ModelsLayerSkip: An End-to-End AI Solution to Speed-Up Inference of Large Language Models (LLMs)Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpointslm-sys/FastChat: An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.Turbocharging Meta Llama 3 Performance with NVIDIA TensorRT-LLM and NVIDIA Triton Inference Server