#TensorRTLLM
NVIDIA just dropped the Vera Rubin NVL72 and it’s already crushing the MLPerf Inference leaderboard. Think massive GPU power, DeepSeek‑R1, Qwen3‑VL and TensorRT‑LLM running at warp speed. Dive into the numbers! #VeraRubinNVL72 #MLPerfInference #TensorRTLLM

🔗 aidailypost.com/news/nvidias...
September 16, 2026 at 7:04 PM
I've been trying for several days to set up TensorRT for accelerating inference of the DeepSeek-R1-Distill-Qwen-32B model in Hugging Face space, but I'm facing a series of dependency conflicts
It’s a different implementation, but it seems like TensorRTLLM is easier to use with TensorRT… huggingface.co ### Accelerated inference on NVIDIA GPUs We’re on a journey to advance and democratize artificial intelligence through open source and open science. github.com ### GitHub - huggingface/optimum-nvidia Contribute to huggingface/optimum-nvidia development by creating an account on GitHub. huggingface.co ### TensorRT-LLM backend We’re on a journey to advance and democratize artificial intelligence through open source and open science. Or try newer torch-tensorrt? github.com/pytorch/TensorRT #### docker/README.md `main` # Building a Torch-TensorRT container * Use `Dockerfile` to build a container which provides the exact development environment that our main branch is usually tested against. * The `Dockerfile` currently uses <a href="https://github.com/bazelbuild/bazelisk">Bazelisk</a> to select the Bazel version, and uses the exact library versions of Torch and CUDA listed in <a href="https://github.com/pytorch/TensorRT#dependencies">dependencies</a>. * The desired versions of TensorRT must be specified as build-args, with major and minor versions as in: `--build-arg TENSORRT_VERSION=a.b` * [**Optional**] The desired base image be changed by explicitly setting a base image, as in `--build-arg BASE_IMG=nvidia/cuda:11.8.0-devel-ubuntu22.04`, though this is optional. * [**Optional**] Additionally, the desired Python version can be changed by explicitly setting a version, as in `--build-arg PYTHON_VERSION=3.11`, though this is optional as well. * This `Dockerfile` installs `cxx11-abi` versions of Pytorch and builds Torch-TRT using `cxx11-abi` libtorch as well. As of torch 2.7, torch requires `cxx11-abi` for all CUDA 11.8, 12.4, 12.6, and later versions. Note: By default the container uses the `cxx11-abi` version of Torch + Torch-TRT. If you are using a workflow that requires a build of PyTorch on the PRE CXX11 ABI, please add the Docker build argument: `--build-arg USE_PRE_CXX11_ABI=1` ### Dependencies * Install nvidia-docker by following https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html#docker ### Instructions - The example below uses TensorRT 10.9.0.34 This file has been truncated. show original
discuss.huggingface.co
April 15, 2025 at 12:56 PM
Just dropped: Step 3.7 Flash now runs on NVIDIA GPUs with SGLang, TensorRT‑LLM and vLLM. Faster inference, smoother dev flow—check out the details! #SGLang #TensorRTLLM #vLLM

🔗 aidailypost.com/news/step-37...
May 29, 2026 at 1:37 AM
LightSeek just dropped TokenSpeed, slashing LLM latency by 50% compared to TensorRT‑LLM. Curious how they pulled it off? Dive into the benchmarks and see the speed boost in action. #TokenSpeed #LLMPerformance #TensorRTLLM

🔗 aidailypost.com/news/lightse...
May 7, 2026 at 10:30 PM
Azure's new ND GB300 VM is crushing it—1.1M tokens/sec on Llama2 70B with FP4, 50% more GPU memory, and TensorRT-LLM tuned for MLPerf v5.1. Curious how this stacks up for your AI workloads? Dive in! #AzureNDGB300 #Llama2_70B #TensorRTLLM

🔗 aidailypost.com/news/microso...
November 4, 2025 at 5:55 AM