#TensorRT
this company needs to be outcompeted in every single area. i want their stock at $1 / share
October 19, 2025 at 8:25 AM
Holy crap, TensorRT is a pain in the ass but it's way faster than base cuda good god
November 11, 2025 at 4:24 AM
🏖️🐻 Les Logiciels Libres de l'été, jour 38 :

@jandotai.bsky.social : une alternative Open Source à ChatGPT qui fonctionne 100% hors ligne sur votre ordinateur. Il supporte de multiples moteurs comme llama.cpp et TensorRT-LLM.
July 28, 2025 at 7:30 PM
📦 NVIDIA / Model-Optimizer
⭐ 3,886 (+22)
🗒 Python

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorR...
GitHub - NVIDIA/Model-Optimizer: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs...
github.com
September 25, 2026 at 12:17 AM
NVIDIA/Model-Optimizer: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc.
September 27, 2026 at 1:15 AM
LightSeek's TokenSpeed, a speed-of-light LLM inference engine

- TensorRT LLM level performance
- vLLM level usability
- Built by a lean and mission-driven team in two months
- MIT license, open-source

Blog: lightseek.org/blog/lightse...
Repo: github.com/lightseekorg...
May 7, 2026 at 4:32 AM
Boost Llama 3.3 70B Inference Throughput 3x with NVIDIA TensorRT-LLM Speculative Decoding

Using in-flight batching, KV caching, custom FP8 quantization, speculative decoding, and more for fast, cost-efficient LLM serving.

developer.nvidia.com/blog/boost-l...
December 18, 2024 at 3:34 AM
📦 janhq / jan
⭐ 24,593 (+91)
🗒 TypeScript

Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM)
GitHub - janhq/jan: Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM)
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM) - janhq/jan
github.com
December 31, 2024 at 4:01 PM
Paperclip just crossed 87k stars and gained 2,589 in a single day—the agent management layer everyone's actually using.
Which one are you forking first—agent management, memory, or model optimization?
September 26, 2026 at 9:55 PM
今日のGitHubトレンド

NVIDIA/Model-Optimizer
NVIDIA Model Optimizerは、量子化やプルーニング、蒸留などの手法を用いてモデルを高速化するためのライブラリです。
Hugging Face、PyTorch、ONNX形式のモデルをサポートしており、Python APIを通じて最適化された量子化チェックポイントを作成できます。
TensorRT-LLMなどの推論フレームワークへの展開を容易にし、効率的なデプロイを実現することを目的としています。
GitHub - NVIDIA/Model-Optimizer: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed. · GitHub
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstre
github.com
September 26, 2026 at 2:57 PM
NVIDIA links TensorRT and Dynamo-Triton for faster AI on multiple GPUs

Read more:
https://quantumzeitgeist.com/nvidia-links-tensorrt-dynamo-triton-faster/
NVIDIA Links TensorRT And Dynamo-Triton For Faster AI On Multiple GPUs
NVIDIA links TensorRT and Dynamo-Triton to accelerate AI performance across multiple GPUs.
quantumzeitgeist.com
September 22, 2026 at 10:30 AM
AI Computing Development Intern, TensorRT-LLM - 2027 - China, Shanghai Job educativ.net/jobs/job/65709...
September 23, 2026 at 7:01 PM
📦 janhq / jan
⭐ 18,929 (+98)
🗒 TypeScript

Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM)
GitHub - janhq/jan: Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM)
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM) - janhq/jan
github.com
May 28, 2024 at 2:51 PM
Running LLMs on encrypted GPUs with near-zero performance cost? NVIDIA Confidential Computing on Blackwell hits 96-98% of baseline throughput with smart TensorRT LLM mitiga… https://developer.nvidia.com/blog/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing
September 23, 2026 at 6:07 AM
Optimizing deep learning is essential to maximize your applications and streamline your workflow. With two powerful tools, torch.compile and TensorRT, we break down which is the best choice for your project: col.la/mltools

#ML #PyTorch
Faster inference: torch.compile vs TensorRT
Pick the right tool for deep learning optimization: torch.compile vs TensorRT
col.la
December 19, 2024 at 4:20 PM
Cortex is a nice way to run LLMs locally. It's like Ollama, but it:
- Is fully C++ and packageable into other applications
- Can swap out engines (llama.cpp, TensorRT, etc.)
- Has better search and easier model quantization discovery
cortex.so
Homepage - Cortex
Cortex is an Local AI engine for developers to run and customize Local LLMs. It is packaged with a Docker-inspired command-line interface and a Typescript client library. It can be used as a standalon...
cortex.so
October 31, 2024 at 6:13 PM
NVIDIA Dynamo-Triton 26.07公開、TensorRTのマルチGPU統合機能に対応

https://localmodelwatch.tsuchitsuchi.com/2026/09/22/nvidia-dynamo-triton-2607-tensorrt-multi-gpu/
September 21, 2026 at 10:09 PM
Recently, vLLM ran a benchmark for DeepSeek's R1 on Nvidia's DGX with 8x H200 GPUs. The surprising part was the performance of Nvidia's TensorRT-LLM — see below.
April 19, 2025 at 4:42 PM
Depth Anything V3 now runs in real-time with our karl.'s camera data predicting metric depth from monocular images. With TensorRT optimization, we’ve wrapped it into a ROS2 inference node that’s ready to drop into your stack.

Github: github.com/ika-rwth-aac...

#Robotics #ROS2 #TensorRT
December 15, 2025 at 2:47 PM
(2/2) more from today's trending ↓
September 24, 2026 at 10:10 PM
🤖 NVIDIA enables native multi-device AI inference for generative models

NVIDIA has introduced native multi device inference support in TensorRT 11.0, allowing generative AI pipelines to scale across multiple GPUs. This development...

#GenerativeAI #AIInference #DeepLearning #AI #AIPulse
Read the full article →
www.synestesia.uk
June 26, 2026 at 1:39 PM
Original post on medium.com
medium.com
February 18, 2025 at 3:27 AM
If you're familiar with Python 3, then Tripy delivers the benefits of TensorRT — including performance-boosting compilation.
NVIDIA's Tripy Provides a "Debuggable Pythonic Frontend" to TensorRT for Deep Learning Projects
If you're familiar with Python 3, then Tripy delivers the benefits of TensorRT — including performance-boosting compilation.
www.hackster.io
February 11, 2025 at 5:47 PM