#AzureNDGB300
Azure's new ND GB300 VM is crushing it—1.1M tokens/sec on Llama2 70B with FP4, 50% more GPU memory, and TensorRT-LLM tuned for MLPerf v5.1. Curious how this stacks up for your AI workloads? Dive in! #AzureNDGB300 #Llama2_70B #TensorRTLLM

🔗 aidailypost.com/news/microso...
November 4, 2025 at 5:55 AM