#NeuralMagic
Have you tried this one with vLLM (which supports structured gen with various options for structured gen)? huggingface.co/neuralmagic/...
neuralmagic/SmolLM-1.7B-Instruct-quantized.w8a16 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
November 11, 2024 at 11:40 AM
You should be able to run vLLM locally. Neural Magic also has several other quantized models that you can try. huggingface.co/collections/...
INT4 LLMs for vLLM - a neuralmagic Collection
Accurate INT4 quantized models by Neural Magic, ready for use with vLLM!
huggingface.co
November 11, 2024 at 11:46 AM
Today, RedHat completed the acquisition of @NeuralMagic, a pioneer in software and algorithms that accelerate #GenAI inference workloads. Read how we are accelerating our vision for #AI’s future: red.ht/408kJ8K.
Red Hat Completes Acquisition of Neural Magic to Fuel Optimized Generative AI Innovation Across the Hybrid Cloud
Red Hat has completed its acquisition of Neural Magic, a pioneer in software and algorithms that accelerate generative AI (gen AI) inference workloads, furthering its vision of high-performing AI work...
red.ht
January 14, 2025 at 12:13 PM
Today, Red Hat completed the acquisition of @NeuralMagic, a pioneer in software and algorithms that accelerate #GenAI inference workloads. Read how Red Hat is accelerating their vision for #AI’s future: red.ht/408kJ8K
Red Hat Completes Acquisition of Neural Magic to Fuel Optimized Generative AI Innovation Across the Hybrid Cloud
Red Hat has completed its acquisition of Neural Magic, a pioneer in software and algorithms that accelerate generative AI (gen AI) inference workloads, furthering its vision of high-performing AI workloads...
red.ht
January 13, 2025 at 3:39 PM
DeepSparse is a CPU inference runtime that takes advantage of sparsity to accelerate neural network inference. Coupled with SparseML, our optimization library for pruning and quantizing your models, DeepSparse delivers exceptional inference performance on CPU hardware.
GitHub - neuralmagic/deepsparse: Sparsity-aware deep learning inference runtime for CPUs
Sparsity-aware deep learning inference runtime for CPUs - GitHub - neuralmagic/deepsparse: Sparsity-aware deep learning inference runtime for CPUs
github.com
November 27, 2023 at 4:12 AM
January 17, 2025 at 8:24 PM
📦 neuralmagic / deepsparse
⭐ 1,899 (+64)
🗒 Python

Sparsity-aware deep learning inference runtime for CPUs
GitHub - neuralmagic/deepsparse: Sparsity-aware deep learning inference runtime for CPUs
Sparsity-aware deep learning inference runtime for CPUs - GitHub - neuralmagic/deepsparse: Sparsity-aware deep learning inference runtime for CPUs
github.com
October 21, 2023 at 10:50 PM
Red Hat adquire startup de otimização de IA Neural Magic - Tecnocrata
…
tecnocrata.com.br
November 12, 2024 at 1:45 PM
Congratulations to @neuralmagic and @ZapataComputing recognized by @BostInno as "Startups to Watch in 2020" - we definitely agree. https://www.americaninno.com/boston/20-Startups-to-Watch-in-2020-Bos
November 17, 2024 at 11:21 AM