#LLMQuantization
September 27, 2026 at 6:32 AM
A developer guide to running local LLMs on 8GB GPUs using llama.cpp, quantization, and GPU offloading for efficient AI performance. #llmquantization
Optimizing Local LLM Inference for 8GB VRAM GPUs
hackernoon.com
March 21, 2026 at 11:34 AM
Comparing parameter selection strategies and discussing the future of AI model compression. #llmquantization
The Future of AI Compression: Smarter Quantization Strategies
hackernoon.com
March 6, 2025 at 7:30 PM
Testing quantization methods to see if AI models can maintain performance with fewer bits. #llmquantization
Quantizing Large Language Models: Can We Maintain Accuracy?
hackernoon.com
March 6, 2025 at 7:30 PM
A deep dive into the latest research on LLM quantization, exploring past methods and new breakthroughs. #llmquantization
Rethinking AI Quantization: The Missing Piece in Model Efficiency
hackernoon.com
March 6, 2025 at 7:30 PM
Discover how a small subset of LLM parameters, called "cherry" parameters, hold a disproportionate influence on model performance. #llmquantization
The Hidden Power of "Cherry" Parameters in Large Language Models
hackernoon.com
March 6, 2025 at 7:30 PM
Discussion on Qwen3-235B's performance on Cerebras raises questions about quantization (FP16, FP8). Does Cerebras use dynamic precision, and how does it impact model quality vs. speed? Crucial for practical deployment. #LLMQuantization 4/6
July 24, 2025 at 1:00 PM
Why do some parameters influence model performance more than others? This blog quantifies their impact. #llmquantization
The Impact of Parameters on LLM Performance
hackernoon.com
March 6, 2025 at 7:30 PM
Investigating the effect of quantization on chat-based LLMs like Vicuna-1.5. #llmquantization
Can ChatGPT-Style Models Survive Quantization?
hackernoon.com
March 6, 2025 at 7:30 PM
Examining how base model quantization impacts perplexity and downstream performance. #llmquantization
The Perplexity Puzzle: How Low-Bit Quantization Affects AI Accuracy
hackernoon.com
March 6, 2025 at 7:30 PM
Unpacking parameter heterogeneity: Why do only 1% of LLM parameters significantly impact performance? #llmquantization
The Science of "Cherry" Parameters: Why Some LLM Weights Matter More
hackernoon.com
March 6, 2025 at 7:30 PM