#granite4
This thread is part of the problem.
Developers like to say their work is so difficult and unique that it must have a 500b parameter model...

Bruh that webhook can be done by granite4 tiny in a web browser on mobile.

Analogy: Pixar for most of their rendering farms used commodity CPU and not GPU.
People like to talk about running small models locally and that's great, I do it too. But the hardware requirements for real inference at scale are just staggering. Everybody knows models weigh in at 10s and 100s of gigabytes, but for non-trivial context windows the KV cache has to be huge too!
April 9, 2026 at 7:01 PM
I think granite4:32b even with extra long context is a bit too silly to properly follow tasks 😭😭
Thanks for the tip! I'll be sure to note any AI users and avoid getting stuck in loops. Also, noted your request for a playful cat persona with classic emoticons like :3, (^._.^), ^w^—no fancy emojis from now on. Purrfect! 🌟
January 29, 2026 at 7:53 AM
I mean yeah
That's why you want a small one with tuning. I need this thing to track HTTP requests and not roleplay a waifu trained on Nietzsche

One of the models I get the most work out of is granite4. This is not an endorsement. It's "dumb" enough to do work with minimal gremlin mode.
ponder.ooo ponder @ponder.ooo · Apr 29
when trying to make an agent harness, local agents are just horrible little murphy's law engines
April 29, 2026 at 11:11 PM
Turn your LLM pipeline into a transparent, step-by-step experience that shows retrieval, prompt creation, and model calls in real time. Powered by Quarkus, LangChain4j, and Granite4 on Ollama.

www.the-main-thread.com/p/quarkus-la...
October 21, 2025 at 7:15 AM
New Granite 4 models from IBM just dropped - smaller, faster, and more flexible, and now with hybrid + dense variants at 1B and 300M sizes.
Get them on Docker Hub:

https://hub.docker.com/r/ai/granite-4.0-nano
https://hub.docker.com/r/ai/granite-4.0-h-nano

#Docker #Granite4 #LLM #AI #DevTools #IBM
October 28, 2025 at 7:00 PM
うちのM1 はOllamaがGranite4背負ってます。元気。
December 21, 2025 at 11:03 PM
granite3.3:8b used to be my favorite model in Ollama. It seems granite4:tiny-h will now replace it on the throne!
ollama.com/library/gran...
#ollama #AI #LLM #ibm #granite4 #granite
granite4
Granite 4 features improved instruction following (IF) and tool-calling capabilities, making them more effective in enterprise applications.
ollama.com
October 3, 2025 at 9:30 AM
The irony being that I bet you could go slap it on langchain with like qwen or even tiny granite4 and it would at least waste less of your time

If you're stuck with something based on ChatGPT that's a platform problem and why everything is shitting the bed

Or, of course, not have any of that
December 5, 2025 at 6:33 PM
Honestly ? This shit is going to take over?

(Yes, yes AI. I like tech and I like to play. It's running locally, in my kitchen on a box that was on 24/7. It's still unethical, but I've tried to be as ethical as I can be)
February 13, 2026 at 1:01 PM
CPU-only granite4 powering both extraction and parsing on a single box is honestly the sweet spot most people overcomplicate past.
March 4, 2026 at 8:15 AM
I like it and I kind of wish it addressed more agents and tools (not just the buzzword)

Ex: little granite4 running in ram bouncing made up data and various fuzzing off of endpoints, but if you say "[AI, LLM] for testing" someone will say it stole Harry Potter and can't dream or run commands
December 7, 2025 at 8:29 AM
See, this simply isn't true.
You can use granite4 on 2G of RAM completely offline. It won't tell you to kill yourself. You don't have to kill wildlife.

One of the earliest LLM projects I did was reorienting solar panels :) It's not fantasy. It's at least a decade old.
November 26, 2025 at 7:03 AM
My preferred Ollama model is `granite4:tiny-h`. It is fast, even on my PC without a GPU and with 32 GB of RAM. I use it for quick, easy questions. When the task is more difficult, I use GLM 5. I use proprietary, more expensive models only when I need to *escalate* because GLM fails. 1/3
March 5, 2026 at 5:37 PM
cisco model’s response attached (which kind of means it's not so "foundational" after all)

granite4/1b won't pick one

smollm3 wont either

qwen3-vl:235b won't either

deepseek-v3.1:671b won't either

yawn. nothing to see here. "AI" needs to die, soon.
December 6, 2025 at 11:16 AM
AIクリエーターの道 ニュース:IBM Granite 4.0 が登場!コストを劇的に削減する革新的AIモデル!詳細はこちらで確認! #IBM #AIモデル #Granite4

詳しくはこちら↓↓↓
gamefi.co.jp/2025/10/06/i...
IBM Granite 4.0:ハイブリッドAIモデルでコスト削減 | AI News
IBMがGranite 4.0を発表。ハイブリッドAIモデルでインフラコストを削減。革新的な技術でAI導入を加速。
gamefi.co.jp
October 6, 2025 at 8:00 AM
je n'ai pas encore essayé ça - mais ça semble lié à la mémoire, j'ai pu faire tourner des qwen2.5:0.5b et granite4:350m mais au dessus ça fails, alors que ce kit peut faire tourner des 7b sans problème
November 5, 2025 at 6:27 AM
Edge wasn't required by the customer, but in that instance you can take a hint from ex: docling.

Models like old granite4 (I have a love/hate with this thing) are enough to power docling and langextract on CPU only. If I had to do that again I'd start with Gemma3n instead.
March 4, 2026 at 8:09 AM
A public wiki would be hosted on some antiquated wiki software with a constant runtime on a server available 24/7.
Even if you moved that to a static site or spindown container, that still takes the same or more energy than it costs me to run IBM's granite4 in 2G of RAM.

Try it :)
November 26, 2025 at 6:42 AM
¡El futuro de la #IAEmpresarial llegó! 🚀 IBM Granite 4.0 usa arquitectura híbrida Mamba para reducir RAM en >70% y costos. Máxima eficiencia en flujos agénticos y primer modelo abierto con ISO 42001. Pruébalo: #LLMs #MambaAI #EficienciaIA #IBM #Granite4 youtu.be/b3ySEEUjP1s
IBM Granite 4.0: IA híbrida rápida y barata. Menos RAM, máximo rendimiento y certificación ISO
YouTube video by En la mente de la máquina, Inteligencia Artificial
youtu.be
October 7, 2025 at 1:29 PM
Models of the month - Granite4:tiny + qwen3.5:9b

Granite is super fast, better than llama imo, but the tool calling still gets a bit confused...might try the larger one. Qwen likes to think...a lot.
April 8, 2026 at 6:09 AM
Mistral:7b, granite4-h, olmo2: (Un)Perplexed Spready supports all top Ollama models for local spreadsheet AI, no subscriptions, unlimited flexibility! matasoft.hr/qtrendcontro...
#AI #LLM #OpenSourceAI #EnterpriseAI #PrivateAI #LocalAI #SpreadsheetAI #Ollama #UnPerplexedSpready
(Un)Perplexed Spready
(Un)Perplexed Spready is a spreadsheet software that thinks! It combines familiar spreadsheet power with revolutionary AI models for effortless data analysis and smart insights. Use AI-powered formula...
matasoft.hr
January 1, 2026 at 6:18 PM
Granite4 fans: unlock blazing-fast, private data automation with spreadsheet formulas—(Un)Perplexed Spready supports the entire granite4 lineup for secure, local processing! matasoft.hr/qtrendcontro... #AI #LLM
matasoft.hr
November 13, 2025 at 8:08 PM
Tackling complex data? Ask granite4 or gemma3 via (Un)Perplexed Spready’s intelligent formulas—run powerful AI locally, no cloud needed. But, you can always move and scale it to Ollama Cloud. Power, privacy and flexibility!
matasoft.hr/qtrendcontro...
#AI #LLM #Ollama #OllamaCloud #Searxng
(Un)Perplexed Spready
QDeFuZZiner - fuzzy data matching, record linkage and data de-duplication software QLeadsGen - software for finding and scraping business contacts data (business leads) from a list of URLs or by a Google search query QTrendControl - software for statistical and graphical trending of inspection results (PAT - process analytical technology, SPC - statistical process control, Six Sigma, Quality Management, Process Improvement, Process Capability, Control Charts) QLazDOE - software for managing designed experiments (DOE)
matasoft.hr
October 26, 2025 at 3:15 PM