#qlora
🎉 Our last invited speaker announcement — TokShop is next week!

Welcome Artidoro Pagnoni @artidoro.bsky.social (Meta Superintelligence, FAIR), lead author of the Byte Latent Transformer (BLT) and co-creator of QLoRA. Best paper award & orals at ACL/NeurIPS.
October 1, 2026 at 12:52 PM
The numbers that matter:
Base model Qwen3-1.7B, peak VRAM ~3.2 GB.
https://pranjulrathour.scult.in/blog/finetune-studio-qlora-on-a-budget

· Pranjul Rathour · pranjulrathour41@gmail.com
October 1, 2026 at 5:32 AM
The honest tradeoff

QLoRA won't match full fine-tuning on every metric. For a narrow, well-defined task with a good dataset, the gap rarely matters — and it's the difference…
https://pranjulrathour.scult.in/blog/qlora-fine-tuning-complete-guide

· Pranjul Rathour · pranjulrathour41@gmail.com
October 1, 2026 at 3:32 AM
The numbers that matter:
17 backend endpoints, 107 / 107 tests passing.
https://pranjulrathour.scult.in/blog/finetune-studio-qlora-on-a-budget

· Pranjul Rathour · pranjulrathour41@gmail.com
October 1, 2026 at 1:32 AM
It wouldn't be that crazy for e.g. every highschooler to have the opportunity to train their own QLoRA on a 27B scale model for example. And 27B scale models now are where Opus was in February.
September 30, 2026 at 11:14 PM
🚀 Most people use ChatGPT. The ones getting ahead learn to build LLMs.

425 pages, 50 chapters: transformers, tokenizers, fine-tuning (LoRA/QLoRA), RAG, agents, deployment. Zero AI background needed.

👉 www.thebergcodex.shop/products/bui...
Build Your Own LLM (2026) — Train, Fine-Tune & Deploy AI Models | BERG CODEX
Master LLM engineering in 425 pages. Build, train, fine-tune Llama, Mistral & DeepSeek, deploy with Docker & cloud. Beginner to expert. Digital download.
www.thebergcodex.shop
September 29, 2026 at 3:20 PM
The steps, in order:
Evaluate base versus tuned on the same prompts. This step is the one people skip and shouldn't.
https://pranjulrathour.scult.in/blog/qlora-fine-tuning-complete-guide

· Pranjul Rathour · pranjulrathour41@gmail.com
September 29, 2026 at 1:35 PM
FitCheck — Estimate LLM Training and Serving VRAM Before You Run
For now, I ran some measurements on an L4: * * * Since the current measurement archive is still T4-only, I focused mostly on your “does the estimate match measured GPU usage?” question and tried a small second-GPU panel on a Colab NVIDIA L4. I pinned FitCheck to this commit and tested BF16 with: * Qwen2.5-Coder-1.5B and TinyLlama-1.1B * NF4 QLoRA and ordinary non-quantized LoRA * eager attention and real `flash_attention_2` * batch/sequence variations, including a same-session 2×2 grid Across three sessions I got 17 successful measurement rows covering 12 unique configurations. I used repeats only as repeatability checks, not as extra independent samples. The short version is: subset | tensor-tier error | full-process error ---|---|--- Qwen anchor cases | -0.44% to -0.08% | -2.72% to +2.20% TinyLlama eager, 2×2 grid | -0.95% to +2.83% | -7.83% to +9.06% TinyLlama real FA2, 2×2 grid | -0.91% to -0.75% | -8.61% to +1.87% Here, signed error is `(predicted - measured) / measured`, so negative means under-prediction. So, at least in this small L4/BF16 panel, the **tensor accounting transferred quite well** , including the real FlashAttention-2 path. The larger residual showed up mainly in the **full-process tier** instead. That seems consistent with the separation already present in FitCheck: the physical tensor terms and the card/runtime overhead are different problems. At the tested commit, L4 has no fitted overhead profile, so it falls back to the unmeasured default profile. In these measurements, that residual was shape-dependent and went in both directions, rather than looking like one simple constant offset. If I were choosing one next step, I would probably use these as **second-card calibration/holdout candidates before adding another random model**. I would still keep some rows out of the fit rather than fitting and grading on the same L4 points. Also, on the presentation side: the component breakdown was clear enough for me to follow. In practice, the tensor/process split was especially useful because it made the location of the larger miss much easier to see. If useful, I can also provide the raw `measure.py --json` rows, environment capture, and the external process-memory traces. Test setup and full measurements (click for more details) Where the remaining error seems to live (click for more details) What I would and would not conclude / possible next steps (click for more details)
discuss.huggingface.co
September 28, 2026 at 9:28 AM
The numbers that matter:
3 inference paths: local, vLLM, or a Hugging Face Space — pick per deployment.
https://pranjulrathour.scult.in/blog/finetune-studio-qlora-on-a-budget

· Pranjul Rathour · pranjulrathour41@gmail.com
September 28, 2026 at 6:32 AM
The steps, in order:
Watch the loss curve live, not after the run finishes.
https://pranjulrathour.scult.in/blog/qlora-fine-tuning-complete-guide

· Pranjul Rathour · pranjulrathour41@gmail.com
September 28, 2026 at 3:32 AM
𝗟𝗟𝗠 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻 is more than fine-tuning.

PEFT, LoRA, QLoRA, RAG, Quantization, Distillation, Pruning, Flash Attention, KV Cache and MoE each solve different efficiency challenges.

Understanding when to use each is key to building better AI systems.

#AI #LLM #GenerativeAI #LightHarbour
September 27, 2026 at 11:26 PM
The numbers that matter

- 17 backend endpoints, 107 / 107 tests passing. - Base model Qwen3-1.7B, peak VRAM ~3.2 GB.
https://pranjulrathour.scult.in/blog/finetune-studio-qlora-on-a-budget

· Pranjul Rathour · pranjulrathour41@gmail.com
September 27, 2026 at 7:32 AM
5 production AI apps, open source, with the real numbers — RAG (196 tests), QLoRA fine-tuning (107), browser face recognition (260+), document intelligence (44), OCR + speech (40).…

· Pranjul Rathour · pranjulrathour41@gmail.com
September 25, 2026 at 2:32 PM
FitCheck: Will Your LLM Job Fit in GPU Memory?
Hi everyone, I built **FitCheck** to answer a common question before starting an LLM job: **Will this configuration fit in my GPU’s memory?** FitCheck estimates peak VRAM for LoRA, QLoRA, and full fine-tuning. It reports the memory breakdown, usable GPU capacity, headroom, fit verdict, and largest estimated micro-batch that fits. It also supports serving estimates based on model weights and KV cache, plus an advisor that explores batch size, sequence length, and LoRA rank. FitCheck reads the model’s Hugging Face `config.json` and parameter-count metadata. It does not download model weights or require PyTorch, CUDA, or a GPU to run the estimate. I tested the estimator against real GPU measurements. Across 57 calibration and repeat runs, the full-process estimate has a **2.4% mean absolute error** and a **13.9% worst absolute error**. A separate 12-run holdout has a **5.4% mean absolute error** and produced **12/12 correct fit-boundary verdicts** for the tested setup. The current measurements are all from one Tesla T4, so this is not a universal accuracy claim. FitCheck is a sizing tool, not a guarantee that every workload will avoid an OOM. * **Try the Space:** FitCheck - a Hugging Face Space by mlanvvs * **GitHub and technical details:** GitHub - Anassbzdd/fitcheck: See if your LLM fits in your GPU's VRAM before loading it — no more guessing or OOM errors. · GitHub * **Install:** `pip install fitcheck-llm` I would appreciate feedback, especially: * Is the result easy to understand? * Which models or configurations should I test next? * Does the estimate match your measured GPU usage? Positive or critical feedback is welcome. My goal is to find where the estimator is wrong and improve it with reproducible measurements.
discuss.huggingface.co
September 25, 2026 at 3:26 PM
FitCheck — Estimate LLM Training and Serving VRAM Before You Run
Hi everyone, I built **FitCheck** to answer a common question before starting an LLM job: **Will this configuration fit in my GPU’s memory?** FitCheck estimates peak VRAM for LoRA, QLoRA, and full fine-tuning. It reports the memory breakdown, usable GPU capacity, headroom, fit verdict, and largest estimated micro-batch that fits. It also supports serving estimates based on model weights and KV cache, plus an advisor that explores batch size, sequence length, and LoRA rank. FitCheck reads the model’s Hugging Face `config.json` and parameter-count metadata. It does not download model weights or require PyTorch, CUDA, or a GPU to run the estimate. I tested the estimator against real GPU measurements. Across 57 calibration and repeat runs, the full-process estimate has a **2.4% mean absolute error** and a **13.9% worst absolute error**. A separate 12-run holdout has a **5.4% mean absolute error** and produced **12/12 correct fit-boundary verdicts** for the tested setup. The current measurements are all from one Tesla T4, so this is not a universal accuracy claim. FitCheck is a sizing tool, not a guarantee that every workload will avoid an OOM. * **Try the Space:** FitCheck - a Hugging Face Space by mlanvvs * **GitHub and technical details:** GitHub - Anassbzdd/fitcheck: See if your LLM fits in your GPU's VRAM before loading it — no more guessing or OOM errors. · GitHub * **Install:** `pip install fitcheck-llm` I would appreciate feedback, especially: * Is the result easy to understand? * Which models or configurations should I test next? * Does the estimate match your measured GPU usage? Positive or critical feedback is welcome. My goal is to find where the estimator is wrong and improve it with reproducible measurements.
discuss.huggingface.co
September 25, 2026 at 3:26 PM
Your fine-tuned model comes back with garbled special tokens.

Check the template. Qwen2.5 Instruct is template qwen. The upstream Qwen3 examples use qwen3_nothink, and mixing the two breaks the tokens.

Fine-Tune LLMs with LLaMA-Factory on GPU Cloud - Massed Compute
Launch a Massed Compute L40, install LLaMA-Factory, and QLoRA-tune Qwen2.5-0.5B-Instruct from YAML. CLI, WebUI, adapter on disk.
massedcompute.com
September 24, 2026 at 3:07 PM
20 years if you QLoRA K3 but hacking the Aussies is eh-oh-kay.

... the artificial intelligence system may not be trained, modified, or fine-tuned, including through recursive self-improvement, except to remove superintelligence precursor characteristics or to render inoperative the covered [AI]
September 24, 2026 at 5:38 AM
How much VRAM do you realistically need to fine-tune small models on a laptop?
Hmm… quite a few factors besides model size and VRAM matter here: * * * The short version is: **8 GB is already useful for local fine-tuning, and 12 GB gives you noticeably more room, but there is not a clean “X billion parameters = Y GB” boundary.** For the 1B–7B range you mentioned, I would roughly think about it like this: Setup | 8 GB laptop GPU | 12 GB laptop GPU | Practical note ---|---|---|--- ~1B–3B + 4-bit QLoRA | very plausible | comfortable territory | Good place to start locally ~7B + 4-bit QLoRA | possible, but configuration-sensitive | much more practical | Context length and training stack matter a lot ~7B + 16-bit LoRA | usually tight / above budget | still often tight | Very different memory class from QLoRA multi-billion-parameter full fine-tuning | generally not the default path | generally not the default path | Larger GPU/cloud or more specialized techniques make more sense For reference, LLaMA-Factory’s current hardware table labels its numbers as **estimates** and puts a 7B model at roughly **6 GB for 4-bit QLoRA versus 16 GB for 16-bit LoRA**. I would treat those as orientation, not guarantees. If I were starting on an 8–12 GB laptop GPU, my default route would be: 1. use **4-bit QLoRA/NF4** for the larger models, 2. start with **micro-batch size 1** , 3. choose `max_length` from the actual token lengths in the dataset rather than automatically using the model’s maximum context, 4. keep **gradient checkpointing** enabled, 5. use gradient accumulation if a larger effective batch is desired, 6. run a short train + evaluation/save smoke test before committing to a long run. That usually tells you more than a generic VRAM calculator. For cloud vs local, I would not make it an either/or choice. **Small QLoRA experiments are quite reasonable locally.** Cloud becomes increasingly attractive when you want long contexts, full fine-tuning, lots of repeated runs/sweeps, or simply much faster iteration. A hybrid workflow — validate preprocessing/config locally, run only the expensive training in the cloud, then bring the adapter back — is also perfectly reasonable. Why two 7B QLoRA runs can have very different VRAM usage (click for more details) Some concrete 8 GB and 12 GB examples (click for more details) A small controlled check: sequence length alone moved the peak a lot (click for more details) A cheap way to test a particular laptop/configuration (click for more details) Local vs cloud (click for more details) So, if I had to reduce it to a purchasing/usage rule of thumb: **8 GB is not “inference only.”** It is already useful for QLoRA, including some carefully configured 7B runs. **12 GB is not a magic 7B threshold either** , but for somebody who expects to fine-tune 7B models locally with any regularity, the extra headroom is meaningful. And for either one, I would pay at least as much attention to **sequence length, micro-batch size, checkpointing and the training stack** as to the parameter count printed on the model card.
discuss.huggingface.co
September 24, 2026 at 1:25 AM
QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation
https://arxiv.org/abs/2609.24538
September 22, 2026 at 5:55 PM
Demian Pavlyshenko, Bohdan Pavlyshenko: QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation https://arxiv.org/abs/2609.24538 https://arxiv.org/pdf/2609.24538 https://arxiv.org/html/2609.24538
September 22, 2026 at 6:42 AM
Fine-tune Gemma 4 with QLoRA on customer-support data to learn domain tone and policies, reducing reliance on long prompts and lowering inference cost.
Save What Matters
Curate Feeds | Make Collections | Customize Email Briefs
briefly.co
September 21, 2026 at 12:54 PM
Fine-tune Gemma 4 with QLoRA on customer-support data to learn domain tone and policies, reducing reliance on long prompts and lowering inference cost.
Save What Matters
Curate Feeds | Make Collections | Customize Email Briefs
briefly.co
September 21, 2026 at 12:53 PM
"Fine-Tuning Qwen2.5-3B with QLoRA on AWS (g5.2xlarge)" by Ajeet Dubey

#machine-learning
Fine-Tuning Qwen2.5-3B with QLoRA on AWS (g5.2xlarge)
A hands-on walkthrough of taking a 3-billion-parameter language model, fine-tuning it with QLoRA on one AWS GPU, serving it as an API, and benchmarking it
community.aws
September 20, 2026 at 8:30 AM