#ModelMerging
How Sakana AI’s new evolutionary algorithm builds powerful AI models without expensive retraining https://venturebeat.com/ai/how-sakana-ais-new-evolutionary-algorithm-builds-powerful-ai-models-without-expensive-retraining/ #AI #ModelMerging
August 30, 2025 at 5:57 AM
🎉 Excited to share that our paper "OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training" has been accepted to EMNLP 2026 Main! See you in Budapest!

📄 Paper: arxiv.org/abs/2603.28858

#EMNLP2026 #ModelMerging #ContinualPretraining #BayesianOptimization
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
Continual pre-training is widely used to adapt LLMs to target languages and domains, yet the mixture ratio of training data remains a sensitive hyperparameter that is expensive to tune: they must be f...
arxiv.org
September 15, 2026 at 5:24 AM
Excited to be at @neuripsconf.bsky.social next week co-presenting our tutorial: "Model Merging: Theory, Practice, and Applications" 🔥

Proud to do this with my PhD advisor, Colin Raffel, our postdoc research fellow @mcicc.bsky.social, and an incredible panel of speakers 💙

#NeurIPS2025 #ModelMerging
November 27, 2025 at 2:43 PM
🔥 Consolidating Corporate LLM Traffic: A New Recipe for Self-Hosted Efficiency

https://pneumetron.com/news/ai_research/consolidating-corporate-llm-traffic-self-hosted-efficiency-c5dc80

#LLM #GRPO #SLERP #ModelMerging
September 5, 2026 at 6:42 AM
ISO: Unlocking Efficient RLVR Through Spectral Inheritance

https://pneumetron.com/news/ai_research/iso-rlvr-spectral-inheritance-0ff3c1

#RLVR #Optimization #MachineLearning #ModelMerging
July 23, 2026 at 3:06 AM
Forget complex alignment! Direct weighted averaging for LLM merging shows promise, but beware the 'seesaw effect' from unbalanced ratios. Limits of simple fusion may bound advanced techniques. #AI #LLMs #ModelMerging

https://www.startuphub.ai/ai-news/ai-research/2026/simple-llm-merging-surprises
Simple LLM Merging Surprises
Direct weighted averaging of LLMs with dimensional adaptation proves surprisingly effective, but ratio control is paramount to avoid capability collapse.
www.startuphub.ai
July 22, 2026 at 6:05 PM
SyMerge introduces a single-layer adaptation technique for synergistic model merging. Read more: https://getnews.me/symerge-single-layer-adaptation-for-synergistic-model-merging/ #symerge #modelmerging #ai
October 8, 2025 at 3:25 PM
Effective noise scale—combining learning rate, weight decay, batch size and augmentation—predicts model‑merging success, with a non‑monotonic optimum. Read more: https://getnews.me/optimizer-noise-shapes-model-merging-success-in-neural-networks/ #modelmerging #effectivenoisescale #optimizers
October 8, 2025 at 3:49 AM
Chain of Merges (CoM) merges weights layer‑by‑layer, updates activation stats to curb covariate shift, and hits state‑of‑the‑art results on several benchmarks. Read more: https://getnews.me/chain-of-merges-new-layer-wise-method-improves-model-fusion/ #modelmerging #chainofmerges
October 3, 2025 at 9:34 AM
Study shows a one‑epoch fine‑tune yields a vector = –η∇L, linking arithmetic to gradients. Vision benchmarks confirm the first‑epoch gradient dominates, enabling merging. Read more: https://getnews.me/task-vectors-match-gradient-descent-theory-improves-model-merging/ #taskarithmetic #modelmerging
October 3, 2025 at 9:26 AM
Researchers unveiled a compact power-law scaling rule linking base model size and number of experts (k), showing gains drop about 1/k. Paper submitted September 2025. Read more: https://getnews.me/new-scaling-laws-reveal-predictable-gains-from-model-merging-in-llms/ #modelmerging #scalinglaws #llms
September 30, 2025 at 7:14 PM
A recent study finds model merging provides little protection against transfer attacks, with a relative success rate over 95% across eight merging methods. Read more: https://getnews.me/model-merging-increases-vulnerability-to-adversarial-transfer-attacks/ #modelmerging #adversarial
September 30, 2025 at 11:58 AM
Model merging blends two pretrained LLMs via weight arithmetic, creating models; the study found Pareto improvements where the merged model beat a parent on accuracy and token usage. Read more: https://getnews.me/model-merging-enables-tunable-reasoning-performance-in-llms/ #modelmerging #llms
September 29, 2025 at 11:13 AM
Linear weight merging of a retrieval model with a domain‑specific one improves search in medicine and Japanese text, matching LoRA fine‑tuning even with limited data. Read more: https://getnews.me/model-merging-boosts-domain-specific-ad-hoc-retrieval-effectiveness/ #modelmerging #adhocretrieval
September 29, 2025 at 10:33 AM
Researchers merged Qwen2.5 models, then tested them on the MMLU benchmark and probing of morphology and syntax, finding stronger linguistic knowledge despite mid scores. Read more: https://getnews.me/evaluation-pipeline-connects-model-merging-behavior-and-internals/ #modelmerging #mmlu #probing
September 26, 2025 at 2:00 PM
PriME merges language models with evolutionary algorithms, boosting task performance by up to 45% on the LaMP benchmark while cutting membership‑inference risk. Read more: https://getnews.me/privacy-preserving-evolutionary-merging-improves-language-models/ #privacy #modelmerging
September 22, 2025 at 10:08 PM
A model‑merging technique superposes task‑specific features with transformation matrices, avoiding fine‑tuning. Benchmarks in NLP and vision showed accuracy exceeding heuristic averaging. Read more: https://getnews.me/new-method-superposes-task-specific-features-for-model-merging/ #modelmerging #ml
September 20, 2025 at 8:21 AM
March 25, 2024 at 7:25 AM