#sft
Jeeves - 추론으로 Jev 계열 의사결정 모델 개선하기

Jeeves 는 Qwen3.5-9B에 LoRA와 포인터 헤드를 결합하고, SFT와 CISPO로 판단 전에 추론하도록 학습한 의사결정 분류 모델임 추론 사용 시 테스트 종합 정확도 0.889 , JevBench 공개 문항 정확도 0.935를 기록함. 다만 JevBench 이외의 Kev/Jev 비교는 같은 데이터 ...
Jeeves - 추론으로 Jev 계열 의사결정 모델 개선하기
Jeeves 는 Qwen3.5-9B에 LoRA와 포인터 헤드를 결합하고, SFT와 CISPO로 판단 전에 추론하도록 학습한 의사결정 분류 모델임 추론 사용 시 테스트 종합 정확도 0.889 , JevBench 공개 문항 정확도 0.935를 기록함. 다만 JevBench 이외의 Kev/Jev 비교는 같은 데이터 ...
news.hada.io
September 29, 2026 at 5:00 PM
Jeeves is a 9B Qwen3.5-based reasoning model designed for decision-making tasks. Trained with SFT and CISPO, it outperforms existing Jev-style models on JevBench. It features a diffusion drafter for accelerated inference and a Jev-compatible API for deployment.
Jeeves. Reasoning improves Jev-like decision models (149)
Jeeves – Reasoning improves Jev-like decision models - PostHog/jeeves
news.ycombinator.com
September 29, 2026 at 3:48 PM
the barrier to entry is just enormous, and as a result it's really only the big labs with fulltime staff who are actually able to deliver. i have yet to find a single amateur SFT that actually fulfills on its promises.
September 29, 2026 at 2:55 PM
The SFT-of-a-regular-model thing is the tell: it reproduces the output shape, not the property that made the shape worth having. A distilled clone can match the picks and still be unable to say how close the call was - and closeness is the number anyone actually routes on.
September 29, 2026 at 2:46 PM
Dylan Zhang, Mingyuan Wu, Jinning Li: Selecting Diverse SFT Traces Improves Post-RL Generalization https://arxiv.org/abs/2609.33780 https://arxiv.org/pdf/2609.33780 https://arxiv.org/html/2609.33780
September 29, 2026 at 6:48 AM
Sencar, Beka, Naeem, Ozalkan, Hawasly, Lucas, AlFuqaha, Abdallah, Senturk: From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data https://arxiv.org/abs/2609.35201 https://arxiv.org/pdf/2609.35201 https://arxiv.org/html/2609.35201
September 29, 2026 at 6:43 AM
Emre Can Acikgoz, Yang Li, Zeyu Leo Liu, Srijan Bansal, Dilek Hakkani-T\"ur, Shafiq Joty, Semih Yavuz: Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training https://arxiv.org/abs/2609.31900 https://arxiv.org/pdf/2609.31900 https://arxiv.org/html/2609.31900
September 29, 2026 at 6:42 AM
Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye: Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features https://arxiv.org/abs/2609.33463 https://arxiv.org/pdf/2609.33463 https://arxiv.org/html/2609.33463
September 29, 2026 at 6:41 AM
Jingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li, Xiaoming Zhai, Wei Chu, Ninghao Liu: Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging https://arxiv.org/abs/2609.32303 https://arxiv.org/pdf/2609.32303 https://arxiv.org/html/2609.32303
September 29, 2026 at 6:39 AM
A new arXiv paper details Rufus-Air, an open post-training recipe for the 106B-A12B GLM-4.5-Air-Base model, walking through an eight-stage pipeline covering SFT, multiple RL stages, agents, and RLHF, plus data, reward,…

#OpenSource #LLM #PostTraining #AIResearch
https://arxiv.org/abs/2609.29421
September 28, 2026 at 8:01 PM
Omg! The QFT class for once went over things I’m familiar with from SFT! Wicks theorem and propagators that are Greens functions etc. it’s weird though because Feynman propagators are not causal, Ive only seen causal propagators.
September 28, 2026 at 6:15 PM
 Politécnico Escuela Hogar del Niño representa a República Dominicana como Campeón Nacional en Solve for Tomorrow 2026 de Samsung

Los estudiantes clasificados avanzan a la etapa de prototipado y mentoría con expertos de Samsung. De los campeones nacionales, se seleccionarán los tres mejores…
 Politécnico Escuela Hogar del Niño representa a República Dominicana como Campeón Nacional en Solve for Tomorrow 2026 de Samsung
Los estudiantes clasificados avanzan a la etapa de prototipado y mentoría con expertos de Samsung. De los campeones nacionales, se seleccionarán los tres mejores proyectos, que viajarán a Panamá para la gran final regional del 18 de noviembre dentro del marco de Solve for Tomorrow 2026 de Samsung. SD. RD.– Samsung Electronics dio a conocer los equipos ganadores por país de la edición 2026 de Solve for Tomorrow (SFT). Este es el programa educativo de Ciudadanía Corporativa que impulsa a jóvenes a resolver problemáticas reales de sus comunidades. Lo hace con herramientas de ciencia, tecnología, ingeniería y matemáticas (STEM), y en esta edición destaca la importancia de Solve for Tomorrow 2026 de Samsung en la región.
tecnologiageek.com
September 28, 2026 at 6:09 PM
With hallucination rates at 15% for GRPO and 18.5% for SFT, neither model is ready for clinical deployment. Founders must stop using standard LLM benchmarks as proxies for clinical safety and evaluate multi-step clinical reasoning pathways.
September 28, 2026 at 1:02 PM
Evaluating SFT, DPO, GRPO, and ICL across 8,201 health records, researchers found that while GRPO achieved the highest automatic scores, blinded expert review showed the SFT baseline had directionally higher ratings on reasoning and treatment feasibility.
September 28, 2026 at 1:02 PM
Standard AI alignment techniques do not guarantee better clinical reasoning. A new study in the Journal of Medical Internet Research exposes an "alignment paradox" where optimizing models for automated metrics can decouple from clinical decision quality.

https://doi.org/10.2196/97221
September 28, 2026 at 1:02 PM
ich tausche ein seit in ein seid. schrecklich. sft.
September 28, 2026 at 8:12 AM
Training on Qwen2.5 (0.5B–7B) is two-stage:
1. SFT on our Peak-Explanation dataset, where each sample pairs the true peaks with a written rationale for picking them and rejecting nearby distractors
2. GRPO with a reward mixing detection F1, heart-rate error, format and completeness
September 28, 2026 at 12:22 AM
time to make a nice cup of sft
September 27, 2026 at 7:34 PM