#language-learning
Fixing GRPO's credit assignment problem without evaluating every step https://arxiv.org/abs/2609.36178 (https://news.ycombinator.com/item?id=49934012)
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
arxiv.org
October 2, 2026 at 6:41 PM
Everyone gets annoyed if you get too specific with describing the internal nature. But it is closer to the math done on raw semantic forms than it is language. Or the more cautious description is: we are still learning a lot about what happens in models and what those numbers mean.
October 2, 2026 at 6:35 PM
Periodontitis Risk Assessment and Prevention Planning: Comparative Study of Multimodal Large Language Models and Periodontist Evaluations
Periodontitis Risk Assessment and Prevention Planning: Comparative Study of Multimodal Large Language Models and Periodontist Evaluations
Background: Periodontitis is one of the most prevalent yet preventable oral diseases, as indicated by multiple clinical and radiographic factors. As these factors are recorded in electronic health records (EHRs), their reuse offers opportunities for personalized risk assessment and targeted prevention. Predictive AI and traditional machine learning models support fragmented detection tasks but lack the integration of textual and imaging predictors. Emerging multimodal large language models (M-LLMs) show promise in combining these data sources for clinical assessment. Evaluating the capabilities of M-LLMs and comparing them against the current clinical standard are therefore essential to determine their potential as digital assistants. Objective: This study aimed to evaluate the ability of M-LLMs to assess periodontitis risk and suggest prevention strategies, based on EHR data and radiographic findings. Each M-LLM was individually evaluated by periodontal experts, benchmarked against other models, and compared with a periodontist as a reference. Methods: A vignette study was conducted following TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines for the evaluation of LLMs. Ten periodontal vignettes were created, each including a panoramic radiograph and textual EHR data. Three LLMs capable of reasoning and handling multimodal data were compared to a periodontist who generated outputs manually, based on the same prompts and input data. Periodontal experts rated all outputs across 6 predefined criteria on a 5-point Likert scale. Statistical analyses evaluated overall performance per model and tested whether performance varied per model, scenario complexity, or rater. Results: GPT o1 Pro and Claude Sonnet 4 showed strong performance, with 86.7% and 85.6% of ratings deemed acceptable—comparable to the periodontist’s output (87.8%). Gemini 2.5 Pro was rated significantly lower than both the periodontist and the other models (59.4% acceptable;
dlvr.it
October 2, 2026 at 6:30 PM
I assume I am learning the general topic, probably the key figures or dates are correct. I also assume I'm missing important context and have no way to know what that is without consulting a native speaker. I do not cite or report on foreign language material w/o a native speaker to backstop.
October 2, 2026 at 6:12 PM
Young brains learning language are so fascinating! 😂
October 2, 2026 at 6:06 PM
YAHOO! Wherever Egg goes I go, so!

I’m Jade! Teal fish twink who streams during the European summertime!

I looove learning new things, handmade crafts, language barrier collabs and suffering through horror games! I also adore playing games w my friends, and I’m currently undergoing a redesign!
October 2, 2026 at 5:40 PM
Planning to enroll with a language school. I want to get my German to C2. I got to B1 with an enshittified language learning app and a Netflix account. I think I owe it to myself to take an intensive course and master this language so I have a decent future here.
October 2, 2026 at 5:33 PM
That one time i liked a character so much, i started learning her native language
October 2, 2026 at 5:20 PM
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

Hyeonmin Lee et al.

#arXiv #cs.AI #cs.HC
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated…
arxiv.org
October 2, 2026 at 5:09 PM
Large Language Models are AI systems trained to understand and generate human-like text. They fall under the umbrella of natural language processing[..]

#llm #nlp #machine #learning #ai

www.ml-nn.eu/a1/71.html
What Are LLMs?
Machine Learning & Neural Networks Blog
www.ml-nn.eu
October 2, 2026 at 5:08 PM
You're learning Swedish? Jätte bra. My wife is Swedish, as you might've picked up from some of my content, and I've been on-off learning it since - God! - 2009, first with evening/conversation classes, then a gap, then pretty much all of Duolingo. I'm nowhere near fluent but it's a great language.
October 2, 2026 at 5:06 PM
When early learners get help with conflict resolution, they're learning how not to be bullies OR victims. Oct. is #bullying prevention month.
youtube.com/shorts/6FEpv...
@PACERcenter #earlylearning #emotionalintelligence
Teaching conflict resolution skills reduces bullying
A short clip from ep. 23 of the Hidden Language of Children podcast, about bullying.
youtube.com
October 2, 2026 at 5:00 PM
Fixing GRPO's credit assignment problem without evaluating every step | Discussion
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
arxiv.org
October 2, 2026 at 4:40 PM
Inside classrooms at Truman Primary School in Norman, the youngest minds are learning their numbers and letters in both English and Spanish.
Norman school sees success in Spanish Language Academy
Inside classrooms at Truman Primary School in Norman, the youngest minds are learning their numbers and letters in both English and Spanish.
kfor.com
October 2, 2026 at 4:32 PM
Coming up in one week - join us to learn about documenting your language, connect with other folks from Indigenous and minoritized language communities around the world, and grow your own skills in language work! Register (free) at: bit.ly/2026langdoc

#langsky #indigisky #linguistics #languages
October 2, 2026 at 4:54 PM
On Monday, Dr. Brian Holzman from Texas A&M University spoke to CREO affiliates about the effects of “newcomer” programs on English language acquisition and high school outcomes!

Findings suggest that immigrant newcomer programs increase English learning and decrease disciplinary action.
October 2, 2026 at 4:08 PM
Norman school sees success in Spanish Language Academy

https://ift.tt/9RSz7MD

Inside classrooms at Truman Primary School in Norman, the youngest minds are learning their numbers and letters in both English and Spanish.

via KFOR.com Oklahoma City https://kfor.com

October 2, 2026 at 08:58AM
October 2, 2026 at 3:59 PM
Code to Control synthesizes real-time Python controllers using language models, enabling rapid adaptive gameplay without planning. This method outperforms traditional reinforcement learning and manages complex tasks with fewer interactions. https://arxiv.org/abs/2609.38733
Code to Control: Synthesizing Parameterized Reactive Controllers
ArXiv link for Code to Control: Synthesizing Parameterized Reactive Controllers
arxiv.org
October 2, 2026 at 3:20 PM
i spent about six months learning ravens' language

recently they drew my attention to symbols (a kind of arrow) they were deliberately making on the road next to my apartment

so i thought they might be able to read; evidently they can, if you write in their language
October 2, 2026 at 3:15 PM
Books in different languages do more than support language learning. They connect families, amplify diverse voices and make reading more inclusive. theconversation.com/national-yea... via @uk.theconversation.com @drhollyjoseph.bsky.social @sophieheywood.bsky.social
National Year of Reading: linguistic diversity should be central to reading for pleasure intiatives
Books in different languages do more than support language learning. They connect families, amplify diverse voices and make reading more inclusive.
theconversation.com
October 2, 2026 at 3:10 PM
Senior iOS Engineer - Tandem Tandem - Berlin, Berlin, Germany
Senior iOS Engineer - Tandem Tandem - Berlin, Berlin, Germany
About Tandem Tandem is the leading global language learning community, where more than 30 million members from over 200 countries practice languages together. The concept of learning from each other is not only an integral part of the Tandem community but it is also at the heart of our everyday lives at Tandem. This shared connection through learning helps us to grow and enhance the language journeys of our members. We are based in Berlin and are backed by top European investors. We are highly international and driven by a shared passion for languages, technology, and culture. Together, we are on a mission to build a product that not only teaches a new skill but also encourages meaningful cross-cultural conversations. As an Senior iOS Engineer you will work as part of an agile, cross-functional team of designers and engineers to architect, develop and maintain the Tandem iOS app. Your role will focus on strengthening and scaling our current architecture to help Tandem keep shipping high-quality, meaningful features that our members love. Your role will be to: Mature and modernize our existing iOS codebase. Build and ship new product features. Ensure we take full advantage of iOS platform capabilities. Help shape technical decisions and drive engineering excellence. You are the right fit for the role if you: Have 5+ years of iOS development experience with Swift (and ideally Objective-C) Have strong knowledge of the iOS ecosystem: UIKit, SwiftUI, etc. Are experienced with reactive programming and building scalable, maintainable solutions (Clean Architecture, modularization, design patterns) Have worked on production apps at scale (500k+ users) and care about performance and reliability Can lead technical design discussions, act as the iOS expert in the room, and provide constructive feedback Have experience with CI/CD and enjoy improving engineering practices Communicate clearly, collaborate well across teams, and are curious about new technologies What would make your application stand out: Experience building realtime chat and media services Experience with Kotlin Multiplatform Experience with iterative design and working with agile cross-functional teams Experience implementing and optimizing iOS accessibility features The benefits Passion for languages: We’re a team of 30 people speaking 20 languages fluently between us. For those wanting to improve their German (or Denglisch!), you can join one of our German classes on us. Take a deep dive: As a Tandem team member, we’ll give you a budget for up to 6 “Deep Dive Days” per year. Pick a topic you’ve always wanted to explore for your professional development that has never made it out of your backlog. Then book a cabin in the woods or whatever matches your individual learning style. Teams who train together, stay together: We offer Urban Sports Club gym memberships to boost our team’s physical and mental health. If you prefer, we can contribute to your monthly public transport ticket instead. It takes a team to Tandem: Bonding with our teammates outside of our normal work routine is important to us. We organize regular team events: from day trips to monthly lunches, summer BBQs, and movie nights. Balance matters: Life isn’t just about work. We know when to snooze Slack notifications, recharge our batteries, and take time out for our personal hobbies. As home office has become the new normal for some of us, it’s even more important to be able to switch off from work and relax. Modern family: As part of our family-friendly company philosophy, we embrace flexible working hours, welcome families to our events, and love for our kids to visit the office and get a taste of life at Tandem. Relocation without the stress: If you’re looking to relocate to Berlin, we’re ready to assist you with the infamous German bureaucracy. Many of us have been through it before, so we have plenty of tips to share. You’ll feel settled in no time. Hybrid by design: We’re all about embracing a flexible, modern way of working. That’s why we focus on a hybrid approach—balancing time between our vibrant Berlin-Mitte office and the comfort of your home. It’s the best of both worlds, and we’ll support you in creating the perfect setup for this dynamic way of working. About us Tandem is the leading global language learning app, where 35 million members in over 200 countries practice languages together. The concept of learning from each other is not only an integral part of the Tandem community but is also at the heart of our everyday lives at Tandem. This shared connection through learning helps us to grow and enhances the language journeys of our members. We are based in Berlin and are backed by top European investors. We are highly international and driven by a shared passion for languages, technology, and culture. Together, we are on a mission to build a product that not only teaches a new skill but also encourages meaningful cross-cultural conversations.
www.devo.zone
October 2, 2026 at 3:00 PM
Books in different languages do more than support language learning. They connect families, amplify diverse voices and make reading more inclusive.
National Year of Reading: linguistic diversity should be central to reading for pleasure intiatives
Books in different languages do more than support language learning. They connect families, amplify diverse voices and make reading more inclusive.
tcnv.link
October 2, 2026 at 2:47 PM
Taco Bell and Duolingo Launch Experiential 'Global Taco Day' Activation
Taco Bell and Duolingo Launch Experiential ‘Global Taco Day’ Activation
Taco Bell Canada has partnered with language-learning giant Duolingo to scale its annual taco celebration from a national event into a worldwide campaign.
www.princanada.com
October 2, 2026 at 2:36 PM