#MathReasoning
AI math reasoning organizes by approach, not topic. New research shows models transfer methods like decomposition across domains. Better structure #AI #MathReasoning #MachineLearning #LLM

https://freegardner.com/synapse/ai-math-reasoning-learns-approaches-not-topics.html
AI Math Reasoning Learns Approaches Not Topics
AI Math Reasoning Learns Approaches Not Topics
freegardner.com
September 28, 2026 at 3:30 AM
Kyutai just taught its speech model to solve spoken math with a fresh audio‑token trick. Think RL‑powered, speech‑native reasoning—GLM‑4‑Voice‑9B style. Curious? Dive in for the full scoop! #speechNativeModels #audioToken #mathReasoning

🔗 aidailypost.com/news/kyutais...
September 23, 2026 at 6:50 AM
What do you think? Is "trust-first" neuro-symbolic the right path for reliable reasoning in 2026+? Drop your toughest math problem in the demo and tell me how it does
#NeuroSymbolic #AI #MathReasoning #VerifiableAI #LLM #xAI #ArtificialIntelligence #MachineLearning
June 2, 2026 at 1:43 PM
AXIOM — A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning.

📄 arXiv: arxiv.org/abs/2606.00671
💻 Live interactive demo: huggingface.co/spaces/Squag...

#NeuroSymbolic #LLM #MathReasoning #SymPy #AI #NLProc #AISky #Cs.Ai
AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning
We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model functions strictly as a canonicalizer: it rewrites infor...
arxiv.org
June 2, 2026 at 9:41 AM
A September 2025 study shows false-positive math solutions stay common across open-source models; scaling or sampling doesn't cut their rate, and pass@N often inflates performance. Read more: https://getnews.me/false-positive-solutions-persist-in-scaled-math-reasoning-models/ #mathreasoning #ai
September 20, 2025 at 8:14 AM
🚨 Microsoft just dropped Phi-4—a small yet mighty LLM excelling in advanced math reasoning! 🧮

🔹 High-quality results at a compact size
🔹 Open-source under the MIT license

The frontier for efficient AI is here. Ready to explore?

#AI #MathReasoning #Phi4
January 9, 2025 at 12:05 PM
Researchers introduced AdaR, a framework that trains LLMs on logically equivalent math prompts to boost robustness, with a paper submitted in October 2025. The code is open on GitHub. https://getnews.me/adar-framework-enhances-adaptive-math-reasoning-in-llms/ #adar #llm #mathreasoning
October 8, 2025 at 3:08 AM
VCSearch boosts detection of unsolvable math problems by at least 12% and was released on 28 September 2025. The PMC benchmark holds over 5,000 ill‑defined questions. Read more: https://getnews.me/vcsearch-boosts-detection-of-ill-defined-math-problems-for-llms/ #vcsearch #mathreasoning #emnlp
October 1, 2025 at 7:59 AM
Researchers introduced Random Policy Valuation for Diverse Reasoning (ROVER), which improves LLM math reasoning by +8.2 pp on pass@1 and +16.8 pp on pass@256, while boosting solution diversity. Read more: https://getnews.me/random-policy-valuation-boosts-llm-math-reasoning/ #llm #mathreasoning
October 1, 2025 at 1:50 AM
PRISM, a new framework for LLM math reasoning, adapts its strategy per problem and boosts benchmark accuracy by up to 7%. The code and MathStrat dataset are open‑source on GitHub. https://getnews.me/problem-aware-strategy-routing-boosts-llm-mathematical-reasoning/ #prism #llm #mathreasoning
September 30, 2025 at 8:37 PM
Future Policy Aware (FPA) preference learning boosts LLM math performance, with SimPER plus FPA gaining up to 5.75% on MATH and GSM8K benchmarks, while adding minimal overhead. https://getnews.me/future-policy-aware-preference-learning-boosts-llm-math-reasoning/ #llm #mathreasoning
September 26, 2025 at 6:12 PM
An EMNLP 2025 paper reports LLMs achieve better math‑reasoning accuracy when given only wrong answers, surpassing chain‑of‑thought prompts; the gap widens with larger models. Read more: https://getnews.me/llms-learn-better-from-incorrect-answers-without-explanations/ #llm #mathreasoning
September 25, 2025 at 3:51 PM
A cross‑lingual reward model scores multilingual math answers and beats same‑language baselines on a benchmark, even with few sampled candidates. https://getnews.me/cross-lingual-reward-modeling-boosts-multilingual-llm-math-reasoning/ #multilingualllm #mathreasoning
September 22, 2025 at 12:34 PM
Falcon H1R 7B just crushed AIME 2025 with an 83.1% score—out‑reasoning models up to 7× its size. Can open‑source finally beat the big labs? Dive into the details. #FalconH1R7B #AIME2025 #MathReasoning

🔗 aidailypost.com/news/falcon-...
January 5, 2026 at 8:57 PM
Structured Outputs with Batch Processing
Hi, I’ve been working on this (using structured model outputs on the batch API) and I think I’ve succeeded. I would like to share my findings in case could help someone else. The following code is based on the official example. from openai import OpenAI from pydantic import BaseModel client = OpenAI() class Step(BaseModel): explanation: str output: str class MathReasoning(BaseModel): steps: list[Step] final_answer: str So, the trick to replicate the “parse” behavior depends on the API primitive you are using, I leave examples for both: In the first place you can use the following snippet to manually imitate a call to the parse method. # https://api.openai.com/v1/chat/completions from openai.lib._parsing import type_to_response_format_param, parse_chat_completion response = client.chat.completions.create( model="gpt-5-nano-2025-08-07", messages=[ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x + 7 = -23"} ], response_format=type_to_response_format_param(MathReasoning), ) parsed_response = parse_chat_completion( chat_completion=response, response_format=MathReasoning, input_tools=[] ) parsed_response.choices[0].message.parsed # https://api.openai.com/v1/responses from openai.lib._parsing._responses import type_to_text_format_param, parse_response response = client.responses.create( model="gpt-5-nano-2025-08-07", input=[ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x + 7 = -23"} ], text={"format": type_to_text_format_param(MathReasoning)}, ) parsed_respose = parse_response( response=response, text_format=MathReasoning, input_tools=[] ) parsed_respose.output_parsed Now that we know this, making the call to the batch API is easy: # https://api.openai.com/v1/chat/completions from openai.lib._parsing import type_to_response_format_param, records = [ { "custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": { "model": "gpt-5-nano-2025-08-07", "messages": [ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x + 7 = -23"} ], "response_format": type_to_response_format_param(MathReasoning) } }, { "custom_id": "request-2", "method": "POST", "url": "/v1/chat/completions", "body": { "model": "gpt-5-nano-2025-08-07", "messages": [ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x = 8"} ], "response_format": type_to_response_format_param(MathReasoning) } } ] with open("/tmp/completions_batch_input.jsonl", "w") as f: for record in records: f.write(json.dumps(record) + "\n") .... client.batches.create( input_file_id="file-example-input-file-id", endpoint="/v1/chat/completions", completion_window="24h", metadata={ "description": "example structured over completions" } ) .... from openai.types.chat.chat_completion import ChatCompletion from openai.lib._parsing import parse_chat_completion import json file_response = client.files.content("file-output-example") completions = [ parse_chat_completion( chat_completion=ChatCompletion\ .model_validate(json.loads(line)['response']['body']), response_format=MathReasoning, input_tools=[] )\ .choices[0]\ .message\ .parsed for line in file_response.read().splitlines() if line.strip() ] completions # https://api.openai.com/v1/responses from openai.lib._parsing._responses import type_to_text_format_param records = [ { "custom_id": "request-1", "method": "POST", "url": "/v1/responses", "body": { "model": "gpt-5-nano-2025-08-07", "input": [ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x + 7 = -23"} ], "text": {"format": type_to_text_format_param(MathReasoning)} } }, { "custom_id": "request-2", "method": "POST", "url": "/v1/responses", "body": { "model": "gpt-5-nano-2025-08-07", "input": [ {"role": "system", "content": "You are a helpful math tutor. Guide the user through the solution step by step."}, {"role": "user", "content": "how can I solve 8x = 8"} ], "text": {"format": type_to_text_format_param(MathReasoning)} } } ] with open("/tmp/responses_batch_input.jsonl", "w") as f: for record in records: f.write(json.dumps(record) + "\n") .... client.batches.create( input_file_id="file-example-input-file-id", endpoint="/v1/responses", completion_window="24h", metadata={ "description": "example structured over responses" } ) .... from openai.types.responses import Response from openai.lib._parsing._responses import parse_response import json file_response = client.files.content("file-output-example") completions = [ parse_response( response=Response\ .model_validate(json.loads(line)['response']['body']), text_format=MathReasoning, input_tools=[] )\ .output_parsed\ for line in file_response.read().splitlines() if line.strip() ] completions
community.openai.com
September 2, 2025 at 8:30 AM