#LongCoT
St Mary the Virgin Church, Longcot, Oxfordshire. #GoodFriday
April 3, 2026 at 8:45 AM
Some moar pics from White horse ride (75km). Superb weather - probably 23C its. Nice coffee in Ashbury at the Baking Bee. Longcot to Uffington road is a good vantage point to make out the whole white horse.

#Oxfordshire #Cycling #Photography #ValeoftheWhiteHorse

bsky.app/profile/andy...
April 8, 2026 at 3:34 PM
The Largest LongCoT Trajectory Dataset with Answers Verified Released Now [35M Trajectories, 1.2M queries [from various open-source resource], using OpenO1-Qwen-7B-SFT]

huggingface.co/datasets/O1-...
O1-OPEN/OpenO1-SFT at main
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 16, 2024 at 11:33 PM
We either are about to discover something amazing that will receive a brotastic name like "abstraction-hyper-grokking" or we have serious case of test poisoning, maybe two fold poisoning. Here are LIMO/s1 results side to side. Same base model, 800/1K SFT on human/LongCoT-Machine highly curated data.
February 9, 2025 at 8:33 PM
The evaluation results show that the best models achieve <10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities.
April 16, 2026 at 3:02 PM
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning

We achieve linear-complexity reasoning. Our "Delethink" decouples thought length from context, matching LongCoT performance with ≈25% of the compute.

📄 arxiv.org/abs/2510.06557
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state i...
arxiv.org
February 3, 2026 at 3:02 PM
Long chain-of-thought works like a molecule: deep steps, self-checks, explorations—held together by different “bonds.” Copy the text and you still miss the structure. go.abvx.xyz/ewbg62
#LongCoT #MechanisticAI #ReasoningModels #Distillation #AIResearch #SyntheticData #ModelDistillation
The Molecular Structure of Thought: Why Long Chain-of-Thought Isn’t “Text” — It’s Topology
Why distillation fails, why “reasoning traces” are a moat, and how MOLE-SYN tries to copy the shape of thought — not the words.
go.abvx.xyz
March 6, 2026 at 8:58 PM
In this paper is introduced LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models.

arxiv.org/pdf/2604.14140
April 16, 2026 at 3:02 PM
via supervised fine-tuning and preference optimization, and the reasoning model leverages Long-Chain-of-Thought (LongCoT) reinforcement learning to improve multi-step code reasoning. Seed-Coder achieves state-of-the-art results among open-source [5/6 of https://arxiv.org/abs/2506.03524v1]
June 5, 2025 at 5:58 AM
"Why it was splendid!" said Frances, and they went on together towards Longcot.

#BevisTheStoryOfABoy #RichardJefferies #Bevis
February 3, 2026 at 8:56 PM
Sketch of the old door, Longcot church, Oxfordshire.
November 20, 2024 at 7:18 AM
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
Large language models (LLMs), such as o1 from OpenAI, have demonstrated remarkable reasoning capabilities. o1 generates a long chain-of-thought (LongCoT) before answering a question. LongCoT allows LLMs to analyze problems, devise plans, reflect, and backtrack effectively. These actions empower LLM to solve complex problems. After the release of o1, many teams have attempted to replicate its LongCoT and reasoning capabilities. In terms of methods, they primarily rely on knowledge distillation with data from existing models with LongCoT capacities (e.g., OpenAI-o1, Qwen-QwQ, DeepSeek-R1-Preview), leaving significant uncertainties on systematically developing such reasoning abilities. In terms of data domains, these works focus narrowly on math while a few others include coding, limiting their generalizability. This paper introduces a novel approach to enable LLM's LongCoT capacity without distillation from o1-like models or expensive human annotations, where we bootstrap LongCoT (BOLT) from a standard instruct model. BOLT involves three stages: 1) LongCoT data bootstrapping with in-context learning on a standard instruct model; 2) LongCoT supervised finetuning; 3) online training to further refine LongCoT capacities. In BOLT, only a few in-context examples need to be constructed during the bootstrapping stage; in our experiments, we created 10 examples, demonstrating the feasibility of this approach. We use Llama-3.1-70B-Instruct to bootstrap LongCoT and apply our method to various model scales (7B, 8B, 70B). We achieve impressive performance on a variety of benchmarks, Arena-Hard, MT-Bench, WildBench, ZebraLogic, MATH500, which evaluate diverse task-solving and reasoning capabilities.
arxiv.org
February 7, 2025 at 5:08 AM
The next branch meeting will be at the King & Queen, Longcot on Tuesday 10th June at 7.45.
June 4, 2025 at 8:47 PM
From @TVPRuralCrime
PC Dollery is out on patrol with rural crime prevention officer Michaela Andrews today 👮‍♂️

The officers have attended a rural burglary non-dwelling in Longcot. Nothing was stolen however it is suspected the offenders were looking for power tools 🧰 Please regularly revie...
September 23, 2025 at 2:49 PM
Shoutout to the authors: Kamran Chitsaz, Milad Aghajohari, @a-kazemnejad.bsky.social. Supervised by: @sarath-chandar@bsky.social, @murefil.bsky.social, AaronCourville and @sivareddyg.bsky.social

🔗 Learn more at: arxiv.org/abs/2510.06557
🔗Build with: github.com/McGill-NLP/the-markovian-thinker
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state i...
arxiv.org
February 17, 2026 at 3:54 PM
🧩 Even state-of-the-art models show Markovian Thinking at zero-shot: both GPT-oss-120B and Qwen3-30B-A3B recover/track LongCoT with no special prompting/training required, and lots of in-distribution positives on initialization, so RL with Delethink is primed to scale!!
February 17, 2026 at 3:54 PM
Markovian Thinking is instantiated by Delethink, an RL enviroment. With it, we trained DeepSeek R1-1.5B and demonstrated:

1️⃣ The same scaling as LongCoT-RL, but at lower costs,
2️⃣ Better test-time scaling, improving past 24K tokens, while LongCoT-RL plateaus.
3️⃣ All this while keeping linear costs!!
February 17, 2026 at 3:54 PM
School praised for 'relentless ambition' to ensure pupils reach potential

#Oxon #Oxfordshire

🔗: https://www.oxfordmail.co.uk/...
School praised for 'relentless ambition' to ensure pupils reach potential
Longcot and Fernham CE Primary School has been praised for its "relentless ambition" to ensure all pupils "fulfil their…
www.oxfordmail.co.uk
September 12, 2025 at 5:02 AM
New benchmark exposes critical AI weakness. LongCoT tests long chains of thought across math, chemistry, and logic. Best models score under 10% despite each individual step being tractable. The gap reveals models cannot reason reliably over extended horizons. https://arxiv.org/abs/2604.14140
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve &lt;10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.
arxiv.org
April 17, 2026 at 11:24 PM
Researchers reveal a critical gap in AI capabilities with LongCoT benchmark. Frontier models achieve under 10% accuracy on problems requiring extended reasoning chains, even though each individual step is tractable. The issue is sustained coherence over long horizons.…
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve &lt;10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.
arxiv.org
April 17, 2026 at 11:24 PM