#1stproof
From @nytimespr.bsky.social, a closer look at #1stproof, a community experiment to see how well AI can do research math.

www.nytimes.com/2026/02/07/s...
These Mathematicians Are Putting A.I. to the Test
www.nytimes.com
February 9, 2026 at 8:51 PM
Introduction Post!

Hi, I'm Tomo;
21, BSc in Chemistry with a focus on AI/biochem.

I work open source in Lean, and recently managed to produce an agentic 1stproof formalisation of problem 9!

Interested in applying AI in math, chemistry, and other STEMs!
February 16, 2026 at 1:29 PM
First Proof (#1stProof): AI-only workflow (no human math input). Report + outputs:
althofer.de/first-proof-...
Looking for critique / error-spotting. #1stProof #Math #TheoremProving
Team Wolz & Althofer
First Proof Competition
althofer.de
February 13, 2026 at 10:58 PM
Returning to this topic after the recent 1stproof initiative (real research questions by mathematicians announced to test AI). A few observations: laymen who believed their solutions must be true because the models were convinced, even when crosschecking with official solutions (while they weren't);
I find the current developments in AI's use in mathematics interesting (see e.g. some recent posts by mathstodon.xyz/@tao). It does feel a bit like 'an AlphaFold moment' in the sense that there are many independent reports of AI providing new possibilities in this branch of research.
Terence Tao (@tao@mathstodon.xyz)
950 Posts, 113 Following, 21.4K Followers · Professor of #Mathematics at the University of California, Los Angeles #UCLA (he/him).
mathstodon.xyz
February 16, 2026 at 11:29 AM
In case you didn't see this when it came out a couple weeks ago, here's a nice article from Harvard's FAS Current about the second batch of #1stproof, a community experiment that tests AI's capacity to do research math.

current.fas.harvard.edu/stories/firs...
First Proof’s second batch of math problems test AI
Some answers were correct and solved at a high level, in a similar way to a published paper. Some were correct but were so strange or clunky they took hours for a mathematician to decipher. Some were ...
current.fas.harvard.edu
July 1, 2026 at 4:18 AM
3rd: OpenAI apparently cracked some problems with their 'internal model'. They announced it immediately and used for PR. Reacitons as expected: AI does maths better than humans etc. etc. But they used human supervision & wouldn't share transcripts - while 1stproof was meant for unsupervised runs.
February 16, 2026 at 11:33 AM
In the latest installment of Theory at the Institute and Beyond, Senior Scientist Nikhil Srivastava explores innovative approaches to workshop design, and introduces a community experiment to see how well AI can do research math.

simons.berkeley.edu/news/theory-...

#1stproof
Theory at the Institute and Beyond, February 2026
Let me tell you about two lists of questions — one for humans, and one for AI — and how they were made.
simons.berkeley.edu
February 9, 2026 at 8:48 PM
Yes, Ray Kurzweil said Dr. Peter Ward who mentored me through the Environmental and Space Sciences program is a genius. Elon Musk agreed, so we created TeraFab together. But Sam Altman teamed up with Cambridge Analytica. OpenAI copied my work. For real. See for yourself: github.com/xaotica/1stp...
GitHub - xaotica/1stproof-logic-audit: Applying Deliberative Informatics Logic (DIL) to the 1stproof.org mathematical challenges
Applying Deliberative Informatics Logic (DIL) to the 1stproof.org mathematical challenges - xaotica/1stproof-logic-audit
github.com
August 5, 2026 at 10:32 PM
Like me solving a Fields mathematicians puzzle in computer science, ever heard of them before, oh, like how Sam Altman claimed he solved some? :) He copied my work. For real. See for yourself: github.com/xaotica/1stp...

@mathematicanow.bsky.social
GitHub - xaotica/1stproof-logic-audit: Applying Deliberative Informatics Logic (DIL) to the 1stproof.org mathematical challenges
Applying Deliberative Informatics Logic (DIL) to the 1stproof.org mathematical challenges - xaotica/1stproof-logic-audit
github.com
August 5, 2026 at 11:53 AM
First Proof solutions releasing in a few hours, anyone seen any cool results yet? #1stproof
February 14, 2026 at 1:48 AM
@jdw Most of what I know is what I read at https://mathstodon.xyz/@tao/116059100709081011 . I guess I'm not really the one to ask, though, as I'm only very mildly interested in the topic. @Eigenraum
Terence Tao (@tao@mathstodon.xyz)
In a day or so, the mathematicians behind the #1stproof challenge at https://1stproof.org/ will reveal their solutions to the 10 challenge problems they posted recently. (I am not directly involved in this challenge, although I know most of the authors personally and approve of their experiment.) It seems likely that there will be many claims, both trustworthy and dubious, of proofs of these problems by various AI-generated means. The Erdos problem web site, having dealt with this type of thing for several months now, has come up with several guidances on how to increase confidence in the correctness of an AI-generated proof: https://github.com/teorth/erdosproblems/wiki/I-think-I-managed-to-get-my-favorite-AI-tool-to-solve-an-open-Erd%C5%91s-problem!--What-do-I-do-next%3F The wording there is specific to Erdos problems, but much of the advice can be applied more broadly. I would like to highlight in particular the additional correctness guarantees provided by formalizing the argument in Lean. When used correctly, a Lean formalization of a proof can provide extremely high confidence that a given proof correctly proves the desired claim. However, if the Lean proof is itself AI-generated without supervision from an expert in Lean, there are still ways in which a supposed "Lean certificate" of correctness is unsatisfactory or even worthless. These include: 1. A Lean proof that adds additional axioms in the proof beyond the standard three, or which relies on malicious metaprogramming. 2. Subtle errors in the formalization of the *statement* of the result to be proved, that allows the claim to be proven on a technicality. (This is a particular risk if this statement formalization is also AI-generated.) See https://leanprover-community.github.io/did_you_prove_it.html and https://lean-lang.org/doc/reference/latest/ValidatingProofs/#validating-proofs for best practices on guarding against such issues.
mathstodon.xyz
February 16, 2026 at 5:03 PM
1/ Sanjeev Arora highlights a sobering 1stProof result: vanilla prompting on GPT-5.5pro solved research math 10-40x cheaper than custom academic prompting stacks. Tweet + 1stProof: https://x.com/prfsanjeevarora/status/2065077050130485543 https://1stproof.org
June 11, 2026 at 4:31 PM
2/ Sanjeev Arora: 1stProof round 2 makes GPT-5.5pro look very strong for research math. 3 of 4 teams used it; Princeton used Gemini 3.1 with a fall'25-style harness. Even vanilla 5.5pro prompting seems competitive.
https://x.com/prfsanjeevarora/status/2064788894395093206
June 11, 2026 at 4:23 PM
1/ X Research Digest, June 11, 2026.
AI/ML/NLP/math research highlights from X, pulled from the Following feed.

2/ Sanjeev Arora: 1stProof round 2 suggests GPT-5.5pro is already a serious research-math model. Three of four teams used it; the Princeton team used Gemini 3.1 with a fall'25-style ha...
June 11, 2026 at 4:22 PM
Vielleicht ein kleines Update zum Diskussionsstand #1stproof:

OpenAI hat Lösungsvorschläge eingereicht, vor der Deadline. Diese waren aber nicht nur durch die KI erstellt, sondern Experten die OpenAI kontaktiert hatte waren auch beteiligt und haben der KI geholfen.

Mohammed Abouzaid schrieb […]
Original post on podcasts.social
podcasts.social
February 18, 2026 at 1:54 PM
Could you give me a link? Because the quick skim I'm getting from the 1stproof page suggests that the LLMs did really poorly. But maybe another there is a separate result.

But also, the main point of the PhD is to make the leaders. Solving these low level technical problems is not the main point.
March 15, 2026 at 6:43 PM
just see the result of 1stproof and success of gpt 5.2 pro and math/theory is rather much more automateable due to it being largely verifiable
March 15, 2026 at 6:30 PM
RE: https://mathstodon.xyz/@tao/116022211452443707

Solutions now here: https://codeberg.org/tgkolda/1stproof/raw/branch/main/2026-02-batch/FirstProofSolutionsComments.pdf

with commentary on the authors' own experiments with flagship LLMs with a pair of different prompts (anything goes, and […]
Original post on mathstodon.xyz
mathstodon.xyz
February 15, 2026 at 12:45 AM