OpenAI Withdraws 3 Math Papers

(github.com)

228 points | by theemathas 6 hours ago

28 comments

  • qoez 2 hours ago
    Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.
    • flowerlad 4 minutes ago
      This is not a problem unique to the mathematical community.

      What about software developer community? AI has eliminated the need for junior software engineers. Almost no one is hiring junior software engineers. But companies still need senior software engineers. Without junior engineers how will there be senior software engineers in the future?

      What is the solution? I don't think the solution is to say AI progress in software, mathematics etc. should be halted.

    • palmotea 13 minutes ago
      > Without a thriving mathematical community to point out these things it would have stayed broken.

      And that's a bad thing. If that community didn't exist or was weak, OpenAI would still benefit from the prestige of these results they were forced to withdraw. Withdrawing these papers has harmed OpenAI's investors, and that's totally unacceptable.

      > With automated math that community as tao pointed out is at risk.

      Good to hear.

    • viraptor 2 hours ago
      Have you got a link to someone pointing it out? It looks like they're still going through formal proofs so likely found the problems that way.
      • sanxiyn 2 hours ago
        • viraptor 2 hours ago
          So someone ran a different LLM to find an issue they'd find anyway during formalisation? That's not the same as relying on thriving community.
          • dominotw 13 minutes ago
            Is that guy just some rando "someone" though?
          • oliculipolicula 2 hours ago
            The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.

            This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.

            • yorwba 1 hour ago
              What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.
              • verdverm 2 minutes ago
                > What makes you think that nobody checked the [ai output] first?

                Because this is what they say all the time. It's like a badge they have to wear and tell everyone they are wearing, even though we see it.

                You can see the same thing with ANT. Had they looked at Mythos output, they would have realized there were only 76 items, not 79 like the bot claimed. Or the ones that were just a "it crashed" and nothing else.

                https://www.youtube.com/watch?v=NnV_cWeoo5Q (Linux Kernel team sharing their side of the Mythos "hacking" story)

              • oliculipolicula 1 hour ago
                Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.
                • dragonwriter 55 minutes ago
                  LLMs are unpredictably complex with potentially sigbificnaly different reaults based on random seed and seemingly insignificant prompt details) in the ideal case and nondeterministic in practice, so someone finding an error with a given LLM is not strong evidence that the result was not checked with an LLM, even with the very same LLM, previously.
                  • computerex 5 minutes ago
                    No, not modern foundation models. This isn’t gpt-3.5-turbo. Although they are causal autoregressive, they have self consistency. You just have to verify multiple times to ensure you have averaged out any sampling errors.
                • Hamuko 52 minutes ago
                  I can throw Opus 5.5 at my code three times for code review and get three different sets of things it considers to be issues. I imagine all of them were checked with Astra at least once, but were they checked enough times?
              • idiotsecant 1 hour ago
                Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO.
            • kolinko 1 hour ago
              My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.
              • oliculipolicula 1 hour ago
                No. Someone received a tip, presumably from inside OAI. He then used OAI Astra to check..
          • verdverm 7 minutes ago
            I would say

            1. This is expected if you only use a single model family like Claude, eg. we use a different model family for code review than authoring, OAI could have done this too for their math dump

            2. Ai needs a good human driver beyond the trivial or mundane, they are expert enhancing machines, not expert creating machines. This is where the community comes in. Reading Tao's ChatGPT session reveals this: https://news.ycombinator.com/item?id=49010345

            3. OAI is not trying to be a member of the/any community, this is not the first story to shows this, nor do I expect it to be the last. Perhaps this is them being effective altruists today? /s

        • afavour 2 hours ago
          That’s not proof though is it? If the original LLM output is fallible surely the LLM review of that output is also very much fallible?
          • rsfern 2 hours ago
            Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).

            I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.

            I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here

          • kolinko 1 hour ago
            Output of LLM can be infallible* even if LLMs themselves make mistakes. Ditto with humans.

            As much as anything can be infallible.

    • thurn 44 minutes ago
      Is the risk here like "people will fund math research less because of AI"? I agree that would be bad, and we should try and stop it (along with e.g. funding for the humanities, which is in a much worse place than math!), but I'm not sure OpenAI are the right people to be mad at.
      • svnt 31 minutes ago
        But they are the ones releasing an unverifiable (no model release) and massive and unchecked body of mathematics into the public while making exaggerated claims about its capabilities to replace human work. And they are the ones doing it in advance of a fractional sale of the company to the public.
      • viccis 31 minutes ago
        Based on what I've heard from my math friends in academia, every talented undergrad who was set on going to grad school for math has switched to something like consulting internships or fintech, even if they were really passionate about math, because they don't want to spend another 7 years or so just to wind up jobless.
        • done_lurking 11 minutes ago
          This has always been the risk with majoring in math. AI is making it worse but there was never a time where studying math didn't have an extremely high opportunity cost.
          • viccis 3 minutes ago
            AI obliterating your field of study was most definitely not considered a risk with majoring in math 10 years ago.
    • nilkn 13 minutes ago
      I'm sorry but you don't actually need the mathematical community for this. You just need autonomous verification, which already exists at scale and speed vastly beyond the entire human mathematical establishment. OpenAI simply rushed these out without completing that for every paper. These mistakes have nothing at all to do with the mathematical community and are 100% just the result of market pressure incentivizing speed at 1,000,000x the pace of human mathematicians. Frankly, a couple mistakes, trivially found not by humans but by humans using AI, is almost completely irrelevant. Human mathematicians have essentially nothing to contribute to this effort aside from prompting verification agents, which any of us can do if we cared to spend the time (most of us don't).
      • tylerhou 10 minutes ago
        > You just need autonomous verification, which already exists at scale and speed vastly beyond the entire human mathematical establishment.

        An absolutely ridiculous statement. There is a vast amount of mathematical knowledge that hasn’t even been written down, much less formalized.

        • nilkn 8 minutes ago
          There is no such thing as knowledge that has never even been written down a single time by anyone or anything. Beyond that, if we assume that such knowledge does exist, its marginal value is clearly nearly zero now, as AI systems without access to it are vastly outperforming all human mathematicians combined by multiple orders of magnitude.
    • caaqil 1 hour ago
      > With automated math that community as tao pointed out is at risk.

      If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.

      • kec 1 hour ago
        This statement assumes local maxima don’t exist and greedy short term optimization always leads to long term benefit.
        • caaqil 13 minutes ago
          > assumes local maxima don’t exist and greedy short term optimization always leads to long term benefit.

          Even if your stated assumption was baked into the original comment, which is doubtful: the historical record shows that we will keep relearning The Bitter Lesson and each community will pretend what they do for a living is exceptional and immune because of xyz. The screams will get louder when the "greedy" and "dumb" automation comes knocking and it turns out nothing was truly immune or "nuanced ".

          Getting some new hobbies may be in order, it's a Brave New World.

      • doc_ick 1 hour ago
        “If they can be automated, they are not necessary.” That’s a pretty interesting take as eventually everything could be automated.
        • palmotea 8 minutes ago
          > “If they can be automated, they are not necessary.” That’s a pretty interesting take as eventually everything could be automated.

          It's great that we're starting to see the light at the end of the tunnel, and will some day achieve a perfect market without humans. If you think about it, all the market really needs is a people to own everything, everything else can be automated, and all those annoying human workers can be eliminated.

        • spidersouris 57 minutes ago
          I guess what OP meant is automated and with 100% accuracy.
      • lkey 1 hour ago
        You've abandoned every part of yourself to the siren's song of efficiency and automation, huh?

        To witness an arson and rejoice reveals an ugly kind of sadism.

        • falcor84 31 minutes ago
          Will it also be "arson" if OpenAI dump on us solutions to more practical problems, for example, a blueprint for a better photolithography machine, or a nuclear power plant?
          • nancyminusone 5 minutes ago
            No that would be irony. Use up all the chip foundries and power plants to make more plans for chip foundries and power plants so we can build more chip foundries and power plants and use them to run AI to design more chip foundries and power plants.
    • aaron695 1 hour ago
      [dead]
    • giancarlostoro 39 minutes ago
      It's amazing that people long long ago figured out so much math.
  • ijustlovemath 41 minutes ago
    I think that on closer inspection, a lot of these fully AI generated proofs will fall apart. Even in Lean, you can build theories which compile but nonetheless state something different than what you actually intend. It's just that the volume of proof is so staggeringly large that it will probably take years before we find the issues, a la abc conjecture
    • nickysielicki 17 minutes ago
      1. why can’t the so called “real mathematicians” (as if mathematics does not belong to all of us) write the lean theorems by hand and then we let the machine fill the rest in? They are very upset that they can no longer contribute to frontier mathematics. This would let them contribute.

      2. Why shouldn’t math progress happen in the open, commit by commit? Why is it so horrible if a proof is 95% of the way there but we later find that it needs to be refined? Mathematics previously was optimizing for an antiquated publishing and distribution scheme. There is no need for the first print to be correct. We have the internet now. We can and should publish incomplete results and correct things on the fly. Maybe mathematicians would have solved some of these problems years ago if they didn’t hide incomplete almost solutions in their filing cabinet because it wasn’t yet ready to be published.

      You don’t hate the pageantry of mathematics and academics enough.

    • cma 33 minutes ago
      Many of the statements were already there and looked over by the community in lean prior to the work though, the statement can get formalized before the proof of it.
      • ijustlovemath 29 minutes ago
        I just think that with the vast amounts of compute involved and the tendency to reward hack, we can't assume the steps towards that formalization are without error until full human understanding of the formalization.
        • mkarrmann 11 minutes ago
          Repeating the above comment: most of the statements were already formalized prior OpenAI's work. So no, the statements were not "reward hacked".
  • ThePhysicist 1 hour ago
    I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.
    • mtlmtlmtlmtl 53 minutes ago
      I'm not familiar with the paper you mention. But it's also worth pointing out that afaik the n log n algorithm itself isn't particularly practical. It's one of these "galactic algorithms" that is asymptotically more optimal, but is so complicated that it's only a real improvement for comically large n. And that's without even considering the mental overhead of implementing and maintaining the thing.

      Of course, that's not to say the research is necessarily useless. It's still theoretically interesting to find "better" algorithms if only to shed some light on lower bounds, and so on. And who knows, maybe the line of research could lead to more practical algorithms later on.

    • IsTom 38 minutes ago
      I'm no expert, but using multitape TM for this feels to me like a wrong level of abstraction. The whole thing has a lot of smell.
    • kolinko 1 hour ago
      I didn’t read the paper but can’t this just be additionally with doing the actual muls? Or was it a nonconstructive proof?
      • peri-cl 1 hour ago
        Surely it's a galactic algorithm that you can't physically run? You wouldn't get a constant as small as 2^{-182} without some other numbers elsewhere being incredibly large.

        From the "Introduction" section of that paper: "The constants and thresholds in the construction are extremely large".

        (And verifying if the algorithm multiplies correctly or not is the less-interesting part of this, anyway. Gets you no closer to verifying the complexity result).

        • hyperbovine 56 minutes ago
          Isn't it possible that all of the integers that have been or will ever be encountered, anywhere, any time, in human history, number less than 2^182? In which case you could argue that integer multiplication is O(1) via LUT :)
        • auspiv 52 minutes ago
          someone else picked up that bit of math and has run with it and has refined it downwards multiple times. believe it is now in the range of 2^-18 or so
  • margorczynski 2 hours ago
    If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.
    • andybak 1 hour ago
      Not a mathematician but surely if a problem I was working on had an AI also working on it, I would want to know as early as possible - even with flaws or gaps. What advantage is it to me to be less informed?
      • baobabKoodaa 1 hour ago
        I can prompt ChatGPT right now and ask for mountains of more "mathematical work"; thousands and thousands of pages of nonsense for you to review. So you can "be informed".
        • cma 30 minutes ago
          But you can't do it with their internal model that is the same or a successor to the one that solved the navier stokes millenium prize problem.

          With 40% formalized they probably have a good idea of how many were found to have fatal issues in the formalization attempt, and they hired some mathematicians to verify some of them, especially the big headline ones.

      • ajs1998 45 minutes ago
        Mathematicians are in no rush. And they are less likely to review math vomit that hasn't even been formalized and verified. Especially among the now thousands of vomit papers out there.
      • ziiinq 1 hour ago
        [dead]
    • DonsDiscountGas 1 hour ago
      Apparently the write-ups are garbage (as in very hard to read). I feel like they could've had AI fix that up at least somewhat. Maybe they'll reinvest more in writing ability now.
      • margorczynski 26 minutes ago
        The write-ups are one thing and more of a cherry on top but I would say the more pressing matter is the lack of Lean formalization which means you can't really say it has been (dis)proven or not.
      • idiotsecant 1 hour ago
        I think the obscurity is a feature, not a bug. They don't want a headline where 250 of these results are invalidated overnight. They want rejections to trickle out and be buried.
        • baobabKoodaa 1 hour ago
          Not sure why you're being downvoted, because this is exactly what I think is happening.
    • shironandonand 1 hour ago
      Almost like they care less about the maths and more about a deadline for marketing purposes.
  • jfyi 10 minutes ago
    This is just OpenAI stealing more work.

    I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.

    All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.

  • kevin42 17 minutes ago
    The papers were all in a 'preprint' directory in GitHub. Academic mathematicians seem to think all research should happen in secret until it's completely proven.

    But how much more progress would have been made in the last 100 years if they collaborated in real time? Watching the work someone is doing and spotting errors, or making suggestions would be a good thing. Unless the goal is simply to claim credit for a discovery vs the discovery itself.

  • blablabla123 24 minutes ago
    I think this is now an interesting part about LLM based automation. The scientific community is quite strict on references. Even with its most advanced model ChatGPT isn't reliably able to tell me if an online shop has an item in stock. Then the other part is peer review which isn't optional either.
  • TrackerFF 2 hours ago
    And added 6 new ones. Might want to add that to the headline.
    • gabrimat 1 hour ago
      Looks like those are new formalizations of existing proofs, not new ones
  • sehw 8 minutes ago
    OpenAI would never lie. HN can't be this stupid.
  • zkmon 1 hour ago
    Let the evolution run its course. Let the stream find its way. Let the forces restore the equilibrium.
  • hmate9 2 hours ago
    3 mistakes (so far) out of ~400 is still a pretty good hit rate
    • afavour 2 hours ago
      I’m conflicted. I guess we’ll see what the final total is once an enormous level of unpaid human effort is expended verifying the AI outputs. A little sad if that’s the future of math.

      It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”

      • ndriscoll 2 hours ago
        Where are you getting unpaid from? Almost all people who are qualified to analyze the results are paid researchers. And if it is unpaid, then it sounds like they're looking it over for their own reasons, and that's fine?

        I also fail to see the issue you have with releasing abandoned source. In what world is that bad? That obviously is a gift and should be encouraged. e.g. id software's history of doing that has meant their work stays alive forever.

        • afavour 1 hour ago
          Unpaid by OpenAI. In service of their PR goals they’re creating a bunch of work for others and not contributing to compensating that work.
          • WarmWash 57 minutes ago
            No one has any obligation to ever look at anything OAI publishes.

            Ironically though, what I imagine will happen, is that the researchers will pay OAI to use chatGPT to help eval the proofs.

            • tffrt 38 minutes ago
              This is a comment in poor taste.

              You have the frontier labs who are marketing that it’s over and they’re building intelligent machines and you’re saying the mathematicians should ignore it? Ok

              • WarmWash 21 minutes ago
                I'm not sure how you construed that, so let me be more precise.

                Mathematicians, as autonomous entities with no formal connection to any AI lab, have zero obligation to do any work for those labs. OAI can't do anything if all the mathematicians band together and say "Sorry, we're not interested".

                If they do chose to engage, they are doing so entirely voluntarily, and it would strongly indicate, if they are voluntarily doing it for free, that there is value (i.e. compensation, payment, barter, worthwhile, whatever) to be had by digging in.

              • charcircuit 19 minutes ago
                Whether you should do something and whether you are free not to do it are different things.
        • _aavaa_ 2 hours ago
          OpenAI isn’t paying mathematicians to verify all these papers. It’s dumping the papers trying to nerd snipe them into checking it for free.
          • ndriscoll 1 hour ago
            Researchers normally don't pay each other to read each other's work. It's a symbiotic relationship. If they think the AI results are nonsense, they could ignore it like any other crank. If they think it seems plausible and it's relevant to them, they can try to understand it. Seems fine?
            • PunchTornado 1 hour ago
              yes, but in this case, openai does a lot of PR how their models solve important math problems. if it goes unchallenged, parts of society would think that is true. what would happen if we scale it and 10k companies dump 10k papers every month claiming solved math problems. how is this scalable?

              We need the companies to humanly review their papers. in the same way as at other companies we use humans to review the papers.

              • ndriscoll 1 hour ago
                Well, per another comment in the thread, some 20% of their solutions come with formalization, so there's a very high chance (probably higher than typical asks of research mathematics) that they did solve the problem. And that also presents a pretty easy solution to the scaling issue: demand formal proofs.

                (If you're going to object that it's difficult to validate the statement of the problem, please first state your level of experience doing so. It's getting tiring seeing people raise this objection and claim that a statement is just as hard as a proof over and over who don't seem to actually know any math and have never tried to write anything in Lean)

                • _aavaa_ 1 hour ago
                  Formalization only helps you if you prove that the formalization is correctly implemented, that the model didn’t subtly mess up or cheat.

                  Someone still has to read the formalization.

            • cowpig 1 hour ago
              > It's a symbiotic relationship

              That's right, and the difference is that this one is parasitic.

              • ndriscoll 1 hour ago
                Why's that? Do you think that institutions and funding agencies won't cover researchers' use of advanced models, or what? The group I worked in in undergrad had millions of dollars of equipment for doing experiments. I'd have to imagine they could get funding for a few thousand dollars in tokens for the theoreticians to have AI assistance.
                • fwip 56 minutes ago
                  That's just money flowing into OpenAI, from the people who used to pay mathematicians.
                  • ndriscoll 36 minutes ago
                    Okay? Again, when I was in my group, money would flow to Nvidia to buy GPUs that we would use to run simulations. The researchers could also request money to use for openrouter or build their own GPU cluster if everyone thinks it's good use of funds.

                    Shouldn't we pay for the best tools if it helps researchers to be more effective?

                    • WarmWash 9 minutes ago
                      You are arguing with people who have AI do 80% of their job yet are still convinced that the labs are the ones capturing all the value.
            • ziiinq 1 hour ago
              [dead]
        • rfgplk 1 hour ago
          > Almost all people who are qualified to analyze the results are paid researchers.

          This _might_ have been true somewhat in the past (although it wasn't), but it's completely false today. Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs. It's no different than reading a codebase you might not be fully familiar with, and checking it for correctness (give an engineering analogy).

          This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.

          • ndriscoll 1 hour ago
            I mean, I have a decent math background, but I would struggle greatly to attempt to even tell you what most (or any) of the conjectures in e.g. number theory are about at even the highest level. I can't imagine a layman would have any hope.

            Like I was reading some about adele rings last night, which is already going to be quite a concept for a layman to be able to even slightly describe. Then you can layer on that apparently they're locally compact, so we can talk about harmonic analysis on the additive group. Like, come on now, 99.99% of people have no hope of ever following along, and this is stuff from 75 years ago.

          • latexr 59 minutes ago
            > Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs.

            But they don’t. What they have is the ability to ask something else to do the analysis. It’s an important distinction. If the asker has the skills to evaluate the results, that’s one thing, but too many don’t and act as if whatever they got is unambiguous truth.

            > This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.

            What must stop is the overuse of the word “gatekeeping”. Anyone is free to study these fields and work on problems. What people rightfully object to is uninformed research flooding everything with hard to verify junk.

          • ziiinq 1 hour ago
            [dead]
      • charlieyu1 1 hour ago
        I’m unsure. The whole academia is built on unpaid human efforts. Journal writers are unpaid, and institutions paid for their papers to be published by for-profit publishers. Journal reviewers are paid the bare minimum, certainly unproportional to their efforts and expertise.
      • rubyfan 2 hours ago
        We are all reverse centaurs now.
    • watinthedeutsch 2 hours ago
      I would say not. For a mathematicians, having to retract more than 2 papers in a lifetime is already a big issue in their career.
      • zeroonetwothree 12 minutes ago
        It’s fairly common to have errors in preprints.
      • jacobstokes 2 hours ago
        But is it equivalent to withdrawing post-publication or is it more akin to not passing peer review with major revisions requested?
      • levocardia 1 hour ago
        But have you formalized every single one of your results? And if you haven't what odds would you put on one of them not working out, if formalized?
      • zzzeek 2 hours ago
        Most mathematicians don't produce 400 papers in a 48 hour window either so I'm not sure comparisons are helpful
        • zerkten 1 hour ago
          The product at the end is important, but so is the process. Few of the things that would happen along the way are happening here, so it's harder to justify the value of the deliverable when there is a failure.
          • gorszon 32 minutes ago
            Exactly, the scientific method is what it is for a reason. Sure the institution of academia around it is not perfect, but in general, science, especially general research such as this is not about just bragging about how many papers you have published, it is a process that might help us find out things we might have not known otherwise.
    • jansport123 2 hours ago
      well, it might be that these proofs are correct or it might be that people aren't bothering to spend a lot of time checking whether they are correct. OpenAI already has a pretty bad reputation in the mathematics community for how they are approaching this process, they seem to be more interested in creating a story for their IPO than advancing math.
    • gorszon 35 minutes ago
      Except Mathematics and science is not done this way, its not code that you can just release bugfixes to, its not a numbers game, its about furthering our shared knowladge. If it is done by flooding everything with a bunch of paper that have not been peer reviewed and verified, and are known to be error prone, that just takes away a bunch of mental capacity from scientists, and time, to manually verify all 400 of them. The issue here is how openAI approaches the science, not their hitrate.
      • lostmsu 28 minutes ago
        I got the impression that indeed mathematics is done exactly this way. People used to write proofs, somebody would find issues in them that don't invalidate the entire work (or they do), the mathematician would work to correct the issues and resubmit.
    • levocardia 1 hour ago
      I bet it's a lower rate of errors than the typical human math paper of 2025
    • matsemann 2 hours ago
      But can the others even be "disproven", given that they apparently are so messy and awful that no humans can follow them? Shouldn't the onus instead be on OpenAI to prove that they're right, instead of hundreds of mathematicians wading through slop?
      • true_religion 2 hours ago
        The onus is on formal verification when it comes to computer generated results. So far only 22% of the papers have it.
    • malux85 2 hours ago
      Exactly, I was glad to see these withdrawals, its a natural part of a healthy ecosystem of scientific review, hypothesis, claim, test, refute, extend, withdraw, its the heart of science.

      IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going

    • an0malous 37 minutes ago
      It has literally not even been a day yet
    • catlifeonmars 2 hours ago
      How many mathematicians need to retract ~1% of their papers?
      • ModernMech 1 hour ago
        Retraction and withdrawal are different. Retraction is when you publish something, it passes peer review, is published, and some time later its publication is undone, often by an editor or some other person because some fraud was uncovered.

        Withdrawal is akin to submitting a paper to peer review and then when you’ve noticed mistakes, you decide to take the paper back and correct it.

        Reject is when someone else notices the mistakes and tells you to take it back and correct it.

        Withdrawal and reject happen all the time in a scientist’s career. They don’t necessarily mean the scientist is doing bad research, just the research was not ready. Retract usually means something more.

        By dumping the papers, OpenAI skipped the typical peer review process, so peer review should be understood as what’s going on now as mathematicians look over the papers and find flaws.

  • fantasizr 1 hour ago
    flooding the system with parts that may be incorrect hurts the whole process and will get people to tune out (like politics). Can't see the International Mathematical Union making a similar mistake because it would do reputational harm. But the models don't care about their reputation. Bad for the layman - like me - to know what to make of all this.
  • fancyfredbot 17 minutes ago
    When You See One Cockroach, There's Probably More.

    This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.

    I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.

    • vlovich123 4 minutes ago
      Fantastic. Let’s also apply this reasoning to math papers written by humans too?

      Humans produce flawed Lean proofs. Indeed LLMs were successful at finding and fixing many issues in the “core” standard library if I recall correctly. Humans regularly produce flawed papers and have minor issues require fixing. And when it happens it often isn’t as prompt and clear as this.

  • soltanov 3 hours ago
    Proof by authority works until human mathematicians actually run the code. Back to prompt engineering.
  • chairhairair 2 hours ago
    This company is just irresponsible. We (at least, Americans that vote and can therefore decide indirectly what’s legal) should not allow them to continue.
    • monideas 2 hours ago
      This type of sentiment is almost always motivated by fear over the potential negative personal economic impact from AI (e.g. losing employment).

      Publishing a math paper and then unpublishing it is not "irresponsible". It's just a math paper.

      • chairhairair 1 hour ago
        Irresponsible is an adjective that means lacking a proper sense of responsibility, or acting without thinking about or caring about the potential consequences of one's actions.

        They do not care if they waste everyone’s time or flood the common with slop. They did not spend their own time to verify that the Lean proofs correspond to the natural language proofs. They did not even spend their own time to verify that all of these proofs are well written.

        They instead are mining the unrenewable resource of open math problems.

        However, seeing it another way is easy if you are financially motivated by their upcoming IPO (see how easy it is to invent motivations for comments?).

        • vessenes 28 minutes ago
          Actually they sat a highly illustrious group of mathematicians down to advise them before releasing -- illustrious enough that I am certain that they have more global standing than you do to make a good call. That group advised open release, which is what they are doing. It lets mathematicians interested and skilled in each area look at what we have. It is by no means irresponsible.
      • idiotsecant 1 hour ago
        It is irresponsible when nobody checked it for correctness first. They just copy pasted what the slop machine spit out.
    • trio8453 2 hours ago
      How is withdrawing a paper irresponsible?
      • jansport123 2 hours ago
        it's like if you vibe coded something and the onus is now on the reviewer and the reviewer tells you your work contains bugs and is messy - that is not acceptable from the reviewers pov - why should the reviewer spend all his human effort, a scarce resource, reviewing your code while you've spent barely a fraction of his effort generating this. OpenAI is a trillion dollar company, surely they can verify stuff before pushing it out? The problem is not that they are solving the problems, they don't care at all about the actual process of doing mathematics. Using compute to mine problems and throwing results in github and letting human reviewers spend effort to correct these is not going to win them any favor. If OpenAI really cared about math, they would have someone on their side who actually understood the results they produced and was able to verify their correctness and educate others.
      • moregrist 2 hours ago
        They only withdrew it after someone pointed out that it was a flawed paper.

        In one case by asking Astra to review it.

        Not exactly encouraging that they did their homework before publishing results.

        • crimsoneer 2 hours ago
          This is how science is meant to work.
          • mahogany 1 hour ago
            No, it's not how math works at all. Math is not experimental science. Results take years to formulate and are generally known to be true by colleagues via correspondence and seminars before a paper is even written. Retraction is very rare.

            I don't know why people are talking about something like this that they are clearly not familiar with.

          • moregrist 1 hour ago
            Most good scientists work hard to prove results to themselves before going public.

            Peer review is then done _in private_ before publication as a check on quality and significance.

            Retracting a paper is pretty embarrassing. And not considered science as usual.

            • nimonian 1 hour ago
              The greatest proof ever written by man, Andrew Wiles' FLT, was published with a flaw before being retracted and reworked. It's _completely_ normal in mathematics to find flaws in the argument. That's what peer review in mathematics is _for_.

              Besides which it's not clear if these papers are considered published or preprint since they appear in no journal, so it's not really a retraction.

              It's also very very normal to post preprints on Arxiv before peer review, so it's not the case that mathematics is kept private during review.

              • mtlmtlmtlmtl 21 minutes ago
                Not true. Wiles presented his proof in a series of lectures in 1993. The error was found by a reviewer, and he spent over a year trying to fix it. Once he did fix it, he asked colleagues to review it again, and only once it passed that review did he actually publish it.
            • rpdillon 1 hour ago
              Retraction is after publication and peer review, no?
          • dwattttt 1 hour ago
            Yes. A researcher who achieves a reputation of making mistakes carries that reputation.
          • idiotsecant 1 hour ago
            No, scientists check their work before releasing it. Slop is still slop if you wrap it up as an academic result.
      • spider-mario 2 hours ago
        Withdrawing it, in itself, might not be, but it kind of skips over the part where they have a paper that warrants it in the first place.
      • mtlmtlmtlmtl 1 hour ago
        Withdrawing it isn't irresponsible. I don't even think releasing it is, because that is often the best way to find problems with it; expose it to the wider research community. Doing PR victory laps based on unverified, unfinished work, on the other hand, is. It is well known, even scientifically demonstrated, that retractions or refutations get much less attention than the initial exuberant PR announcement. And thus this practice contributes to misinformation.

        And just to preempt the kneejerk whataboutism: Yes, all of academia does this to varying extents, and the mainstream media are also complicit. And that is also irresponsible. And no, that is not an excuse for OpenAI. Especially when you consider that OpenAI actively portray themselves as some kind of moral arbiter on AI and "doing good for humanity". They should be held to the extraordinarily high moral standards they purport to hold themselves to.

  • xvxvx 1 hour ago
    Picture a remake of Good Will Hunting, where Will is an AI and, instead of getting the mathematical formulas correct, he just mass dumps a bunch of nonsense and the professors have to go through it all, pointing out where it is wrong. The professors know that AI Will isn’t as smart as people say, but their funding depends on it, and if they prove it, which they can do easily, it may just tank the whole economy, causing a depression, all because they have history’s dumbest President in office.
  • quantum_state 2 hours ago
    It would turn out to be a pure energy and time wasting exercise … the math community would want to keep away from it.
    • OutOfHere 1 hour ago
      Some authors in the math community already use AI to write papers.
  • fluidcruft 2 hours ago
    Whose names are on these papers?
    • kant2002 1 hour ago
      They try to be clever and put company name as the author.
      • baobabKoodaa 1 hour ago
        Not a great sign when nobody wants to attach their own name to words thrown into the world.
        • fluidcruft 1 hour ago
          I guess I feel two ways about this. It does could be a "working with the garage open" sort of thing where these are known to be unverified and they're just letting people see the sausage being made.

          But that also means marketing and PR should shut the fuck up until things are verified. And they should more clearly annotate lack-of-verification status on their repo.

          • kant2002 50 minutes ago
            I believe this is nice marketing trick. Don't have human names have following effects - AI can work alone and produce value - There no person whose career on the line, and who is responsible. - You can skip lot of time on verification of papers, and just show as it is.

            Goals of marketing people obviously don't aligned with science goals.

            • fluidcruft 34 minutes ago
              Sure, unverified claims are unverified so there's a real risk of reputational tarnish. Why is it anyone else's job to verify a papermill? Would any journal accept an article that doesn't name people who are accountable? It should correctly be viewed as abuse of peer review. Let them rot.
  • gt565k 1 hour ago
    All that AI compute and resources and they couldn't hold onto releasing the papers until a legitimate peer review was performed?

    Just had to get that PR stunt out to bump their valuation.

    Vibe coding math research is just next-level AI slop.

    Mind you, this is a competent AI company that's making these mistakes.

    I can only imagine what non-technical people are putting out in production via vibe-coded AI slop.

    • WarmWash 45 minutes ago
      People keep calling this "a PR stunt", but from a PR perspective, it seems they went about it in the least PR way possible.

      No announcement, no press blast (big orgs have press contacts that they feed big story tips to), they just quietly dropped a blog post and a link to the repo on a tuesday evening.

      I get we really want to have some kind of negative framing here, but lets at least not lose track of the ground truth around it. In a parallel universe this was easily headline news that made it to water cooler chat the next day.

    • surgical_fire 1 hour ago
      These are not mistakes.

      Those results were not dumped to advance math. Those results were dumped to generate positive press for OpenAI.

      Now it's on actual mathematicians to figure out whether the proofs are bogus or not. But the mistakes will never reach the same level of public attention as the original positive press, so for OpenAI this is good anyway, consequences be damned.

      In a sense this is a microcosm of AI usage in the wild, ignorance is laundered through LLMs, and it is left for those that still have knowledge to figure out what makes sense.

      OpenAI is a horribly negligent company. The threat it poses to humanity is not that their models will be a superintelligent singularity that will take over the world, is that their negligence and greed has real world consequences that they really don't give a fuck about. Like right now, their shitty math bruteforce is just keeping actual specialists occupied trying to figure out what is bullshit from what is not. For free, mind you.

    • lptome 1 hour ago
      I think it's equal parts PR stunt and negging.

      People won't want to seriously peer review an AI study unless it's already out there potentially spreading misinformation

  • bamboozled 46 minutes ago
    Vibing, the literal definition of it.
  • renyicircle 2 hours ago
  • dlisboa 1 hour ago
    Imagine how many wasted hours in peer review these AI math papers will cause for the 1% chance of reaching a transformative idea.
    • hyperbovine 53 minutes ago
      Less than the countless hours thousands of mathematicians spent over the course of decades unsuccessfully attacking these problems...?
    • unddoch 16 minutes ago
      They literally proved the UGC?
  • gyanchawdhary 1 hour ago
    Tao and most anti/critical ai math folks remind me a bit of the brhamins in the hindu caste hirarchy system... whilst others fought (kshatriyas), farmed/traded (vaishyas), built things, cleaned roads etc (shudras), the brahmins were the high priest doing science, stronomy, religion ...

    AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...

    • idiotsecant 1 hour ago
      This is weirdo pseudo-neo-religious slop.

      This metaphor might be applicable if the AI slop machine was in fact producing novel output. It seems to be getting invalidated as people dig through the wall of meaningless text surrounding the actual results.

  • iamniels 2 hours ago
    "a sign error" LOL
  • Jeeetendra 6 minutes ago
    [dead]
  • falconBrisk47 3 hours ago
    [flagged]
  • rfgplk 2 hours ago
    This is pretty standard. Important to note that these "errors" (* not really errors) themselves were caught by an LLM, further proving their usefulness.

    * The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.

    • mahogany 1 hour ago
      It's so funny to see people on Hacker News talk about academic communities that they clearly are not part of. To be clear to anyone reading, no, retracting papers is _not_ standard in math. In math, people submit papers after they have checked with colleagues, they give seminars, they submit it for peer review, etc. It is not shotgunning papers and retracting incorrect ones on the regular. Retracting papers is rare and embarrassing.
      • vessenes 26 minutes ago
        Wait, what? Well, yes, you're correct, AND flipping a sign in a draft paper that your buddy catches in a pre-read is SUPER common. People don't do math in a vacuum. Well, many do not. Some do.
      • rfgplk 32 minutes ago
        First of all, people retract papers all the time. I can't give you the exact figures, but surely at least half of all academic researchers have retracted a paper at least once in their lives... Second, is it so hard to believe we are going through a complete paradigm shift in academia? People can now produce hundreds of high quality papers in a matter of weeks/days (with enough compute). Even if the error rate goes up (which it likely won't) we're still being exposed to a monstrous amount of novel information.

        Even Einstein retracted one of his earlier papers re the cosmological constant... And he is one of the greats. Although not purely math related, it still counts.

      • empath75 59 minutes ago
        It is wild to me how much stock mathematicians put into the social aspects of math. The whole social framework you are describing _did not exist_ for almost the entire history of math. It might very well be seen as a historical anomaly, like renaissance mathematicians challenging each other to solve cubic equations.
    • mtlmtlmtlmtl 1 hour ago
      Would you care to share your research? What was your methodology? Did you write it up? Is it published?

      Or is it the kind of research you'd prefer to keep shrouded in mystery?

      • rfgplk 34 minutes ago
        Thanks for the question. It's difficult to quantify but it basically comes down to how many obvious errors LLMs produce amortized by the processing time.
  • breezybottom 3 hours ago
    So much for the "it's lean verified" defense.
    • n2d4 2 hours ago
      The withdrawn papers were not lean verified nor claimed to be.
      • hckrme 2 hours ago
        I wonder why the heck were they provided / uploaded then. Perhaps just a fast and loose play-out on their part. What I don't understand is how come engineers / scientists working on these are okay with this kind of attitude.
        • amelius 2 hours ago
          I suspect they are not okay with this way of working.
          • mcmcmc 2 hours ago
            I’m sure the outsized pay packages help quite a bit
          • ModernMech 2 hours ago
            They’re okay enough to do the work and collect a paycheck and stock options. I don’t think their arms are being twisted that hard.
        • empath75 57 minutes ago
          Because they have hundreds of potential proofs they were sitting on, and lack the internal expertise to judge them. It seems like they're working to formalize all of the in Lean, but that takes time.
          • n_e 19 minutes ago
            This sounds suspiciously like non-developers dumping their vibe-coded projects on online communities
        • boxed 2 hours ago
          How come mathematicians are ok with human mathematicians are ok with that kind of sketchy submissions? That's pretty bad too.
          • dwattttt 2 hours ago
            ... they're not? Why would you think mathematicians are happy about sketchy or bad submissions.
        • xjdirkdn 2 hours ago
          Lies travel around the world before the truth has time to put it’s sneaker’s on…

          The headlines keep the hype train arunnin

      • breezybottom 2 hours ago
        But that's the argument that was used when people here were skeptical about the results.
        • yreg 2 hours ago
          not these results
    • nkmnz 2 hours ago
      Can you point to where that defense has been made?