@PokerPirate's banner p

PokerPirate


				

				

				
0 followers   follows 1 user  
joined 2022 October 06 22:32:38 UTC
Verified Email

				

User ID: 1504

PokerPirate


				
				
				

				
0 followers   follows 1 user   joined 2022 October 06 22:32:38 UTC

					

No bio...


					

User ID: 1504

Verified Email

You will find more greek letters than English ones in a COLT open problem, and never anything empirical. Here's a representative example I just googled: http://proceedings.mlr.press/v125/van-erven20a/van-erven20a.pdf


When googling for the above, I also stumbled on a 3mo old repo that has some AI resolutions to some COLT open problems: https://github.com/Pengbinghui/pipeline-math. The problems they resolve are:

  • Shuffled SGD — the SS–RS–GD inequalities (Yun, Sra, Jadbabaie, COLT 2021)
  • Learning measured-output quantum circuits (Kun and Reyzin, COLT 2015) (Partial solution).
  • Unweighted data selection for linear regression (Hanneke, Moran, Shlimovich, Yehudayoff, COLT 2025) (Problem 3)
  • Robust conditional probability estimation (Langford, COLT 2010)
  • Fixed-parameter tractability of zonotope problems (Froese, Grillo, Hertrich, Skutella, COLT 2025)

Of these, I've personally spent some time thinking about shuffled SGD and robust conditional probability estimation. The SGD problem is vaguely related to LLM training in a "that's interesting" kind of way, but not something that has any real world impact. The robust estimation result I could see being actually implemented in the backend of FAANG companies to improve performance of various algorithmic feeds and help them make a few more million/year, but no impact on LLM training.

I teach basic NP hardness results in intro CS classes now specifically so that students can make arguments like this. I feel like everyone these days should know that AI will never solve a traveling salesman problem[*], even approximately! Of course, neither will any human :(

[*] the usual caveats of P!=NP and general graphs (not euclidean/metric) apply.

It's frustrating how no one impressed by these accomplishments seems to have made any effort whatsoever to explain why any of this should matter to a layman.

It's been <24 hours. If you want discussion from older results, see any of the famous math youtubers like 3blue1brown for layman-accessible intros: https://youtube.com/watch?v=TfyPshgMbug.

ML also doesn't have a meaningful bank of rigorous conjectures

Sort of true. (I am a ML professor.)

COLT is the main conference for ML theory, and to ML researchers is considered more prestigious than ICML/NeurIPS/etc. It has a much lower impact factor, however, because it is truly a math venue about proving statistical theorems and does not accept any applied/experimental work. It has always had a track for papers presenting open problems. You can find the latest year's CFP at: https://learningtheory.org/colt2026/openproblems.html#cfp

I expect most of these problems to be "much easier" than standard math open problems and that these AI systems could prove most of them. The reason is that they don't receive the same amount of attention as "traditional" math problems. The number of researchers who can meaningfully even understand any one of these problems averages <100. That's similar to many math problems, but the difference here is that the ML researchers do not actually spend time thinking about the ML open problems because they are spending most of their time thinking about realworld ML applications and chasing $$$. Traditional mathematicians don't have either of these distractions.

There are a handful of classes of open problems where a resolution could meaningfully improve user experience with LLMs somehow. For example, there are open problems about improving the sample efficiency of reinforcement learning, automatic hyperparameter search, and A/B testing; and all of these are subproblems that the major labs have to implement to train their models. The actual math is abstract enough, however, that I don't see a lab keeping a result as a trade secret if they do resolve any of these problems.

dammit, yes

Story time.

I have 4 small kids and a SAH wife. Even with this modest family, the government has made basic errors with our taxes every year the youngest has been alive (3 years in a row now). For example:

  1. I live in California, and the California tax form 540 only has space for 3 children. For larger families, the official instructions are to write a letter explaining the situation to get the extra tax credits. Apparently this is so infrequent that there isn't even a form for it! Two out of the 3 years, California has lost this letter and not applied the tax credit for the 4th child until I do dozens of back-and-forths with the tax board.

  2. This year, the IRS had a mistake reading one of the SSNs of the children, and so they didn't count any of the children as deductible. (Apparently the OCR system they use didn't like the font of my pdf writer, but it's the same font as before... and they should already know all the SSNs...) We've been doing back-and-forths with the IRS since April and they still haven't resolved this for us.

So I don't trust the government at all to properly handle any sort of tax breaks/etc for large families.