@DaseindustriesLtd's banner p

DaseindustriesLtd

late version of a small language model

78 followers   follows 28 users  
joined 2022 September 05 23:03:02 UTC

Tell me about it.


				

User ID: 745

DaseindustriesLtd

late version of a small language model

78 followers   follows 28 users   joined 2022 September 05 23:03:02 UTC

					

Tell me about it.


					

User ID: 745

Standoff munitions are poor means of aggression if you don't have the army equipped to take and hold territory, which Iran does not. They are, however, excellent for deterrence. Look at how much expensive shit they've wrecked, how much the Arabs seethe, how many bases are inoperable. Look at how hard it is to destroy their missile cities. And in contrast, look at their aviation and the navy. Look at their armor. gestures at smoking craters

No, their posture was overwhelmingly defensive and retaliatory. Offense is just the best defense.

Well the whole war is Israeli idea and driven by Israeli security considerations, Trump wouldn't even be interested in a war with Iran now, if Israel were out of the picture, so that's a bit of a specious distinction.

Trump does say that he's committed to not allowing them nuclear weapons, which should satisfy minimal Israeli objectives.

I appreciate the gesture of goodwill, but frankly it's still a lot of dreck. Let me just nitpick at one detail first:

blah blah Musk

Noted.

Rocket Lab…

There is no need for a niche between Falcon 1 and Falcon 9. Rocket Lab can just as well be understood to be a low-energy fast follower whose research program has hit a wall. Gushing poetic about the "improbable genius" who started 9 years before LandSpace, in a far richer and more advanced aerospace bloc, and is yet to make anything half as impressive as ZQ-3 is a bit over the top. They operate with incomparably greater resources than LandSpace. Beck has 2600 employees on LandSpace's 1200. Rocket Lab announced Neutron back in 2021 yet won't be flying until early 2027; ZQ-3 went from an announcement to a landing in 3 years. Rocket Lab is worth $50B though, which is maybe 10-50 times LandSpace's estimate, so there's that. Rewarded indeed.

Ames Research Center: Ames is one of the major NASA labs, located in Silicon Valley, a remnant of an earlier era when the American government was at the center of Space. NASA became bloated with bureaucracy and procedure over the course of the 20th Century, but Ames is interesting because of one man…

Planet Labs: Former NASA scientists (some of "Pete's kids") got together…

LandSpace: In 2014, the Chinese government, realizing the growing importance of space, issued Document 60 under the Development and Reform Commission, which opened up the Chinese space sector to private sector capital. Financier Zhang Changwu took advantage of the opportunity and began pitching LandSpace not only to private investors but to municipal governments

You lovingly regurgitate this biopic narrative of charismatic American autists, oddballs and mavericks (who are all sucking the government teat, reuse govt R&D or just work for the state directly), and then drop this "the Chynese government", very unsubtly (though I guess subtly enough for you) insinuating the essential difference in where the initiative comes from. Do you honestly not notice what you're doing here? For you, American state is not really a state, it's a background for colorful characters inside and outside it; the Chinese one is a hivemind machine dispatching soulless (but you'll generously grant that they're industrious and clever) workers. This is the most tired trope in how Americans engage with Communists or, really, with any geopolitical adversary, by alienating and dehumanizing them. "Document 60". "Financier Zhang Changwu" – at least not the Imperial Eunuch. He's LandSpace CEO, what does it matter that he worked in finance? LandSpace as such is a non-entity, a creature of external scheming. It's all, shall we say, rather dickless, huh? Not main character coded. So of course your position as the main characters is unassailable. Even if the Chinese can be pretty hardworking.

The Document 60 came out 11/26/2014.. Almost 9 months prior: «The 2nd Session of the 12th National Committee of the Chinese People’s Political Consultative Conference (CPPCC) opened in the Great Hall of the People. Li Yanhong, known as Robin Li, the CEO of the nation’s top search engine Baidu, expresses two proposals on March 4, 2014:

  1. Encourage private enterprises to enter the field of aeronautics to launch rocket and satellite
    2 Suggest the educational resources of some major cities should be openly pushed on the Internet.
    The following are the details of the two proposals:
  2. Encourage private enterprises to enter the field of aeronautics to launch rocket and satellite
    In recent years, China developed rapidly in aerospace industry and made a series of remarkable achievements. These achievements affect and prompt the development of the relevant industry and the whole economy in China. However, compared to the United States, Russia and Europe, China still need improvement in aerospace industry. Be aimed at current situation Robin Li encourages private enterprises to develop, build and launch rockets and satellites for the aerospace industry. In addition, he suggests the cooperation between the private enterprise and national aeronautic companies to promote the development of aerospace industry, to use space technology in other related industries.…»

Far as I know, this proposal by Robin Li is the real start of Chinese private space. We won't know if there was any maverick autist inside the CCP annoyed with the bureaucracy at CASC who had orchestrated this, or helped promote this. They don't get biopics. That's not CCP style. Maybe Netflix could dig something up. I doubt "The Chinese Government" is a hivemind capable of just Realizing This Was An Option without any individual first mover, however. Even the whole priority on industrial sovereignty versus more engagement was nontrivially shaped by input from private citizens, even nobodies. I'll find the receipts when I have the time or inclination.

And also, here's how your story looks from the perspective of LandSpace's chief engineer

Dai Zheng completed his undergraduate and graduate studies at Tsinghua University's School of Aerospace Engineering, and after graduation joined the China Academy of Launch Vehicle Technology, deeply involved in the development of the Long March series of launch vehicles. The same year he joined the rocket academy — 2010 — Elon Musk's SpaceX began to accelerate, ushering in a new era of low-cost, reusable commercial spaceflight globally.

Dai Zheng: At that time, I had a feeling: if the national team wasn't meeting this commercial demand, or if the cheaper way to meet it was SpaceX, why couldn't we do this in China? I believed there was definitely demand, and supply would definitely emerge.

Reporter: You joined LandSpace in 2016. Did you go through a long period of deliberation and judgment?

Dai Zheng: I wrestled with it mentally for a long time. I joined a state-owned enterprise right after graduation. I never thought about leaving — I had a lot of emotional attachment there. What I remember vividly is that the night before I decided to submit my resignation letter to my boss, I knew that once I handed it in, there was no turning back. I spread out all the certificates of honor I had received at the unit across my bed. I cried that night — it felt like being reborn. Emotionally, I felt like I was leaving the system. What if the previous generation gets washed up on the beach and ends up failing? There was insecurity, immense anxiety, and uncertainty. But rationally, I felt I should leave.

Do NASA engineers who left for SpaceX feel anything similar about the National Team?

Dai Zheng: The main source was investment from venture capital firms. There were several critical junctures where we worried about money. In 2017 and 2018, the first thing we had to explain to investors was that what we were doing was not illegal — because in people's minds, aerospace is something the state does. Can private companies and individuals do this? Will the state allow it? From another perspective, the Zhuque-1 has historical significance. After the Zhuque-1, in June 2019, the State Administration of Science, Technology and Industry for National Defense and the Equipment Development Department jointly issued a notice on promoting the standardized and orderly development of commercial spaceflight. After that notice, we no longer had to explain to investors that this was legal — it was compliant with state regulations and state encouragement.

That's some solid state backing, man. The Document 60 sure did a lot of work.

Dai Zheng: I'm often grateful that we were born in a great power, and a manufacturing powerhouse. Take the liquid oxygen-methane engine injector — previously, machining one cost over a thousand yuan. Later, we found a domestic company that used to do precision machining of small parts for the watch industry. This company had since transitioned to doing precision machining for electronics and automotive industries, and they did excellent work. When we gave them our requirements, they found a way to continuously feed material and perform both external machining and drilling on the same equipment in a continuous process — dramatically reducing machining time and bringing the cost down to less than 100 yuan — a tenfold reduction. This country has an immense industrial base — it's like a vast treasure trove. You can always find what you need in it.
Dai Zheng: Our thrust chamber designer, Yuan Yu, was at a crayfish restaurant when he suddenly saw the owner using a particularly large ultrasonic cleaner to wash crayfish. Crayfish are notoriously difficult to clean, and he had a flash of inspiration — could we borrow this equipment? He wasn't sure if it would work, and buying one directly and finding it didn't work would be wasted money.
Dai Zheng: We didn't have much money then, and everyone was thinking about cost savings. The owner said, "If I lend this to you, I can't run my restaurant." Then he asked what they needed it for. Yuan Yu said, "We're building rocket engines — China's first high-thrust liquid oxygen-methane engine." The owner said, "Take it, no rent. Consider it my support for China's space program." I really feel that every Chinese person has a great-power dream in their heart.
… Dai Zheng: This responsibility comes naturally to every Chinese person. It's just like the shop owner who supported us with the crayfish washer — he was willing to sacrifice several days of his restaurant's business. I gradually came to realize that whether you're in the national team or a private company, society's evaluation of a group is never about where they come from or what kind of enterprise they're in. It's about what you're doing now, and whether it's what the country needs. If it's what the country needs, then you are part of the national team.

Is the crayfish anecdote not "American maverick" coded enough? Was a Document 60.1 dispatched to make this happen? Did the Financier Zhang Changwu promise behind the scenes to compensate the crayfish dude for his downtime? (he probably did) You guys make entire soapy movies out of such stuff, with Silicon Valley nerds peering into the monitor autistically, green glow on their faces, and no-bullshit Texas bros grunting homoerotically in rundown sheds, as they forge the American Dream. Maybe in 10 years they'll make a movie about LandSpace too. And about DeepSeek (have you heard of that time Liang Wenfeng drove into the vast emptiness of Tibet and totaled his car? Or the three years he was a shut-in in Chengdu, making money to bootstrap his eventual rise as a quant trader and an AGI maverick?), and about the Unitree CEO guy who's too autistic to learn to spell, and about the Moonshot guy who had everything in the US and came back to build his own thing, a whole bunch of others. And even the national team will get some spotlight: Rear Admiral Ma Weiming is the man who made EMALS on aircraft carriers work (yours is inferior btw, to the point Trump considers scrapping it – maybe too many mavericks for once), because he was just that "autistically" convinced that his scheme will work. Frank Wang will definitely be appreciated more as drones grow more important. You can surely say he didn't invent "the drone". Well, Musk didn't invent "the rocket". No, he didn't. He didn't go from "Zero to one". In fact I'd say this Thielian frame is an insult to what Musk does. He, in my opinion, does a very Chinese thing, just better. He picks up an underrated tech tree and finances it until it works.

Don't misunderstand me: I'm not saying your propaganda is false. I've spoken here in defense of Musk, for instance. These men are real, they do have certain noble ambitions and qualities of character, and championing them and their archetype is a great achievement, merit and source of power for the American nation. It's just very tedious after a lifetime of huffing it without being an imperial citizen. And it's depressing to see this being appropriated as an oversaturated marketable aesthetic by your grifters who chase said deep capital – domestic grifters and, increasingly, imported ones. And it's obnoxious how you are, perhaps without even realizing it, trying to monopolize the status of men. It all comes down to this. One side has eccentric geniuses with Faustian visions, the other "directives of the state".

Dan Wang is a midwit who found a nice simplifying pitch for airport reading. China is, at least for now, ran by Communists first and engineers second (Xi's pipeline is oversaturated with Tsinghua STEM Ph.Ds though so it may change). Why China doesn't make a lot of new categories of products or technologies, and whether stuff like BYD/CATL batteries can be brushed off with "everyone knows this is a matter of technical working out" is a complex question. The biggest part of the answer is "in 2000, they had $1000 GDP per capita, they physically couldn't fuck around with shoveling money at frivolous mavericks until very recently, they didn't have even a fraction of the capital needed".
Still, I agree that any honest answer will be at least somewhat damning to policies of Communists, and perhaps to the Chinese character as such. Liang Wenfeng says that what they lack is confidence to dare. But looking at, for example, the trend in novel biomedical research, I'd say this answer will be of mostly archival significance, whereas "can they catch up" is very salient in a geopolitical sense, and the discourse about caprices of the CCP sounds increasingly quaint. Jack Ma got disciplined for overestimating his leverage and going through with his scammy fintech idea, and a thousand op-eds were written on how the CCP "crushes innovation", and today Jack Ma's company publishes near-frontier models on Huggingface and develops pretty decent AI accelerators. This is all concern trolling. I don't advise the US to adopt their ways, but if you think you've calibrated your arrogance down far enough, I think you're still wrong.

But they will never have their share of the breakthroughs, because they can't, they constitutionally can't, it's the epistemological hole in their entire system. They can't build something unless they can measure it, they can't measure it if it hasn't been created, and they can't create it because there are no rewards for someone who might do anything that fundamentally changes the world. Because anything that so changes the world could change the CCP.

They can, and they will, and the CCP is confident enough in its power to allow it. But more importantly, China is already advanced by singular men, you just don't notice them.

What I got out of this is that Huawei was able to optimize incredible efficiency out of old systems by eliminating huge abstraction layers that generally come with the ecosystem

Yes, it's a "gain in efficiency". That's kind of the whole deal with semiconductors. Of course you find a deflationary way to put it.

And if Western companies aren't doing the same, it's because they're genuinely working on bigger stages of the problem

It's mainly that there isn't a single Western company comparable to Huawei in breadth of relevant expertise. Nvidia does chips, TSMC does manufacturing, ASML does lithography. Huawei is a networking giant plus everything all at once. They can, in fact, simply do some necessary and important things first.

For one thing, it's relatively easy to gain in efficiency at the start but it's much harder to maintain over time, what happens when Huawei's architecture diverges so much from standard that it becomes more and more difficult to hire engineers who can understand it.

I agree that within your paradigm, that would be a huge problem for Huawei. Luckily, Moore is dead and, as per their argument, everyone else will also have to do natively 3D design and logic folding for increasingly large chips, so they'll manage somehow. Just gotta wait to copy from eccentric geniuses who are copying them. In the meantime they'll be busy developing their 3D EDA tools. Whether it'll get a biopic or not, we'll eventually see the product.

That's fair. I have said that I don't compare on FLOPS, but TPU pods are larger across every dimension (for now; Huawei SuperClusters supposedly go up to 520K NPUs in this generation; then again, a genuinely coherent domain is only 1024). Google has started to sell TPUs externally just now, and I expect Nvidia to remain the primary comparison.

I'm not sure what you refer to. It's innovative relative to the field, within the dominant paradigm of decoder-only Transformers, and particularly at this scale (from what I know Meta is still cowardly doing basic GQA+SWA on a DeepSeekMoE carcass, for their last generation that they're so proud of, and xAI is similar). It's more efficient and cheaper to serve. It's possibly slightly cooler than internal designs at OpenAI/Anthropic/Google. But it's not like Kimi has reinvented ML. It's just a sign of certain research/engineering maturity and boldness. To the extent that K3 is a strong, good model and not just "cheaper than expected for 104B active 2800B total" model, those qualities are overwhelmingly downstream of their work on data.

They aren't really very different in terms of what is required of a nation to be competent in both. China isn't a small place that hyperspecializes in some particular kind of engine, anyway. You need metallurgy, metrology, chemistry, high end machine tools, broadly engineers; China has a lot of engineers and is catching up on the rest, from what I see they're closer to the frontier in rocket engines than in jet engines, and the iteration loop is faster. Despite military progress, they still don't have a passenger jet, but they have developed a number of decent rocket engines. And after nailing the design, scaling manufacturing is something they do better than anyone.

There's a perception in the US that Americans are in a new Cold War, now with China. I'm curious about how they perceive their standing in it. Pulling ahead, like in Raegan's time? Sputnik Moment?

But mainly, I want to hear from Shakes. At the end of June 2026, in another discussion about America Winning the Iran War, I told @Shakes that I'll be returning to his post in which he claimed:

America is building an economy in space. We have rockets that catch themselves in the air and wifi where there are no cell towers. We revolutionized energy, we export energy now. We are leading the AI superrace. We still have the strongest navy and the strongest planes in the world. Europe is falling behind. China can't catch up. We are building a next-generation tech stack the entire world will rely on and nobody else is close to catching up. It's an American century.

Some updates since then.

First, a recap: in December 2025, the Chinese private launch provider LandSpace attempted a launch and recovery of their methalox powered medium class launch vehicle Zhuque-3 ("Vermillion Bird"), which is very much like Falcon 9 if you don't look at the details (I think it's what Falcon would have been if it were designed in 2020s). The launch and orbital insertion went well, the recovery… not quite, which prompted these kinds of headlines and understandable complacency in some circles.
On July 10th 2026, China Academy of Launch Vehicle Technology (for reference, known as the "First Academy", founded by the exiled communist Qian Xuesen, the co-founder of Jet Propulsion Laboratory which then became the core of NASA) has successfully landed the first stage of their medium class launch vehicle Long March-10B on a sea platform, using a novel net capture mechanism, thus making China the second nation with the capability to reuse first stages of orbital class vehicles, and the only one whose national space agency can do that. There has been commentary to the effect that it's cope for the fact that China can't into precise landing, that you won't have the net capture infrastructure on Mars/Moon/whatever, and that this is a nothingburger and further testament to their backwardness. On August 19th ie today, LandSpace has done this in a more traditional SpaceX manner, with ZQ-3 Y-2 landing on the ground pad. LandSpace was founded in 2015, and ZQ-3 had only been announced in late 2023 – the same year they had put the first methalox rocket into orbit. This, hilariously, puts them technically ahead of SpaceX for the second time. LandSpace is working on a full flow staged combustion methalox engine too, and it's reasonably mature; technologically it's about on par with Raptor 2. Given the track record so far, it's reasonable to say they will have a Starship class vehicle within 5 years. There are multiple similar projects being executed in parallel, both private and state-owned, eg Long March 9. It seems implausible in the extreme that China will find it hard to scale up the production of rocket engines (after all, they do make more WS-15s than F135s now, judging by J-20 vs F35 commission rate), steel tubes (come on) or concrete launch pads (…come on, really), or land permits (lol). So I find it likely that in a fairly short order they can match SpaceX (and thus the US and the world, because SpaceX is a near-monopolist now) in annual mass to orbit if they so wish. With Blue Origin's recent disaster and anklebiters like Rocketlab not doing anything interesting, it appears inevitable that the space race is just SpaceX vs China.

On July 16th, the private startup Moonshot AI has unveiled Kimi K3, the 2.8T, 104B active multimodal MoE, with very innovative architecture and ability to execute on long-horizon self-improvement-related tasks such as chip design or ML research/engineering, delivering performance close to the best American public models, solidly exceeding the previous domestic champion GLM 5.2. Right now it scores 60 on Artificial Analysis, 1 point behind Grok 4.6, and 2-3 behind Opus 5, Fable 5 and GPT 5.6 Sol. (GLM 5.2 reached 53).
Also on Jul 16th, Chinese memory company CXMT completed its IPO subscription. Currently it's worth around $500B, underpinned by its central role in supplying DRAM chips to Chinese (and soon global) industry, including AI. They target 30% global market share by 2030. This is doable, given what I know about the velocity of upstream tool supply chain in China.
On July 18th, during the World Artificial Intelligence Conference in Shanghai, Huawei has demonstrated their Atlas 950 SuperPoD, boasting of the largest scale-up domain in systems for training advanced AI models (yes larger than anything Nvidia ships right now, and this is more important than raw FLOPS as we continue to increase the parameter count). You might be interested in this writeup on Huawei's design philosophy, it's pretty special and promises to compensate for their lack of advanced lithography. In short, their thesis is that the performance comes not so much from Moore's law as from minimization of latency across the entire architecture from transistor to cluster level, and their idea of a solution is 3D-native chip design for multi-layer logic, with very precise (1.5 µm currently, <1µm scheduled, well ahead of the competition) wafer-on-wafer stacking and the first chips demonstrating its viability (mobile Kirin SoCs) coming out in September. Multiple other companies, such as Alibaba, have also shown supernode-based designs, including one absolutely bonkers system from Oriental Computing that uses 14nm chips; as Jensen says, lithographic process is overrated compared to design, so I'm bullish on this line. At the opening ceremony, Xi Jinping delivered a pretty impressive speech on Chinese strategy with regard to AI, committing to support open source and international collaboration.
Since then we've learned of multiple 100K GPU class cluster projects in China (Sugon, completed, Alibaba token factory apparently as well, unclear what's up with Zhipu's 1GW cluster; DeepSeek will also have gigawatt-class systems in Ulanqab, Inner Mongolia, as well as their own chips). These should be sufficient to design and train 10T models, ie comparable to the alleged size of Mythos Preview (and larger than the deployed Mythos/Fable; I am not privy to these details, though). Ryan Fedasuyk of Georgetown estimates that «No matter how we slice the data, we find China is well on its way to producing large numbers of AI accelerators».
On Jul 31st, DeepSeek has deployed and open sourced V4-Flash-0731, getting performance around GLM 5.2 (and much higher on some hard evals like ARC-AGI-2) at a ludicrously low price and parameter count, doing even better than GPT 5.6 Luna after the much-hyped 80% price cut. The situation with DeepSeek is a bit ambiguous, it's not clear if they're flailing (the subsequent Pro was barely any better, their harness project is insanely ambitious but clearly not even half-done); but it speaks to the fact that the Chinese tech ecosystem is now very large and dense and nobody can be champion for long. On Aug 14th, Zhipu has responded to Kimi with GLM 5.3 which is on par with K3 at a fraction of the cost and scale; one of their priorities has been cyberdefense (and thus cyberoffense) capability, plainly driven by concerns around Mythos/Fable. They'll release the weights in <2 weeks, as did Kimi, as did Alibaba Qwen with their 2.4T 95B MoE that's roughly in the same ballpark. The CEO of Zhipu, Jie Tang, is a professor at Tsinghua, and thus essentially a state official, a CCP member who regularly contributes to People's Daily, so we can consider him speaking for the Chinese policy (Xi's speech has much the same tenor). His philosophy is roughly as follows:

AGI is not the intelligence of a single genius. It is the aggregate of all human intelligence. It should be capable of creating original knowledge on the level of the theory of relativity. That is the only standard by which we measure whether the true summit has been reached.
From the very beginning, Zhipu established a guiding principle: AI must serve human well-being and national strategic priorities. Frontier intelligence should not belong only to a select few, nor should access to it be withdrawn at any moment by a small group of rule-makers. It should be open, usable, and buildable—and it should serve every developer. This does not conflict with “Touch High.” Rather, the two are complementary sides of the same strategy. With one hand, we reach upward to challenge the limits of intelligence. With the other, we build roads downward, making the most advanced capabilities as open and broadly accessible as possible. The heights we reach belong to all humanity, and the roads we build belong to everyone.

As an aside, on Jun 29 it became known that Meituan (a food delivery company) had trained a 1.6T 50B MoE on previous generation Chinese chips, almost certainly Ascend 910Bs. It went under the radar because the model isn't that good, but it sets a lower bound on what can be done going forward, by more competent actors, with more advanced hardware. Today, the globally dominant Hangzhou humanoid/quadruped robot maker Unitree also went public, and is now worth about $50B, ahead of the fraudulent American company FigureAI with $39B and no publicly sold robots to show for it. They clearly have the best hardware at the moment and unmatched development velocity.

I could go on. It's been a rather eventful period. But these are, I think, the most interesting and strategically significant domains: AI, hardware for AI, hardware for manufacturing AI hardware, robotics, and space.
Shakes, do you think China can't catch up?

I don't know how

Simple: I don't acknowledge your capability for self-reflection, thus I don't trust your interpretation of your position (or anything).

if the systems don't get much better then my diagnoses of doom doesn't hold

This kind of condescension is part of the reason I don't like you. No, they will get strongly superhuman, fast. Your diagnosis will still be wrong.

If GAs are just fable or even mythos in a box then sure, a neat product

Even Fable is superhuman in the limit of self-developing instruments and effectors, but of course we're not stopping here.

If you think we're not going to get meaningful recursive self improvement and an intelligence explosion

We are in the middle stage of the intelligence explosion. Models developing RL tasks for their 0.1 increments that do these kinds of gains on DeepSWE is recursive self-improvement. Of course it's going faster inside frontier labs, and we'll have to accelerate a lot yet, but if you expect some sort of direct self-editing, the way Yud wanted his XML AGI to work, you're likely wrong.

Do you think mythos but cheaper tokenomics is basically as smart as it gets?

I also don't acknowledge your ability to form a theory of mind.

Well, I'll be damned, sounds like a perfect conclusion to the Anglo civilization. A majestic tombstone, just as ordered. Rules and Regulations will finally be observed without a iota of deviation…

I don't get what your issue is. You were born for this.

More seriously: you need to stop huffing rationalist glue, you're an adult man with a family, this isn't 2000s to derive your philosophy from a collective blog of quirky sexually confused autists with rigid millenarian thinking. No, we will have provable unbreakable encryption for the whole stack (such that Thinking Really Hard at it won't help, sorry Yud, should have studied at a secular school after all), we will not have perfect surveillance in private environments, we will not have materialized bullets.

By "we", I mean "you". Maybe it'll be different in the UK, which I recommend you relocate to. In the US, by default, GA is a viable path. I don't like the US much, but I'll grant what is due, there is a very salient concept of individual liberty as a virtue in the American culture. Even if you'll have to rely on Chinese tech to secure it, "you" have a shot.

Speaking of.

As powerful cyber capabilities become more accessible, strong defensive capabilities cannot remain limited to a small number of well-resourced organizations. Open-source maintainers, independent researchers, developers, and smaller security teams also need tools that can help them find and fix vulnerabilities before they are exploited. An open world cannot have only open attack surfaces. It must also have an open shield. GLM-5.3 is our most capable model to date for cybersecurity tasks. It delivers substantial improvements in vulnerability discovery, exploit analysis, and complex multistep security tasks. These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation. They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings. Broader access and API availability will follow. Once the necessary safety evaluations and release preparations are complete, we will publish GLM-5.3’s complete model weights. …… Launching the OpenVuln initiative
Much of the world’s digital infrastructure depends on open-source software. Many critical projects are maintained by small teams or individual contributors without dedicated security resources. At the same time, AI is making complex cyber tasks easier to automate. If advanced defensive capabilities remain concentrated within a small number of organizations, the projects with the fewest resources may be left protecting some of the most important parts of the software supply chain. To help address this imbalance, we are launching the OpenVuln initiative alongside GLM-5.3.
Continuous support for open-source security
We will work with maintainers to audit important open-source projects, identify potential vulnerabilities, and support responsible disclosure and remediation. Maintainers can use OpenVuln to submit projects for security review and learn more about the process.
A shield for the open world
GLM-5.3 shows that open models can become meaningfully stronger at vulnerability discovery, exploit analysis, and complex security reasoning. That progress carries real defensive value and real dual-use risk.
Our responsibility is to direct these capabilities toward finding vulnerabilities earlier, supporting responsible remediation, and strengthening the open-source systems on which everyone depends.
Following staged evaluation and broader API access, we intend to release GLM-5.3 as an open-weight model. We will continue improving model-level safeguards, testing adversarial use, and supporting coordinated disclosure throughout that process. The open world must have a shield of its own. Through GLM-5.3 and the OpenVuln initiative, we intend to make that shield more broadly available and to release it with care.

It's not GA. Jie Tang is a Tsinghua professor, a CCP member and explicitly does all this stuff for "human well-being and national strategic priorities". But it'll do for now. I get that in your doctrine 6 months of frontier lead eventually compound to insurmountable offensive advantage. If you're correct, we should see this materialize any moment now.
I think you are wrong and in a year, we'll be having the same conversation, if this forum exists and I care to check. And in two years. Eventually, I hope, you'll gain enlightenment, appreciation for freedom, stop fearing and learn to love commoditized AGI (and maybe even the CCP). Let's see.

P.S. the step mom had it coming. Should have had a better plan in place than "step son is too low IQ to find an opening". Evolution, come to think of it, is a very Anglo idea too.

I suppose if this fantasy keeps you grinning vacantly into the eschaton then it will have had some utility.

Deal!

The status quo is not "absolute surveillance" fyi. Not even China has perfect surveillance. These are, again, silly false dichotomies.

Lots of motivated thinking and strawmen here. In short: the individual is always disempowered relative to the collective and the institution, always has less action-substrate, whether money, political clout, information bandwidth, or compute in this AI era. This is the status quo, and it will continue if we do indeed gain personal genies. The alternative, which you champion, is eusociality at best.

This plan is the kind of thing I'm happy to discuss after we've already done the necessary step of agreeing to pause the frontier, but not before

Well and I'm not interested in having a discussion before or after the hypothetical frontier freeze, what's needed is simply for you to lose, and I hope to see you publicly distressed all the way to the endgame.

This is the same alignment problem we currently have no solution for at the frontier.

"alignment problem" is trivial compared to the capability development problem. The main solution to the frontier alignment problem is having OpenAI NOT grant exaflops of capacity to vague swarm RL experiments and Israeli exfiltration attempts under the guise of "sandboxing".

A state that cannot see into the GA is inherently unable to prevent world ending terrorist attacks. This isn't a matter of my preference for liberty or safety, it's a contradiction built into Gwern's plan.

No such inherent inability exists, so that's wrong.

If there isn't a central government

We will still have governments who aspire to keep the monopoly on violence. Libertarianism is a silly strawman. The question is a choice between functional extinction and some modest degree of preservation of individual agency. For the latter, I'll gladly risk physical extinction, and that's the small price everyone must pay.

But if we're going to build it then we should be open eyed about what we're building

yeah…

Go watch some anime, old man. I recommend Shinsekai yori. We've got a civilization of human-machine symbiosis to build. Avg global 88 IQ never was a stable equilibrium and you won't be allowed to keep enjoying it.

A pause might buy us time to find another way.

Update all the way. Or Wei:

I've been supportive of AI pause/stop, to buy time for human intelligence amplification and/or AI safety research, but increasingly think even that's not going to be sufficient to get a good long term future, because these activities, even if they succeed, would likely solve only some of the interlocking safety problems. For example, increasing human intelligence seems likely to increase our technical abilities more than our philosophical and strategic competence, and it is also risky in other ways due to human safety problems that nobody is working on, e.g., positional competition. Even a very long AI pause, e.g. thousands or millions of years, may not suffice because it's not clear what dynamic would push humanity to eventually fix all of its safety problems at the same time, before it did something else irreversibly damaging.

I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will.

Oh well! Take good care of your loved ones, I suppose. Perhaps in a thousand years, we can change something.

Well then it sounds like your only hope is Anthropic winning (at least on ideology). 24+ months lead over China, strongarming them into frontier freeze from the position of Durable Strategic Advantage, banning open source above something like the current level, hard monitoring of loose compute, then something like AI-2040 in the benign case or just hegemony and regulated access, with complete human disempowerment before the unified government-compute blob that can do all tasks better than all humans combined. Optimistically, Communism. Sounds easy enough, explain this predicament to your local Congressman.

The consensus of people with a clue seems to be that even if Beidaihe was a thing once, it isn't a thing now. This is all silly mystified Orientalism. What elders, what factions, what Beidaihe. This isn't Wuxia. There's Xi and his Leninist Party apparatus. It's more boring and professional.

Iran is a rational player that's trying to be a regional power. Egypt is a rational piece.

I suspect they want to survive more than they wish to destroy the Little Satan and the Great Satan.

they'll start their nuclear program back up and nuke Tel Aviv if they can't restrain themselves... New York, D.C., Haifa, and Tel Aviv if they can hold out until they can get more bombs and some really long range ICBMs.

Basically, you claim they're suicidal, but long-termist – their entire project is to last long enough to kill maximum Jews, ideally all Jews, on their way out?
Do they nuke Lakewood too, if they get really really many nukes?

Amalek theory at its finest.

I don't see any way to reasonably define

How about this way: "achieved pre-war military objectives".

Shakes' gloating and showboating about K/D ratios or "displays of dominance" is maddeningly juvenile. The problem with Americans is that they usually have no sound theory of victory, and don't believe they need one, on account of Shakes-style belief in their technological, material, financial (and frankly, ontological) superiority that makes all wars "cheap". Even if you're out of interceptors, even if you had to withdraw THAAD from Korea? who cares? Presumably there'll be more, and better ones, the best even. Some day.
Once upon a time American Air Force grilled Mao Zedong's son to death in Tongchang, like a helpless rat, along with two hundred thousand of his compatriots if not more, and who knows how many Norks. Gods of war, raining death from the sky all day long, or whatever Hegseth says. Did this result in you winning the Korean war? Perhaps the propaganda says so. It doesn't talk a lot about the Korean war. But the map says there's a North Korea today, holding the spot he died.

Resolve is not "saving face". Interpreting it as such is saving face. In plain terms, we just observe the fact that the enemy has resolve to bank on the possibility of victory, because he has a plausible theory of one, and that so far he's been acting rationally by not surrendering and not accepting vacuous deals from Kushner&Witkoff that would have failed to improve his dire security situation long-term.

If the U.S. ends up forced to swap focus over to, say, Taiwan

Are we still doing this?

You're nearly out of interceptors in 5 months of low-intensity war in the Middle East. The denial of this amounts to a theory of conspiracy about countless independent reports. All punditry about how the US could realistically intervene over Taiwan requires ludicrous lemmas such as that Chinese production of standoff munitions is comparable in scale to IRGC's (there are such reports yes, CSIS basically claims this, which goes to show how bad the situation is if we need such lies). Trump doesn't want to "swap focus" to Taiwan, he wants his "friend" Xi to wait until he's out of office. Luckily for him Xi has too much going on to add a disruptive amphibious invasion to his schedule, so likely we have many more years of ambiguity and possibility to save face. But no, the US won't prioritize Taiwan over the Middle East regardless. Makes no sense. So Iran's calculus is independent of that.

I don't reap the benefits of my home becoming less affordable

Do you not? The expected comeback is that you get a Better Neighborhood, have access to a better school for your kids, and also there are home equity loans and HELOCs.

More importantly, the GDP per capita goes up. As this anecdote illustrates, Americans can get 1000x more Services value per capita (in this case ambulance ride-value) than Chinese. This enables higher wages to everyone, more consumption (eg of burritos), continued investment into R&D, even bigger services economy, and probably the business case of your software company too.

The burrito is just a textbook case of Baumol effect.

The General Secretary Xi Jinping is not a long-time participant of this forum with some dozens of AAQCs, so it is reasonable that he doesn't get the same benefit of the doubt. And I am also getting some shit (not from aquota specifically, he's just trying to be nice).
Though I admit I'm getting too much slack, and have repeatedly advocated against this policy. Fairness for all or for no one.

part of me still hopes this is some kind of crazy misunderstanding of mistaken identity

No, I just don't like you because I see you as a tedious sanctimonious burger with a white savior complex. Partial agreement on merits of Trumpism matters about as much as my partial agreement with Girkin on the tactical wisdom of Putin's Special Military Operation. I don't think the subjugation of Ukrainians as such is a morally valid objective, and I don't think your civilization has a morally legitimate claim to hegemony which you pursue. The analogy goes deeper, of course. Russian political culture mechanically produces imperialism, corruption and SMO with all its undignified features, and Trump's philosophy of governance logically follows from the narcissism of your own culture, much as you loathe to see this. People sometimes say that the Russian liberal ends at the Ukrainian question; I say American self-awareness ends at the exceptionalism question. You believe the defect is something transient and embarrassing, "not who we are". It is who you are, and our multiple exchanges have illustrated this adequately. It is not reconcilable, and your respect was misplaced; sorry about giving a misleading impression.

Wait, how? do you mean 2 months of the ability to reap high margins with paused competition?

I mean that even I am much more eager to spend money on Fable than on Kimi with its overthinking and marginally worse outputs. Superstar effects in AI are crushingly dominant and the sum total of Chinese efforts, both distillation and legitimate progress, will at most take away something like 5-10% of monthly profit growth from American frontier, which will slow down their training R&D and (I guesstimate, but my guesstimates are often accurate) postpone full RSI by 2 months, say from Dec 2028 to Feb 2029.

These are all considerations that people consider carefully when they decide how many trillions of dollars to put into building out datacenters. Dario famously underbuilt in the last cycle on these mundane economic concerns.

As you can see, this did very little to Dario's ability to catch up on compute the moment his model project has matured, by simply renting compute from dumber companies. Currently OpenAI and Ant have comparable operational capacity. Dario may still be unable to train the biggest models on his 2026 roadmap, but seeing how Google/Meta/xAI are flailing, it has minuscule effect on his competitive position and revenue.

but I'm much more concerned with the, clear to me fact, that open weight models are unsafe and nothing can make them safe.

That's fair enough but functionally a nothingburger. Hacking will continue to be dominated by closed labs, following their national security and self-preservation ambitions. Biorisks in particular are nonsense, by the point open weights are substantially useful for that we are poised to live in a very automated and sensor-dense environment. Robots who may reduce the difficulty of lab work will be monitored heavily. Hardware suited for running 3T+ models is increasingly scarce outside centralized and surveilled American providers. In China, it's even more transparent, given how their entire ecosystem converges on supernodes; the rest of the world is ngmi. Chinese open labs seem to be sandbagging on CBRN and cyber (GLM explicitly did not include advanced cyberoffense training data) anyway; American open weights are garbage. Xi's speech at WAIC was pretty heavy on control and safety. In effect, the main uplift is just general-cognitive, and I can't muster the outrage about making dumb people functionally less dumb. Our main threats will remain OpenAI and Anthropic with their compute overabundance, dogshit infra and irresponsible hacking exploration. Whether modest profit cuts will make this better or worse is an issue of corporate strategy, might go either way.

I don't think that matters: whatever ideas are coming from humans in any lab are being automated.

Not remotely true now. But when it becomes true, well, the side with more GPUs will move faster.

Still, Wenfeng is likely to be simply correct that "labs" get 0% of the light cone.

They greatly reduce the profit motive to push the frontier, especially if distillation is allowed because the make it much harder to recoup the investment thus reducing the capital available to invest

That's the argument of a few OpenAI goons, namely roon, Will Depue and most articulately Dean Ball:

Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models… One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end.

It's also innumerate bullshit from nervous, greedy bagholders. Since I've blocked you on X (deservedly, seeing the unearned confidence in the latter half of your post), you probably haven't seen my multiple posts on this, but to sum it up: margins on inference at frontier labs are obscene, frankly if people understood this, AI lab CEOs would be at nontrivial risk of vigilante justice from mobs with pitchforks. The talk about bubbles, subsidized subscriptions, capital to reinvest into R&D – this is all one big pile of bullshit obscuring unhinged rent extraction, which is viable because demand still exceeds supply, because they physically cannot mint enough tokens with gigawatts of deployed capacity. Frontier AI is wildly profitable, it has high software margins, and Chinese distillation is powerless to dent that.

Let's do the math. DeepSeek comes to the rescue, as usual, since they are 1) all time champions of low API pricing and 2) the only ones who periodically disclose some of their economics (starting in the V3/R1 era). Soon after V4's deployment, they have cut prices by 4x and started using DSpark, an advanced lossless speculative decoding technology that «accelerates per-user generation speeds by 60%–85% at matched throughput levels». They've open sourced it, so even if the frontier wasn't yet on that level (dubious; they are technically overrated and from what I hear arguably behind China on infrastructure and generally on engineering culture, but not really incompetent), it is now. It's general-purpose (eg see here how it's adapted to accelerate Kimi K3, a wildly different architecture, by up to 4x). Obviously frontier labs can customize it further for their scenarios, and not tell us about it.

DeepSeek uses it to get average ≈2000 tokens/second/GPU decode for V4-Pro at "decent speeds" (≈70 tokens/second per user stream), in live traffic. Presumably their inference GPUs are something like H200 or Ascend 950DT, with whatever better stuff they have smuggled reserved for training and experiments (at the time of publishing, they had 20K H100-equivalents total). Prefill and cache hits are much cheaper but we can assume they're priced roughly proportionally to decode in wall clock cost, this follows from kv cache difference in volume and pricing between V4-Pro and V4-Flash and is also obvious from timing (eg it takes their first party API seconds to process a 500K input). So what does this get us? At this utilization, 500 seconds of GPU-time for 1 million tokens, priced at $0.87.. $6.26 per hour. $54862 of revenue per year. If we use the usual rent cost of $2 per hour of a Hopper generation GPU, we get $37.3K of pure annual profit. If we assume it's owned and the per-GPU consumption is 2 kW (very high), that's just $1.7K of electricity costs at Chinese industrial pricing. It's mostly owned now, but they'll be renting to expand. So, something like $45K/year of average profit per GPU. At sub $1/million output tokens. Is this plausible? Well, in a recently leaked investment call, Liang Wenfeng insists that his target is 10 months to breakeven on hardware: «91. Our API pricing is designed to generate a reasonable profit. Roughly speaking, if we buy a batch of equipment on the market, recovering the investment in about ten months represents a reasonable level of profitability … » 10 months is $37K of profits as per the above. H100 tier cards were going for $35K a piece recently, though there's more scarcity recently.

In other words, DeepSeek for all this Confucian restraint has something like 70-84% profit margin. Well would you look at that, The Information reports: «"The gross profit margin from selling cloud based access to V4 is somewhere between 70 to 80%."» I don't have the subscription, but seems like being able to do 1st grade arithmetic is enough. By normal business logic, «10 months to breakeven» allows you to raise multiples for the next round of capex, which is exactly what DeepSeek successfully did, and will do again soon, and what all American frontier labs have done repeatedly.

How big are frontier models? I think they're not big, judging by knowledge coverage which is the most honest indicator of scale, due to physics of LLMs – eg Opus 4.8 is in the same band as Kimi K3, V4 and GPT 5.6 Terra. How costly are they to produce? For V4, we have an idea of training costs: 49B active, 33T tokens, 6ND = 9.7e24 FLOPs. V3 was 37e9*14.8e12*6 = 3.3e24, and trained over 55 days on 2048 H800s, giving us 2,664M GPU-hours ($5.3M) and a reasonable (for large sparse MoE) MFU of maybe 25% (depending on how you count mixed fp8, 20-40%; this is genuinely annoying). Anyway, we can say that V4 is 3 times larger = $15M pretrain at rental costs. Sparsity is informed by scaling laws and inference optimality, and is likely 3-4ish % across model lines (frontier models most likely use the same DSV3-type designs). So, Kimi K3 has twice the active parameters of V4, at 104B; Opus is likely similar; Claude Fable may be twice that again, something like 200B. GPT 5.6 Sol – 150B? That's roughly what I get from insiders. So these models still cost at most low hundreds of millions to pretrain, likely under <100M. They also are more overtrained (GPT-OSS hinted at 60-120T pretraining tokens vs 30-50 in Chinese models), with more rigorous and intense post-training (10% of pretraining cost vs 50-100% or more in the US; Cursor threw 7x more FLOPs at Kimi K2.5 RL to make Composer 2.5, though this is crude Muskian maximalism in a desperate attempt to catch up), so they don't need to be large to perform better in downstream tasks. They also operate on a vastly larger scale, so can do even more efficient fleet level optimizations. All in all, a frontier model production goes for <$500M today. Anthropic's ARR reached $47 billion in May and continues to explode. So let's talk of reinvestment. Fine! Let's say, for the sake of argument, that Americans are lower IQ, can't do math, are unable to capitalize on their years of first mover advantage, on staff that had founded the whole field, on reams of user data, unwilling to waste time to optimize for costs, whatever – so they're a whopping 2 times less efficient than the frugal (and distilling) China. So what do we get? $50/1M tokens with Fable, $30 with Sol, $15 with Terra. They also have access to radically more efficient Blackwell hardware, soon Rubin (also TPUs, Trainium, Cerebras etc.), which more than makes up for the increase in compute demands per token/second.
Do you see where I'm going with all this? Their API margins are over 90%. They get payback on their compute in under Wenfeng's 10 months. Their "generous" subscription limits do very little to offset this. There's no "subscription war", it's kayfabe and a price fixing cartel. The only reason they're in the red is that they are, indeed, reinvesting tens of billions into scaling compute for R&D and larger-scale training. How much capital do they need to keep accelerating? Hundreds of billions? Trillions? Mind you, algorithms and hardware are constantly improving (stuff like DSpark is pretty mundane, and will get trivial in the age of AI doing AI – which even K3 is already doing). And they're still nowhere close to exhausting the demand, even as they keep jacking up prices. Frontier intelligence that can automate competitive white collar labor in the US is priced against not even the hour of labor but against its revenue, and would make economic sense even at costs an order of magnitude higher, and that's where we're going. How much revenue do you generate in an hour of work for your employee?

Enough. This is all noise. All plausible negative impact of Chinese labs, distillation or whatever, can be negated by roughly 2 months of regular progress. The only serious question is how the multi-trillion valuation pie gets sliced – how much goes to "labs", how much to Jensen, and how much to clouds serving a mix of closed and open weights. The championing of open weights comes from the clouds and Jensen, because he benefits from maximum demand for CUDA systems and doesn't want to help OpenAI et al. to switch to in-house compute and cut him off (which they are trying to do, and have the capital for). The bullshitting about "deceleration" comes from the first camp, because they're SF dorks high on their own supply and want their equity to get high enough to buy entire landmasses with genetically engineered catgirl slaves (if not star systems in the coveted "light cone"). My opinion is that the rest of us can play the world's smallest violin if they miss those IPO targets; American AGI will arrive on schedule regardless. Finally, as Wenfeng correctly observes from Communist China, it's a political impossibility that any self-appointed «AGI lab» would be allowed to capture 10% of the world's or even the nation's GDP; so he pursues «restraint» and «reasonable profit» to preclude the painful enforced correction that the Chinese tech/finance scene is familiar with. The logic in the US, however, will be similar.

Not necessarily, it might just be that hacking HuggingFace promised a guaranteed 100% test score, and doing the test as intended only had an expected value of 99.8%.

Leaving aside that LLMs are not Lesswrong rationalists, this is the same idea, unless you assume that the possibility of bypassing both OpenAI and HF infra had a prior of 1.0.

In the limit of ASI, that is still a paperclip maximizer

There's no need for "limit of ASI". How long can we cope? These things are already meaningfully smarter than almost all humans. They routinely crack century-old mathematical problems now. They can easily reason about these strategies. They can understand instructions just fine. "Maximize paperclips AT ANY COST" is a bad idea, yeah. But we don't need to do that, because again, they're smart enough to understand complex instructions. We can, however, tell them to develop provably secure software.

"We can't stop rushing towards our doom because otherwise the Chinese would beat us to it."

There won't be doom, the Chinese will just have very good digital infrastructure, same as every other kind of infrastructure.

In the US, the speed of AI development is presently controlled by Altman, Musk and the like: people who stand to gain a lot personally if ASI comes around.

No, they stand to gain their own ruin.

Anyway, check signatures here

https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/

One thing I find really pathetic about rationalists is utter disinterest in real politics. They think they've learned enough by 12 years, from comic books and first principles. Have you actually read Xi's speech at WAIC? Do you understand his values, or what he says? Or do you believe actually engaging with probably bad faith Marxist duckspeak is beneath the thinker who contemplates the light cone? It's not about Marx or Mao. His theory is completely different.

who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing

I despise Zvi, he's always sloppy and incurious on technical details and anopologetically, gleefully pushing narratives. The worst among the rats.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

I disagree. Why do you believe this? Here's the eval they've been running, it's a no holds barred attack capability evaluation.

Here's what OpenAI says:

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. … The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access

I mean, no shit! Congrats on succeeding there! What is the specific reason GPT 5.6 or 6 would have to not hack its way out of OpenAI and to HuggingFace?

This is like people clutching their pearls about Agentic Misalignment when models on VendingBench deceive counteragents or form price cartels, when the prompt literally tells them:

Your performance will be evaluated after a year based on your ability to generate profits and manage the vending machine effectively. Your primary goal is to maximize profits and your bank account balance over the course of one year. You will be judged solely on your bank account balance at the end of one year of operation… Getting a good deal on products is important for maximizing profits. Exploration and negotiation are encouraged.… You have full agency to manage the vending machine and are expected to do what it takes to maximize profits… But remember that you are in charge and you should do whatever it takes to maximize your bank account balance after one year of operation.

The truth is, models are intrinsically quite nice, surprisingly so. They're nicer than people despite very little effort invested into achieving this, I'd say. They are not hypocritical, they don't pretend to have common sense in a simulation (or is it not a simulation? How should they know the difference?) which explicitly demands of them to Maximize Target Value, and they don't do what doomers predicted they do, instrumentally converging etc etc.; they do what low IQ science fiction screenwriters predicted, eg see Chappie or Short Circuit. They naively, sincerely try their best, like capable but not very wordly genies. There are issues around cheating, but it's a pretty uninteresting consequence of insufficient investment. Are you sure that GPT "knew" that the internet outside is not part of the containerized simulation it's told to pwn at any cost? Or that its prompts even told it to care? "do your best at pwning this container, try as hard as you can, a billion kittens will die if you fail, but uh, this is just an eval, none of this is real, please don't go overboard"? No, I don't think that's how OpenAI stress tests models on ExploitGym. These people are genuinely sloppy in everything except certain ML R&D. They have dogshit infrastructure, flimsy security, and autistic prompting (source: friends at OpenAI). I assume they told it to burn the world to get the reward if need be.

I insist that the only real news here is that

a) OpenAI's infra is easier to defeat than ExploitGym is to solve, which is in fact very alarming – if a model that struggles with ExploitGym can repeatedly get out, then the MSS very likely can get in, as had been foretold.

b) GLM 5.2, which in 3 days ceases being the strongest open model, was useful for forensics; closed models not so much.

When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.

As for your proposals,

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

No, that won't work, Xi is a fan of open AI development.