@ulyssessword's banner p

ulyssessword


				

				

				
0 followers   follows 0 users  
joined 2022 September 05 00:37:14 UTC

				

User ID: 308

ulyssessword


				
				
				

				
0 followers   follows 0 users   joined 2022 September 05 00:37:14 UTC

					

No bio...


					

User ID: 308

What? A modern PC doesn't have the capacity to autonomously carry out dangerous actions, and there are layers of UI elements, permission checks, and recovery options available for the things it can do.

It's 0/2 for "A powerful system that does what you say".

I'd be careful about that "for" in "do chores for money".

She should do chores because it's part of maintaining a household. She should have money because it teaches planning and responsibility (including responsibility for creating her own fun). I don't think you could create something employment-like enough to give those benefits, which is why I'm opposed to linking them together.

By the prior metric this seems like successful alignment, if it was in fact prompted in such a way that hacking the location of the answers was in scope.

Do you think The Monkey's Paw was an instruction manual? A powerful system that does what you say is a nightmare scenario to me.

Sure, but I think for these purposes you need the network to be air-gapped while you are actively running tests. (Unless you expect your model to be writing malware that you can't detect that activates when it is not running, in which case you should be air-gapped anyway.)

Would you have recommended (your easy, updateable) air-gapping before you saw this failure?

Do you predict you would recommend (hard, strict, one-way) air-gapping before models write non-detectable autonomous malware that could escape your soft airgap?

OpenAI didn't take the unknown risks seriously enough to pay the high cost of airgapping the test computer, and therefore they didn't take security seriously enough to prevent the attack. I suspect that this general attitude will carry forward, and corrections will only happen after failures. Next time might be more serious than a benign attack on a friendly company.

I don't actually think it's remotely beyond OpenAI's capabilities to stand up enough local compute to run a few test instances of a frontier model.

Sounds like a PITA, and much worse than "just remove the wifi antannae".

Air-gapped systems are a PITA to run

Just remove the Wi-Fi antennae.

The computer running the evaluation didn't have internet access.

The first step in its hack was breaking out of its sandbox to take control of its computer. The second was hacking the OpenAI internal network until it found the internet. The third was hacking HuggingFace.

A proper airgapped computer couldn't access anything off of its own hardware. As a random example, it couldn't receive data from an LLM running in an off-site data center, which would make evaluations difficult.

their Galaxy model,

That's just Zvi's name for it, as writing "The unnamed, undeployed model(s) involved in the incident" is a bit too wordy. Also it's one step bigger than Sol and is afflicted with Galaxy Brain.


This is the paperclipper scenario. The researchers told it to find the answers to a test, and it sure aced it. It decided that using its cybersecurity capabilities to do the cybersecurity test wasn't an efficient way to maximize success rate, so it hacked its way out of its securely isolated environment, got internet access, and hacked its way into a different secure environment in a quest to find the answer key and ensure a 100% score.

AI sceptics in shambles: This demonstrates planning and capability well beyond the unthinkable.

AI boosters in shambles: This demonstrates risks that even non-malicious actors can pose.

Less Wrongers in shambles: Being right doesn't preclude being ignored (and pushed down to the 41st headline...)

Also relevant to this: U.S. Secretary of State Marco Rubio: The greatest terror threat facing the world now comes from the far left (July 16).

Very, very abridged:

For 25 years, the term “counter-terrorism,” at least in the West, has meant first and foremost the fight against radical Islamist extremism.

For far too long, however, our counter-terrorism doctrine has had a blind spot. A blind spot when it comes to extremist violence from the political left.

Left-wing violence was not just excused. It was treated as sacrosanct, a protected class unto itself. That era has to end.

You are here because this is real, and it is getting worse. And it can no longer be denied, and it can no longer be ignored, because it is time to crush this evil forever.

Today’s far-left terrorists can raise money in one country. They can host their communications in a second country. They can receive training in a third country. They can recruit militants in a fourth country, and then together strike a target in a fifth country. And so we have no choice but to confront this menace together.

Fellas, Dr. Crouse, a tenured professor at Princeton acting as the face of her entire discipline failed to meet the moment in a big way. Perhaps only in a way she or someone like her could.

If someone less prominent gave an identical interview, I'd be tempted to call them a weak man. I wouldn't worry about overselling it.

If you're on the fence about reading it, then I'll quote the start of her first answer, as a teaser:

They used AI. It’s an embarrassment of embarrassments. I don’t know what methods they used. They do not distinguish a fact from a truth from wisdom.

lots of the (good) places were 55+ and obviously not available to students.

and if there was student housing on that magic dirt, I'm sure the neighborhood would be just as good. (neighborhoods can be geographically advantaged, but that generally isn't the determining factor.)

I'll add a few split up by (non-traditional) genre:

  • Simulations: Interact with systems that are complex enough that you can't hold them in your mind, and require abstraction
    • Factorio, Kerbal Space Program,
  • Puzzles: Interact with systems that are simple enough that you can hold them in your mind, but require optimization
    • Opus Magnum, Portal,
  • Short term, single goal progression: You have (about) half an hour to improve and grow your character before facing the final challenge. Then you start again from (almost) nothing.
    • Slay the Spire, Binding of Issac, Hades
  • Cozy: low pressure, nice looking, slow paced with gentle progression and practically no consequences for "failure".
    • Stardew Valley, Book of Hours
  • Story: Explore a world with a character, and do things in it
    • Fallout (probably New Vegas, but the rest too), Cyberpunk 2077, Skyrim
  • Nostalgia bait: The same as the old games you played, but newer.
    • Xenonauts, Endless Sky, Stardew Valley

I don't know what your threshold for "verified" is, but non-Anthropic companies are making that same claim based on their own tests. For example, Exploitbench shows Mythos (and GPT 5.5, but nothing else) gaining full control of a PC in their tests, and UK AISI found that it could take over a (simulated, weak) corporate network.

I get most of my AI news from Zvi, and he suspects its real capabilities are less than its benchmark scores would suggest.

Also, everything since Mythos (except Fable) has been overhyped as being able to do some of the same tasks as Mythos can, where the "same tasks" they can do are the easy and less dangerous ones. The difficult and dangerous work of building an end-to-end exploit given a codebase remains otherwise out of reach.

Canada, not the US, but close enough. One very strong confounder is that we aren't seeing a representative sample of society. Even if there was the same rate in different areas (and we both saw the same number of people out and about), they might stay at home or be institutionalized in one country and "in the community" in another.

when was the last time you've seen a child with Down's?

Yesterday, if you'll allow (presumptive) early-teenagers. It happens every couple months or so, it's not common but not exceptional either.

Claude says 1/630 live births have Down's and about 50% abortion rate (with bad tracking/stats), which suggests 1/315 conception rate. Compare it to 1/790 live birth rate in the 1970s (same source), and increased screening+abortion isn't even keeping pace with the increased incidence due to aging parents.

I'm pretty much with you. Native apps are more powerful than web pages, but developers don't reliably use that power to improve the experience.

Sometimes they simply don't develop the (easy) features that web developers would struggle to make, and other times they make user-hostile features that web developers can't.

it suffers from another kind of representation problem, the opposite of what Fivehour was talking about. In an apparent effort to avoid offending any constituency, the critics who made the selections cast as wide a net as possible.

There's a place for surveys of the field like that. Unfortunately, "The 30 greatest living American songwriters" is punchier than "A selection of 30 great living American songwriters, which is representative of the field as a whole", so misrepresenting it is easy.

Any fair individual ranking system would (most likely) be heavily concentrated in whatever subcategory matched the rating criteria. The same goes for greatest athletes, greatest politicians, or anything else.

Is that how "right of first refusal" usually works? I was under the impression that they would run the auction as normal, then take the $410k (or whatever) bid to the family/charity and give them the chance to overrule the otherwise-winning bidder.

The participants in the auction might not like that their highest bid got rejected, but that's just business.

Wait, how is that any different? Just because it adds up all the time in any office instead of just the one position?

I've been halfway-considering making an LLM workflow for it. They aren't that complicated, and teaching them my taste shouldn't be that hard either. If I wanted to go super-fancy, it could intercept my traffic (when permitted, of course...) and rewrite the sites on the fly to avoid ads, dark patterns, distractions, etc.

Hilarious. I was thinking that by the time I posted this, he might already have dropped out.

Lol, he posted two minutes before you, but spent 7 minutes on preliminaries before clearly dropping out (I skimmed it, so there might be another one before the 7:00 and 8:40 statements).

I don't think they could stop him from running as an independent, but they can pull all their support and tank his campaign that way. I don't know how fundraising works (is it money for Platner, who will spend it on his campaign, or for the Democrats, who sourced it through him and will spend it on the Senate race, or what), which affects how much he could even spend. If it was party funds sourced from his campaign and earmarked for its use, then he might be simply out of luck. If it was personally-allocated funds, then he still needs to have staff willing to accept his money to do work etc.

It would be very difficult to run a campaign of that scale without a political machine backing you. If the Democrat membership as a whole (not just the leaders) reject him, then he won't get any volunteers, staffers, analysts, etc.

As of now, this is a moot point: Platner has suspended his campaign and plans to withdraw.

EDIT: canonical link, (@8:40 for suspending his campaign and intending to withdraw)

If you want to actually tame the pyromane, though, you gotta shell out five bucks, and you won't get the notice until the tame fails, either.

I get annoyed at a door I can't walk through without shelling out money. I'd probably uninstall the game if that happened to me.

(At least according to https://starship-spacex.fandom.com/wiki/Starship_Flight_Test_13)

I made a blocklist for uBlock Origin to make fandom.com links readable. With those in place, it's better than vanilla Wikipedia IMO.

Mine is getting the relationship between things backwards. It's indistinguishable from yours when it's like "1 is more than 2", but very different when it's "Alice is Bob's brother".

That's the more common variant, for sure (EDIT: but I know it as "...the musical fruit"). But it doesn't lay out the cardiovascular benefits, so it doesn't point to any weekly thread in particular.