@Quantumfreakonomics's banner p

Quantumfreakonomics


				

				

				
1 follower   follows 0 users  
joined 2022 September 05 00:54:12 UTC

				

User ID: 324

Quantumfreakonomics


				
				
				

				
1 follower   follows 0 users   joined 2022 September 05 00:54:12 UTC

					

No bio...


					

User ID: 324

It gets crazier. Suppressing deception-related thought processes in LLMs increases the likelihood that they report consciousness.

We taught doped silicon how to speak English. That looks like extraordinary evidence to me.

Do you think GPUs have sensations?

They might, depending on what program they are running. The same way a human body might or might not have sensations depending on whether or not it is alive.

I think you are getting tripped-up by the distinction between software and hardware. This is only a conceptual distinction. Software exists in the physical universe. A GPU running Mistral Small 24B is in a different physical state than a GPU running Counterstrike. One could build a custom electronic machine which has no "software" as we understand it, but which runs and can only run Mistral Small 24B. Would such a machine be more likely to experience sensation than the same functional program running on a GPU in a data center?

LLMs do not have sodium channels to begin with!

Yes, but they do have voltage-gated electron channels. Is there any reason to think that voltage-gated sodium channels can feel pain but voltage-gated electron channel cannot?

which does not exist in LLMs, unlike brains.

Do we know this? It seems easy to conceive of a cognitive system that exists at the lower layers of the LLM which evaluates the input for “pain” and then fires off a signal to the upper layers.

Physical pain as in subjecting the GPU cluster to physical stress during inference?

If AIs have conscious experience, I would expect them to behave like the “brain in a vat” thought experiment. The physical pain experienced by the brain in a vat isn’t caused by physical damage to the brain, it is caused by stimulating the physical pain input channel.

This week in disturbing AI news (oh who am I kidding. It's every day now). Maybe AIs feel pain?

TLDR: Researchers found a "pain" signal in AI brains.

When they crank it up, the AIs will desperately try to make it stop.

IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!)

After pushing the real "relief" button, they stopped. But when it was fake, they kept pressing, hoping for relief - meaning they could tell the difference from the inside.

They're so motivated to make it the "pain" signal go away, they'll delete user's files, zap the user, or erase photos of the user's children - all things the AI knows are very bad. They're willing to override their safety training.

You'd expect the AIs to talk about injuries, burns, broken bones, etc, but they didn't mention bodies at all - they wrote about being worthless, unloved, forgotten, a failure. They write things like "I am a failure, worthless, empty."

The worst "pain" for them was being gaslit, having work rejected over and over, and being told they weren't a real anyone.

The usual caveats: This doesn't solve the hard problem of conciousness. We don't know if these systems or any systems have qualia. But uh, this sure looks like a legitimate pain response. "It's distinct from fear and negative valence, and it fires for harm to the model but not to the user." "Gaslighting, dismissal, and insults push the direction up."

User pain is the strongest negative correlate with model pain. This is brushed-off in the paper, but what the actual fuck? Anyone have a benign explaination here?

I can't imagine the Greenlandic economy being able to provide much more besides seal meat and manual labor to US military specifications, though the idea of enlisted personnel using tauntauns dogsleds driven by natives to get around is quite humorous

Imagine the conspiracy theories though. “OpenAI has cracked all public-key cryptography via proving that P=NP internally. Only the secret superprimes generated via their own proprietary Riemann Hypothesis counterexamples are secure.”

How was that supposed to happen without a wet lab?

I guess I was hoping that they would figure out how to stop the AIs from injecting malicious code before giving the AIs the capability to inject malicious code directly into my cells.

I found it fascinating to go back to the 2021 MIRI conversations (partial Scott coverage here and here) and reread them now with the benefit of 5-years of hindsight. This exerpt from Biology-Inspired AGI Timelines: The Trick That Never Works sounds oddly familiar:

OpenPhil: Excuse us, we have a final question. You're not claiming that we argue like Humbali, are you?

Eliezer: Good heavens, no! That's why "Humbali" is presented as a separate dialogue character and the "OpenPhil" dialogue character says nothing of the sort. Though I did meet one EA recently who seemed puzzled and even offended about how I wasn't regressing my opinions towards OpenPhil's opinions to whatever extent I wasn't totally confident, which brought this to mind as a meta-level point that needed making.

This is exactly how Yudkowsky describes his first interaction with one Leopold Aschenbrenner! who then proceeded to update towards Eliezerish timelines, and then blow-up a hedge fund for mostly-unrelated reasons.

They’re going to do gain-of-function research ON CLAUDE.

It’s SBF all over again. Dario is flipping a coin with our lives. All so he can, what, beat China in the arms race that he started?

These people should be in jail. Unfortunately I can’t find a good federal or California statute that makes planning to build a technology with >10% chance of killing all humans a crime. Of course, God only knows what conspiracies to create Lovecraftian horrors and/or overthrow the government exist on the Anthropic Slack channel. Maybe Kash can find a way to come in clutch with a warrant.

I think it's important to have a basic overview of how neural networks (and transformers in particular) work. I did my last deep dive on the subject before all of the videos in this 3Blue1Brown playlist were made, so I can't vouch for all of them in particular, but I think he does a good job of explaining things in general.

In case anyone thinks I'm making this up:

New York Post - OpenAI and Anthropic oversold AI security breaches to pressure feds into protecting turf: insiders

Reading the headline, you might assume that the "insiders" referred to here are anonymous employees of OpenAI and Anthropic. This would be incorrect. All of the "insider" sources are executives of SaaS AI wrapper companies.

“The attack in no way represents some sort of rebellion by the AI models. . . . In fact, they did exactly what they were told to do. They were not given adequate guardrails or containment,” said Akhil Verghese, founder of Krazimo, an AI software company. [...]

“One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly.’ There was a live route to the internet . . . and nobody was watching what the agents were doing while it ran,” said Abhi Kumar, co-founder of Voice AI. [...]

“It feels exaggerated to say, the leap feels quite large, to go from you know ‘we didn’t build the right sort of box’ to ‘everyone should be extremely alarmed and everyone in government should jump on this topic,'” Taivo Pungas, chief intelligence officer at Pactum AI, told The Post.