william

Posts
Comments

Posts

Principles for the AGI Race 2024-08-30T14:29:41.074Z

Transformer Circuit Faithfulness Metrics Are Not Robust 2024-07-12T03:47:30.077Z

William_S's Shortform 2023-03-22T18:13:18.731Z

Thoughts on refusing harmful requests to large language models 2023-01-19T19:49:22.989Z

Prize for Alignment Research Tasks 2022-04-29T08:57:04.290Z

Is there an intuitive way to explain how much better superforecasters are than regular forecasters? 2020-02-19T01:07:52.394Z

Machine Learning Projects on IDA 2019-06-24T18:38:18.873Z

Reinforcement Learning in the Iterated Amplification Framework 2019-02-09T00:56:08.256Z

HCH is not just Mechanical Turk 2019-02-09T00:46:25.729Z

Amplification Discussion Notes 2018-06-01T19:03:35.294Z

Understanding Iterated Distillation and Amplification: Claims and Oversight 2018-04-17T22:36:29.562Z

Improbable Oversight, An Attempt at Informed Oversight 2017-05-24T17:43:53.000Z

Informed Oversight through Generalizing Explanations 2017-05-24T17:43:39.000Z

Proposal for an Implementable Toy Model of Informed Oversight 2017-05-24T17:43:13.000Z

Comments

Comment by William_S on Richard Ngo's Shortform · 2025-03-30T14:36:14.716Z · LW · GW

Would be interested in a quick write-up of what you think are the most important virtues you'd want for AI systems, seems good in terms of having things to aim towards instead of just aiming away from.

Comment by William_S on William_S's Shortform · 2025-03-15T22:07:17.250Z · LW · GW

Initial version for firefox, code at https://github.com/william-r-s/MindfulBlocker, extension file at https://github.com/william-r-s/MindfulBlocker/releases/tag/v0.2.0

Comment by William_S on Daniel Kokotajlo's Shortform · 2025-03-12T18:27:27.641Z · LW · GW

Maybe there's an MVP of having some independent organization ask new AIs about their preferences + probe those preferences for credibility (e.g. are they stable under different prompts, do AIs show general signs of having coherent preferences), and do this through existing apis

Comment by William_S on Daniel Kokotajlo's Shortform · 2025-03-12T18:22:17.996Z · LW · GW

I think the weirdness points are more important, this still seems like a weird thing for a company to officially do, e.g. there'd be snickering news articles about it. So if some individuals could do this independently might be easier

Comment by William_S on Daniel Kokotajlo's Shortform · 2025-03-12T18:07:47.503Z · LW · GW

How large a reward pot do you think is useful for this? Maybe would be easier to get a couple of lab employees to chip in some equity vs. getting a company to spend weirdness points on this. Or maybe could create a human whistleblower reward program that credibly promises to reward AIs on the side.

Comment by William_S on Principles for the AGI Race · 2025-02-16T23:10:23.205Z · LW · GW

I think it's somewhat blameworthy to not think about these questions at all though

Comment by William_S on Principles for the AGI Race · 2025-02-16T22:23:30.418Z · LW · GW

On reflection there was something missing from my perspective here, which is that taking any action based on principles depends on pragmatic considerations, like if you leave are there better alternatives? How much power do you really have? I think I don't fault someone who thinks this through and decides that something is wrong but there's no real way to do anything about it. I do think you should try to maintain some sense of what is wrong and what the right direction would be, look out for ways to push in that direction. E.g. working at a lab but maintaining some sense of "this is how much of a chance it looks like pause activism would need before I'd quite and endorse a pause".

Comment by William_S on Principles for the AGI Race · 2025-02-16T22:13:26.403Z · LW · GW

I think I was just conflating different kinds of decisions here, and imagining arguing with people with very different conceptions of what are important to count in costs and benefits, and a bit confused. On reflection I don't endorse 10x margin in terms of like percentage points of x-risk. And like maybe margin is sort of a crutch, maybe the thing I want more is like "95% chance of being net-positive, considering possibility you're kind of biased". I still think you should be suspicious of "the case exactly balance lets ship'

Comment by William_S on Principles for the AGI Race · 2025-02-16T22:09:21.102Z · LW · GW

Yeah this part is pretty under-defined, I was maybe falling into the trap of being too idealistic, and I'm probably less optimistic about this than I was when writing it before. I think there's something directionally important here, are you trying at all to expand the circle of accountability at all, even if you're being cautious about expanding it because you're afraid of things breaking down?

Comment by William_S on 6 (Potential) Misconceptions about AI Intellectuals · 2025-02-16T21:23:50.693Z · LW · GW

Would be nice to have a llm+prompt that tries to produce reasonable AI strategy advice based on a summary of the current state of play, have some way to validate that it's reasonable, be able to see how it updates as events unfold.

Comment by William_S on 6 (Potential) Misconceptions about AI Intellectuals · 2025-02-16T21:20:39.647Z · LW · GW

A couple advantages for AI intellectuals could be:
- being able to rerun based on different inputs, see how their analysis changes function of those inputs
- being able to view full reasoning traces (while also not the full story, probably more of the full story than what goes on with human reasoning, good intellectuals already try to share their process but maybe can do better/use this to weed out clearly bad approaches)

Comment by William_S on William_S's Shortform · 2025-02-16T19:47:25.837Z · LW · GW

Yep, I've used those, with some effectiveness but also tend to just like get used to it over time, form a habit of mindlessly jumping through the hoops. Hypothesis here is that having to justify what you're doing would be more effective at changing habits.

Comment by William_S on William_S's Shortform · 2025-02-16T18:32:39.294Z · LW · GW

LLM-based application I'd like to exist:
Web browser addon for firefox that has blocklists of websites, when you try to visit one you have to have a conversation with Claude about why you want to visit it in this moment, convince Claude to let you bypass the block for a limited period of time for your specific purpose (let you customize the claude prompt with info about why you set up the block in the first place).
Wanting to use for things like news, social media where it's a bit too much to try to completely block, but I've got bad habits around checking too frequently.
Bonus: be able to let the LLM read the website for you and answer questions without showing you the page, like is there anything new about X.

Comment by William_S on I found >800 orthogonal "write code" steering vectors · 2024-07-15T19:25:28.880Z · LW · GW

Hypothesis: each of these vectors representing a single token that is usually associated with code, vectors says "I should output this token soon", and the model then plans around that to produce code. But adding vectors representing code tokens doesn't necessarily produce another vector representing a code token, so that's why you don't see compositionality. Does somewhat seem plausible that there might be ~800 "code tokens" in the representation space.

Comment by William_S on Habryka's Shortform Feed · 2024-07-05T23:56:55.319Z · LW · GW

Absent evidence to the contrary, for any organization one should assume board members were basically selected by the CEO. So hard to get assurance about true independence, but it seems good to at least to talk to someone who isn't a family member/close friend.

Comment by William_S on Habryka's Shortform Feed · 2024-07-05T17:53:58.045Z · LW · GW

Good that it's clear who it goes to, though if I was an anthropic I'd want an option to escalate to a board member who isn't Dario or Daniella, in case I had concerns related to the CEO

Comment by William_S on 80,000 hours should remove OpenAI from the Job Board (and similar EA orgs should do similarly) · 2024-07-05T17:33:15.665Z · LW · GW

I do think 80k should have more context on OpenAI but also any other organization that seems bad with maybe useful roles. I think people can fail to realize the organizational context if it isn't pointed out and they only read the company's PR.

Comment by William_S on Habryka's Shortform Feed · 2024-07-01T18:59:50.564Z · LW · GW

I agree that this kind of legal contract is bad, and Anthropic should do better. I think there are a number of aggrevating factors which made the OpenAI situation extrodinarily bad, and I'm not sure how much these might obtain regarding Anthropic (at least one comment from another departing employee about not being offered this kind of contract suggest the practice is less widespread).

-amount of money at stake
-taking money, equity or other things the employee believed they already owned if the employee doesn't sign the contract, vs. offering them something new (IANAL but in some cases, this could be a felony "grand theft wages" under California law if a threat to withhold wages for not signing a contract is actually carried out, what kinds of equity count as wages would be a complex legal question)
-is this offered to everyone, or only under circumstances where there's a reasonable justification?
-is this only offered when someone is fired or also when someone resigns?
-to what degree are the policies of offering contracts concealed from employees?
-if someone asks to obtain legal advice and/or negotiate before signing, does the company allow this?
-if this becomes public, does the company try to deflect/minimize/only address issues that are made publically, or do they fix the whole situation?
-is this close to "standard practice" (which doesn't make it right, but makes it at least seem less deliberately malicious), or is it worse than standard practice?
-are there carveouts that reduce the scope of the non-disparagement clause (explicitly allow some kinds of speech, overriding the non-disparagement)?
-are there substantive concerns that the employee has at the time of signing the contract, that the agreement would prevent discussing?
-are there other ways the company could retaliate against an employee/departing employee who challenges the legality of contract?

I think with termination agreements on being fired there's often 1. some amount of severance offered 2. a clause that says "the terms and monetary amounts of this agreement are confidential" or similar. I don't know how often this also includes non-disparagement. I expect that most non-disparagement agreements don't have a term or limits on what is covered.

I think a steelman of this kind of contract is: Suppose you fire someone, believe you have good reasons to fire them, and you think that them loudly talking about how it was unfair that you fired them would unfairly harm your company's reputation. Then it seems somewhat reasonable to offer someone money in exchange for "don't complain about being fired". The person who was fired can then decide whether talking about it is worth more than the money being offered.

However, you could accomplish this with a much more limited contract, ideally one that lets you disclose "I signed a legal agreement in exchange for money to not complain about being fired", and doesn't cover cases where "years later, you decide the company is doing the wrong thing based on public information and want to talk about that publically" or similar.

I think it is not in the nature of most corporate lawyers to think about "is this agreement giving me too much power?" and most employees facing such an agreement just sign it without considering negotiating or challenging the terms.

For any future employer, I will ask about their policies for termination contracts before I join (as this is when you have the most leverage, if they give you an offer they want to convince you to join).

Comment by William_S on Buck's Shortform · 2024-06-25T21:01:29.801Z · LW · GW

Would be nice if it was based on "actual robot army was actually being built and you have multiple confirmatory sources and you've tried diplomacy and sabotage and they've both failed" instead of "my napkin math says they could totally build a robot army bro trust me bro" or "they totally have WMDs bro" or "we gotta blow up some Japanese civilians so that we don't have to kill more Japanese civilians when we invade Japan bro" or "dude I'm seeing some missiles on our radar, gotta launch ours now bro".

Comment by William_S on Buck's Shortform · 2024-06-24T23:43:23.683Z · LW · GW

Relevant paper discussing risks of risk assessments being wrong due to theory/model/calculation error. Probing the Improbable: Methodological Challenges for Risks with Low Probabilities and High Stakes

Based on the current vibes, I think that suggest that methodological errors alone will lead to significant chance of significant error for any safety case in AI.

Comment by William_S on Buck's Shortform · 2024-06-24T23:13:24.608Z · LW · GW

IMO it's unlikely that we're ever going to have a safety case that's as reliable as the nuclear physics calculations that showed that the Trinity Test was unlikely to ignite the atmosphere (where my impression is that the risk was mostly dominated by risk of getting the calculations wrong). If we have something that is less reliable, then will we ever be in a position where only considering the safety case gives a low enough probability of disaster for launching an AI system beyond the frontier where disastrous capabilities are demonstrated?
Thus, in practice, decisions will probably not be made on a safety case alone, but also based on some positive case of the benefits of deployment (e.g. estimated reduced x-risk, advancing the "good guys" in the race, CEO has positive vibes that enough risk mitigation has been done, etc.). It's not clear what role governments should have in assessing this, maybe we can only get assessment of the safety case, but it's useful to note that safety cases won't be the only thing informs these decisions.

This situation is pretty disturbing, and I wish we had a better way, but it still seems useful to push the positive benefit case more towards "careful argument about reduced x-risk" and away from "CEO vibes about whether enough mitigation has been done".

Comment by William_S on Richard Ngo's Shortform · 2024-06-22T08:22:24.241Z · LW · GW

Imo I don't know if we have evidence that Anthropic deliberately cultivated or significantly benefitted from the appearance of a commitment. However if an investor or employee felt like they made substantial commitments based on this impression and then later felt betrayed that would be more serious. (The story here is I think importantly different from other stories where I think there were substantial benefits from commitment appearance and then violation)

Comment by William_S on Richard Ngo's Shortform · 2024-06-22T08:19:12.869Z · LW · GW

Everyone is afraid of the AI race, and hopes that one of the labs will actually end up doing what they think is the most responsible thing to do. Hope and fear is one hell of a drug cocktail, makes you jump to the conclusions you want based on the flimsiest evidence. But the hangover is a bastard.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-22T00:37:05.137Z · LW · GW

Really, the race started more when OpenAI released GPT-4, it's been going on for a while, this is just another event that makes it clear.

Comment by William_S on On OpenAI’s Model Spec · 2024-06-22T00:16:18.663Z · LW · GW

Would be interesting philosophical experiment to have models trained on model spec v1 then try to improve their model spec for version v2, will this get better or go off the rails?

Comment by William_S on What distinguishes "early", "mid" and "end" games? · 2024-06-22T00:13:18.895Z · LW · GW

You get more discrete transitions when one s-curve process takes the lead from another s-curve process, e.g. deep learning taking over from other AI methods.

Comment by William_S on What distinguishes "early", "mid" and "end" games? · 2024-06-22T00:11:39.150Z · LW · GW

Probably shouldn't limit oneself from thinking only in terms of 3 game phases or fitting into one specific game, in general can have n-phases where different phrases have different characteristics.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-21T04:47:43.433Z · LW · GW

If anyone wants to work on this, there's a contest with $50K and $20K prizes for creating safety relevant benchmarks. https://www.mlsafety.org/safebench

Comment by William_S on Richard Ngo's Shortform · 2024-06-21T03:29:40.465Z · LW · GW

I think that's how people should generally react in the absence of harder commitments and accountability measures.

Comment by William_S on Richard Ngo's Shortform · 2024-06-21T03:25:23.198Z · LW · GW

I think the right way to think about verbal or written commitments is that they increase the costs of taking a certain course of action. A legal contract can mean that the price is civil lawsuits leading to paying a financial price. A non-legal commitment means if you break it, the person you made the commitment to gets angry at you, and you gain a reputation for being the sort of person who breaks commitments. It's always an option for someone to break the commitment and pay the price, even laws leading to criminal penalties can be broken if someone is willing to run the risk or pay the price.

In this framework, it's reasonable to be somewhat angry at someone or some corporation who breaks a soft commitment to you, in order to increase the perceived cost of breaking soft commitments to you and people like you.

People on average maybe tend more towards keeping important commitments due to reputational and relationship cost, but maybe corporations as groups of people tend to think only in terms of financial and legal costs, so are maybe more willing to break soft commitments (especially, if it's an organization where one person makes the commitment but then other people break it). So for relating to corporations, you should be more skeptical of non-legally binding commitments (and even for legally binding commitments, pay attention to the real price of breaking it).

Comment by William_S on Richard Ngo's Shortform · 2024-06-21T02:30:45.558Z · LW · GW

Yeah, I think it's good if labs are willing to make more "cheap talk" statements of vague intentions, so you can learn how they think. Everyone should understand that these aren't real commitments, and not get annoyed if these don't end up meaning anything. This is probably the best way to view "statements by random lab employees".

Imo would be good to have more "changeable commitments" too in between, statements that are "we'll do policy X until we change the policy, when we do we commit to clearly informing everyone about the change" which is maybe more the current status of most RSPs.

Comment by William_S on William_S's Shortform · 2024-06-20T23:50:39.003Z · LW · GW

I'd have more confidence in Anthropic's governance if the board or LTBT had some fulltime independent members who weren't employees. IMO labs should consider paying a fulltime salary but no equity to board members, through some kind of mechanism where the money is still there and paid for X period of time in the future, even if the lab dissolved, so no incentive to avoid actions that would cost the lab. Board salaries could maybe be pegged to some level of technical employee salary, so that technical experts could take on board roles. Boards full of busy people really can't do their job of checking whether the organization is fullfilling its stated mission, and IMO this is one of the most important jobs in the world right now. Also, fulltime board members would have fewer conflicts of interest outside of the lab (since they won't be in some other fulltime job that might conflict).

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T22:20:15.944Z · LW · GW

Like, in Chess you start off with a state where many pieces can't move in the early game, in the middle game many pieces are in play moving around and trading, then in the end game it's only a few pieces, you know what the goal is, roughly how things will play out.

In AI it's like only a handful of players, then ChatGPT/GPT-4 came out and now everyone is rushing to get in (my mark of the start of the mid-game), but over time probably many players will become irrelevant or fold as the table stakes (training costs) get too high.

In my head the end-game is when the AIs themselves start becoming real players.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T22:12:39.994Z · LW · GW

Also you would need clarity on how to measure the commitment.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T22:06:26.969Z · LW · GW

It's quite possible that anthropic has some internal definition of "not meaningfully advancing the capabilities frontier" that is compatible with this release. But imo they shouldn't get any credit unless they explain it.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T22:04:06.448Z · LW · GW

Would be nice, but I was thinking of metrics that require "we've done the hard work of understanding our models and making them more reliable", better neuron explanation seems more like it's another smartness test.

Comment by William_S on Zach Stein-Perlman's Shortform · 2024-06-20T21:52:36.348Z · LW · GW

IMO it might be hard for Anthropic to communicate things about not racing because it might piss off their investors (even if in their hearts they don't want to race).

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T21:05:10.854Z · LW · GW

https://x.com/alexalbert__/status/1803837844798189580

Not sure about the accuracy of this graph, but the general picture seems to match what companies claim, and the vibe is racing.

Do think that there are distinct questions about "is there a race" vs. "will this race action lead to bad consequences" vs. "is this race action morally condemnable". I'm hoping that this race action is not too consequentially bad, maybe it's consequentially good, maybe it still has negative Shapely value even if expected value is okay. There is some sense in which it is morally icky.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T20:51:33.356Z · LW · GW

To be clear, I think the race was already kind of on, it's not clear how much this specific action gets credit assignment and it's spread out to some degree. Also not clear if there's really a viable alternative strategy here...

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T19:20:13.122Z · LW · GW

In my mental model, we're still in the mid-game, not yet in the end-game.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T19:19:18.699Z · LW · GW

Idk there's probably multiple ways to define racing, some of them are on at least

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T19:09:14.497Z · LW · GW

I'm disappointed that there weren't any non-capability metrics reported. IMO it would be good if companies could at least partly race and market on reliability metrics like "not hallucinating" and "not being easy to jailbreak".

Edit: As pointed out in reply, addendum contains metrics on refusals which show progress, yay! Broader point still stands, I wish there were more measurements and they were more prominent.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T19:06:13.342Z · LW · GW

The race is on.

Comment by William_S on Claude 3.5 Sonnet · 2024-06-20T19:01:29.344Z · LW · GW

IMO if any lab makes some kind of statement or commitment, you should treat this as "we think right now that we'll want to do this in the future unless it's hard or costly", unless you can actually see how you would sue them or cause a regulator to fine them if they violate the commitment. This doesn't mean weaker statements have no value.

Comment by William_S on Ilya Sutskever created a new AGI startup · 2024-06-19T18:12:16.764Z · LW · GW

If anyone says "We plan to advance capabilities as fast as possible while making sure our safety always remains ahead." you should really ask for the details of what this means, how to measure whether safety is ahead. (E.g. is it "we did the bare minimum to make this product tolerable to society" vs. "we realize how hard superalignment will be and will be investing enough to have independent experts agree we have a 90% chance of being able to solve superalignment before we build something dangerous")

Comment by William_S on Ilya Sutskever created a new AGI startup · 2024-06-19T18:06:59.795Z · LW · GW

I do hope he will continue to contribute to the field of alignment research.

Comment by William_S on Ilya Sutskever created a new AGI startup · 2024-06-19T18:02:39.548Z · LW · GW

I don't trust Ilya Sutskever to be the final arbiter of whether a Superintelligent AI design is safe and aligned. We shouldn't trust any individual, especially if they are the ones building such a system to claim that they've figured out how to make it safe and aligned. At minimum, there should be a plan that passes review by a panel of independent technical experts. And most of this plan should be in place and reviewed before you build the dangerous system.

Comment by William_S on Boycott OpenAI · 2024-06-19T17:49:15.486Z · LW · GW

In my opinion, it's reasonable to change which companies you want to do business with, but it would be more helpful to write letters to politicians in favor of reasonable AI regulation (e.g. SB 1047, with suggested amendments if you have concerns about the current draft). I think it's bad if the public has to play the game of trying to pick which AI developer seems the most responsible, better to try to change the rules of the game so that isn't necessary.

Also it's generally helpful to write about which labs seem more responsible/less responsible (which you are doing here), what you think labs should do instead of current practices. Bonus points for designing ways to test which deployed models are more safe and reliable, e.g. writing some prompts to use as litmus tests.

Comment by William_S on Non-Disparagement Canaries for OpenAI · 2024-06-03T23:46:54.631Z · LW · GW

Language in the emails included:

"If you executed the Agreement, we write to notify you that OpenAI does not intend to enforce the Agreement"

I assume this also communicates that OpenAI doesn't intend to enforce the self-confidentiality clause in the agreement

Comment by William_S on Non-Disparagement Canaries for OpenAI · 2024-06-03T23:41:32.821Z · LW · GW

Evidence could look like 1. Someone was in a position where they had to make a judgement about OpenAI and was in a position of trust 2. They said something bland and inoffensive about OpenAI 3. Later, independently you find that they likely would have known about something bad that they likely weren't saying because of the nondisparagement agreement (instead of ordinary confidentially agreements).

This requires some model of "this specific statement was influenced by the agreement" instead of just "you never said anything bad about OpenAI because you never gave opinions on OpenAI".

I think one should require this kind of positive evidence before calling it a "serious breach of trust", but people can make their own judgement about where that bar should be.

User info

Posts

Comments