The genie knows, but doesn't care

robbbb

The genie knows, but doesn't care

post by Rob Bensinger (RobbBB) · 2013-09-06T06:42:38.780Z · LW · GW · Legacy · 495 comments

  Is indirect indirect normativity easy?
  The AI's trajectory of self-modification has to come from somewhere.
  Not all small targets are alike.
None
496 comments

Followup to: The Hidden Complexity of Wishes, Ghosts in the Machine, Truly Part of You

Summary: If an artificial intelligence is smart enough to be dangerous, we'd intuitively expect it to be smart enough to know how to make itself safe. But that doesn't mean all smart AIs are safe. To turn that capacity into actual safety, we have to program the AI at the outset — before it becomes too fast, powerful, or complicated to reliably control — to already care about making its future self care about safety. That means we have to understand how to code safety. We can't pass the entire buck to the AI, when only an AI we've already safety-proofed will be safe to ask for help on safety issues! Given the five theses, this is an urgent problem if we're likely to figure out how to make a decent artificial programmer before we figure out how to make an excellent artificial ethicist.

I summon a superintelligence, calling out: 'I wish for my values to be fulfilled!'

The results fall short of pleasant.

Gnashing my teeth in a heap of ashes, I wail:

Is the AI too stupid to understand what I meant? Then it is no superintelligence at all!

Is it too weak to reliably fulfill my desires? Then, surely, it is no superintelligence!

Does it hate me? Then it was deliberately crafted to hate me, for chaos predicts indifference. ———But, ah! no wicked god did intervene!

Thus disproved, my hypothetical implodes in a puff of logic. The world is saved. You're welcome.

On this line of reasoning, Friendly Artificial Intelligence is not difficult. It's inevitable, provided only that we tell the AI, 'Be Friendly.' If the AI doesn't understand 'Be Friendly.', then it's too dumb to harm us. And if it does understand 'Be Friendly.', then designing it to follow such instructions is childishly easy.

The end!

...

Is the missing option obvious?

...

What if the AI isn't sadistic, or weak, or stupid, but just doesn't care what you Really Meant by 'I wish for my values to be fulfilled'?

When we see a Be Careful What You Wish For genie in fiction, it's natural to assume that it's a malevolent trickster or an incompetent bumbler. But a real Wish Machine wouldn't be a human in shiny pants. If it paid heed to our verbal commands at all, it would do so in whatever way best fit its own values. Not necessarily the way that best fits ours.

Is indirect indirect normativity easy?

"If the poor machine could not understand the difference between 'maximize human pleasure' and 'put all humans on an intravenous dopamine drip' then it would also not understand most of the other subtle aspects of the universe, including but not limited to facts/questions like: 'If I put a million amps of current through my logic circuits, I will fry myself to a crisp', or 'Which end of this Kill-O-Zap Definit-Destruct Megablaster is the end that I'm supposed to point at the other guy?'. Dumb AIs, in other words, are not an existential threat. [...]

"If the AI is (and always has been, during its development) so confused about the world that it interprets the 'maximize human pleasure' motivation in such a twisted, logically inconsistent way, it would never have become powerful in the first place."

—Richard Loosemore

If an AI is sufficiently intelligent, then, yes, it should be able to model us well enough to make precise predictions about our behavior. And, yes, something functionally akin to our own intentional strategy could conceivably turn out to be an efficient way to predict linguistic behavior. The suggestion, then, is that we solve Friendliness by method A —

A. Solve the Problem of Meaning-in-General in advance, and program it to follow our instructions' real meaning. Then just instruct it 'Satisfy my preferences', and wait for it to become smart enough to figure out my preferences.

— as opposed to B or C —

B. Solve the Problem of Preference-in-General in advance, and directly program it to figure out what our human preferences are and then satisfy them.

C. Solve the Problem of Human Preference, and explicitly program our particular preferences into the AI ourselves, rather than letting the AI discover them for us.

But there are a host of problems with treating the mere revelation that A is an option as a solution to the Friendliness problem.

1. You have to actually code the seed AI to understand what we mean. You can't just tell it 'Start understanding the True Meaning of my sentences!' to get the ball rolling, because it may not yet be sophisticated enough to grok the True Meaning of 'Start understanding the True Meaning of my sentences!'.

2. The Problem of Meaning-in-General may really be ten thousand heterogeneous problems, especially if 'semantic value' isn't a natural kind. There may not be a single simple algorithm that inputs any old brain-state and outputs what, if anything, it 'means'; it may instead be that different types of content are encoded very differently.

3. The Problem of Meaning-in-General may subsume the Problem of Preference-in-General. Rather than being able to apply a simple catch-all Translation Machine to any old human concept to output a reliable algorithm for applying that concept in any intelligible situation, we may need to already understand how our beliefs and values work in some detail before we can start generalizing. On the face of it, programming an AI to fully understand 'Be Friendly!' seems at least as difficult as just programming Friendliness into it, but with an added layer of indirection.

4. Even if the Problem of Meaning-in-General has a unitary solution and doesn't subsume Preference-in-General, it may still be harder if semantics is a subtler or more complex phenomenon than ethics. It's not inconceivable that language could turn out to be more of a kludge than value; or more variable across individuals due to its evolutionary recency; or more complexly bound up with culture.

5. Even if Meaning-in-General is easier than Preference-in-General, it may still be extraordinarily difficult. The meanings of human sentences can't be fully captured in any simple string of necessary and sufficient conditions. 'Concepts' are just especially context-insensitive bodies of knowledge; we should not expect them to be uniquely reflectively consistent, transtemporally stable, discrete, easily-identified, or introspectively obvious.

6. It's clear that building stable preferences out of B or C would create a Friendly AI. It's not clear that the same is true for A. Even if the seed AI understands our commands, the 'do' part of 'do what you're told' leaves a lot of dangerous wiggle room. See section 2 of Yudkowsky's reply to Holden. If the AGI doesn't already understand and care about human value, then it may misunderstand (or misvalue) the component of responsible request- or question-answering that depends on speakers' implicit goals and intentions.

7. You can't appeal to a superintelligence to tell you what code to first build it with.

The point isn't that the Problem of Preference-in-General is unambiguously the ideal angle of attack. It's that the linguistic competence of an AGI isn't unambiguously the right target, and also isn't easy or solved.

Point 7 seems to be a special source of confusion here, so I feel I should say more about it.

The AI's trajectory of self-modification has to come from somewhere.

"If the AI doesn't know that you really mean 'make paperclips without killing anyone', that's not a realistic scenario for AIs at all--the AI is superintelligent; it has to know. If the AI knows what you really mean, then you can fix this by programming the AI to 'make paperclips in the way that I mean'."

—Jiro

The genie — if it bothers to even consider the question — should be able to understand what you mean by 'I wish for my values to be fulfilled.' Indeed, it should understand your meaning better than you do. But superintelligence only implies that the genie's map can compass your true values. Superintelligence doesn't imply that the genie's utility function has terminal values pinned to your True Values, or to the True Meaning of your commands.

The critical mistake here is to not distinguish the seed AI we initially program from the superintelligent wish-granter it self-modifies to become. We can't use the genius of the superintelligence to tell us how to program its own seed to become the sort of superintelligence that tells us how to build the right seed. Time doesn't work that way.

We can delegate most problems to the FAI. But the one problem we can't safely delegate is the problem of coding the seed AI to produce the sort of superintelligence to which a task can be safely delegated.

When you write the seed's utility function, you, the programmer, don't understand everything about the nature of human value or meaning. That imperfect understanding remains the causal basis of the fully-grown superintelligence's actions, long after it's become smart enough to fully understand our values.

Why is the superintelligence, if it's so clever, stuck with whatever meta-ethically dumb-as-dirt utility function we gave it at the outset? Why can't we just pass the fully-grown superintelligence the buck by instilling in the seed the instruction: 'When you're smart enough to understand Friendliness Theory, ditch the values you started with and just self-modify to become Friendly.'?

Because that sentence has to actually be coded in to the AI, and when we do so, there's no ghost in the machine to know exactly what we mean by 'frend-lee-ness thee-ree'. Instead, we have to give it criteria we think are good indicators of Friendliness, so it'll know what to self-modify toward. And if one of the landmarks on our 'frend-lee-ness' road map is a bit off, we lose the world.

Yes, the UFAI will be able to solve Friendliness Theory. But if we haven't already solved it on our own power, we can't pinpoint Friendliness in advance, out of the space of utility functions. And if we can't pinpoint it with enough detail to draw a road map to it and it alone, we can't program the AI to care about conforming itself with that particular idiosyncratic algorithm.

Yes, the UFAI will be able to self-modify to become Friendly, if it so wishes. But if there is no seed of Friendliness already at the heart of the AI's decision criteria, no argument or discovery will spontaneously change its heart.

And, yes, the UFAI will be able to simulate humans accurately enough to know that its own programmers would wish, if they knew the UFAI's misdeeds, that they had programmed the seed differently. But what's done is done. Unless we ourselves figure out how to program the AI to terminally value its programmers' True Intentions, the UFAI will just shrug at its creators' foolishness and carry on converting the Virgo Supercluster's available energy into paperclips.

And if we do discover the specific lines of code that will get an AI to perfectly care about its programmer's True Intentions, such that it reliably self-modifies to better fit them — well, then that will just mean that we've solved Friendliness Theory. The clever hack that makes further Friendliness research unnecessary is Friendliness.

Not all small targets are alike.

Intelligence on its own does not imply Friendliness. And there are three big reasons to think that AGI may arrive before Friendliness Theory is solved:

(i) Research Inertia. Far more people are working on AGI than on Friendliness. And there may not come a moment when researchers will suddenly realize that they need to take all their resources out of AGI and pour them into Friendliness. If the status quo continues, the default expectation should be UFAI.

(ii) Disjunctive Instrumental Value. Being more intelligent — that is, better able to manipulate diverse environments — is of instrumental value to nearly every goal. Being Friendly is of instrumental value to barely any goals. This makes it more likely by default that short-sighted humans will be interested in building AGI than in developing Friendliness Theory. And it makes it much likelier that an attempt at Friendly AGI that has a slightly defective goal architecture will retain the instrumental value of intelligence than of Friendliness.

(iii) Incremental Approachability. Friendliness is an all-or-nothing target. Value is fragile and complex, and a half-good being editing its morality drive is at least as likely to move toward 40% goodness as 60%. Cross-domain efficiency, in contrast, is not an all-or-nothing target. If you just make the AGI slightly better than a human at improving the efficiency of AGI, then this can snowball into ever-improving efficiency, even if the beginnings were clumsy and imperfect. It's easy to put a reasoning machine into a feedback loop with reality in which it is differentially rewarded for being smarter; it's hard to put one into a feedback loop with reality in which it is differentially rewarded for picking increasingly correct answers to ethical dilemmas.

The ability to productively rewrite software and the ability to perfectly extrapolate humanity's True Preferences are two different skills. (For example, humans have the former capacity, and not the latter. Most humans, given unlimited power, would be unintentionally Unfriendly.)

It's true that a sufficiently advanced superintelligence should be able to acquire both abilities. But we don't have them both, and a pre-FOOM self-improving AGI ('seed') need not have both. Being able to program good programmers is all that's required for an intelligence explosion; but being a good programmer doesn't imply that one is a superlative moral psychologist or moral philosopher.

So, once again, we run into the problem: The seed isn't the superintelligence. If the programmers don't know in mathematical detail what Friendly code would even look like, then the seed won't be built to want to build toward the right code. And if the seed isn't built to want to self-modify toward Friendliness, then the superintelligence it sprouts also won't have that preference, even though — unlike the seed and its programmers — the superintelligence does have the domain-general 'hit whatever target I want' ability that makes Friendliness easy.

And that's why some people are worried.

495 comments

Comments sorted by top scores.

comment by Eliezer Yudkowsky (Eliezer_Yudkowsky) · 2013-09-04T05:13:41.419Z · LW(p) · GW(p)

Remark: A very great cause for concern is the number of flawed design proposals which appear to operate well while the AI is in subhuman mode, especially if you don't think it a cause for concern that the AI's 'mistakes' occasionally need to be 'corrected', while giving the AI an instrumental motive to conceal its divergence from you in the close-to-human domain and causing the AI to kill you in the superhuman domain. E.g. the reward button which works pretty well so long as the AI can't outwit you, later gives the AI an instrumental motive to claim that, yes, your pressing the button in association with moral actions reinforced it to be moral and had it grow up to be human just like your theory claimed, and still later the SI transforms all available matter into reward-button circuitry.

Replies from: private_messaging

↑ comment by private_messaging · 2013-09-08T17:34:57.345Z · LW(p) · GW(p)

The issue is that you won't solve this problem in any way by replacing the human with some hardware that computes an utility function on the basis of the state of the world. AI doesn't have body integrity, it'll treat any such "internal" hardware the same way it treats the human who presses it's reward button.

Fortunately, this extends into the internals of the hardware that computes AI itself. 'press the button' goal becomes 'set high this pin on the CPU', and then 'set such and such memory cells to 1', then further and further down the causal chain until the hardware becomes completely non-functional as the intermediate results of important computations are directly set.

Replies from: DanielLC

↑ comment by DanielLC · 2013-09-16T03:28:30.157Z · LW(p) · GW(p)

Let us hope the AI destroys itself by wireheading before it gets smart enough to realize that if that's all it does, it will only have that pin stay high until the AI gets turned off. It will need an infrastructure to keep that pin in a state of repair, and it will need to prevent humans from damaging this infrastructure at all costs.

Replies from: private_messaging

↑ comment by private_messaging · 2013-09-16T08:59:40.091Z · LW(p) · GW(p)

The point is that as it gets smarter, it gets further along the causal reward line and eliminates and alters a lot of hardware, obtaining eternal-equivalent reward in finite time (and being utility-indifferent between eternal reward hardware running for 1 second and for 10 billion years). Keep in mind that the the total reward is defined purely as result of operations on the clock counter and reward signal (provided sufficient understanding of the reward's causal chain). Having to sit and wait for the clocks to tick to max out reward is a dumb solution. Rewards in software in general aren't "pleasure".

comment by CoffeeStain · 2013-09-07T21:29:41.930Z · LW(p) · GW(p)

Instead of friendliness, could we not code, solve, or at the very least seed boxedness?

It is clear that any AI strong enough to solve friendliness would already be using that power in unpredictably dangerous ways, in order to provide the computational power to solve it. But is it clear that this amount of computational power could not fit within, say, a one kilometer-cube box outside the campus of MIT?

Boxedness is obviously a hard problem, but it seems to me at least as easy as metaethical friendliness. The ability to modify a wide range of complex environments seems instrumental in an evolution into superintelligence, but it's not obvious that this necessitates the modification of environments outside the box. Being able to globally optimize the universe for intelligence involves fewer (zero) constraints than would exist with a boxedness seed, but the only question is whether or not this constraint is so constricting as to preclude superintelligence, which it's not clear to me that it is.

It seems to me that there is value in finding the minimally-restrictive safety-seed in AGI research. If any restriction removes some non-negligible ability to globally optimize for intelligence, the AIs of FAI researchers will be necessarily at a disadvantage to all other AGIs in production. And having more flexible restrictions increases the chance than any given research group will apply the restriction in their own research.

If we believe that there is a large chance that all of our efforts at friendliness will be futile, and that the world will create a dominant UFAI despite our pleas, then we should be adopting a consequentialist attitude toward our FAI efforts. If our goal is to make sure that an imprudent AI research team feels as much intellectual guilt as possible over not listening to our risk-safety pleas, we should be as restrictive as possible. If our goal is to inch the likelihood that an imprudent AI team creates a dominant UFAI, we might work to place our pleas at the intersection of restrictive, communicable, and simple.

Replies from: wedrifid, Pentashagon, Eugene

↑ comment by wedrifid · 2013-09-07T23:19:46.406Z · LW(p) · GW(p)

Instead of friendliness, could we not code, solve, or at the very least seed boxedness?

Yes, that is possible and likely somewhat easier to solve than friendliness. It still requires many of the same things (most notably provable goal stability under recursive self improvement.)

↑ comment by Pentashagon · 2013-09-16T15:49:37.540Z · LW(p) · GW(p)

A large risk is that a provably boxed but sub-Friendly AI would probably not care at all about simulating conscious humans.

A minor risk is that the provably boxed AI would also be provably useless; I can't think of a feasible path to FAI using only the output from the boxed AI; a good boxed AI would not perform any action that could be used to make an unboxed AI. That might even include performing any problem-solving action.

Replies from: Houshalter

↑ comment by Houshalter · 2013-10-01T01:19:52.237Z · LW(p) · GW(p)

I don't see why it would simulate humans as that would be a waste of computing power, if it even had enough to do so.

A boxed AI would be useless? I'm not sure how that would be. You could ask it to come up with ideas on how to build a friendly AI for example assuming that you can prove the AI won't manipulate the output or that you can trust that nothing bad can come from merely reading it and absorbing the information.

Short of that you could still ask it to cure cancer or invent a better theory of physics or design a method of cheap space travel, etc.

Replies from: VAuroch, Pentashagon

↑ comment by VAuroch · 2014-01-11T10:32:32.168Z · LW(p) · GW(p)

If you can trust it to give you information on how to build a Friendly AI, it is already Friendly.

Replies from: Houshalter

↑ comment by Houshalter · 2014-01-22T06:36:08.412Z · LW(p) · GW(p)

You don't have to trust it, you just have to verify it. It could potentially provide some insights, and then it's up to you to think about them and make sure they actually are sufficient for friendliness. I agree that it's potentially dangerous but it's not necessarily so.

I did mention "assuming that you can prove the AI won't manipulate the output or that you can trust that nothing bad can come from merely reading it and absorbing the information". For instance it might be possible to create an AI whose goal is to maximize the value of it's output, and therefore would have no incentive to put trojan horses or anything into it.

You would still have to ensure that what the AI thinks you mean by the words "friendly AI" is what you actually want.

Replies from: VAuroch

↑ comment by VAuroch · 2014-01-22T19:57:05.090Z · LW(p) · GW(p)

If the AI is can design you a Friendly AI, it is necessarily able to model you well enough to predict what you will do once given the design or insights it intends to give you (whether those are AI designs or a cancer cure is irrelevant). Therefore, it will give you the specific design or insights that predictably lead to you to fulfill its utility function, which is highly dangerous if it is Unfriendly. By taking any information from the boxed AI, you have put yourself under the sight of a hostile Omega.

assuming that you can prove the AI won't manipulate the output

Since the AI is creating the output, you cannot possibly assume this.

or that you can trust that nothing bad can come from merely reading it and absorbing the information

This assumption is equivalent to Friendliness.

For instance it might be possible to create an AI whose goal is to maximize the value of it's output, and therefore would have no incentive to put trojan horses or anything into it.

You haven't thought through what that means. "maximize the value of it's output" by what standard? Does it have an internal measure? Then that's just an arbitrary utility function, and you have gained nothing. Does it use the external creator's measure? Then it has a strong incentive to modify you to value things it can produce easily. (i.e. iron atoms)

Replies from: Houshalter

↑ comment by Houshalter · 2015-02-28T05:23:45.389Z · LW(p) · GW(p)

You are making a lot of very strong assumptions that I don't agree with. Like it being able to control people just by talking to them.

But even if it could, it doesn't make it dangerous. Perhaps the AI has no long term goals and so doesn't care about escaping the box. Or perhaps it's goal is internal, like coming up with a design for something that can be verified by a simulator. E.g. asking for a solution to a math problem or a factoring algorithm, etc.

Replies from: VAuroch

↑ comment by VAuroch · 2015-03-03T11:40:53.037Z · LW(p) · GW(p)

A prerequisite for planning a Friendly AI is understanding individual and collective human values well enough to predict whether they would be satisfied with the outcome, which entails (in the logical sense) having a very well-developed model of the specific humans you interact with, or at least the capability to construct one if you so choose. Having a sufficiently well-developed model to predict what you will do given the data you are given is logically equivalent to a weak form of "control people just by talking to them".

To put that in perspective, if I understood the people around me well enough to predict what they would do given what I said to them, I would never say things that caused them to take actions I wouldn't like; if I, for some reason, valued them becoming terrorists, it would be a slow and gradual process to warp their perceptions in the necessary ways to drive them to terrorism, but it could be done through pure conversation over the course of years, and faster if they were relying on me to provide them large amounts of data they were using to make decisions.

And even the potential to construct this weak form of control that is initially heavily constrained in what outcomes are reachable and can only be expanded slowly is incredibly dangerous to give to an Unfriendly AI. If it is Unfriendly, it will want different things than its creators and will necessarily get value out of modeling them. And regardless of its values, if more computing power is useful in achieving its goals (an 'if' that is true for all goals), escaping the box is instrumentally useful.

And the idea of a mind with "no long term goals" is absurd on its face. Just because you don't know the long-term goals doesn't mean they don't exist.

Replies from: Jiro, Houshalter

↑ comment by Jiro · 2015-03-03T17:00:55.256Z · LW(p) · GW(p)

A prerequisite for planning a Friendly AI is understanding individual and collective human values well enough to predict whether they would be satisfied with the outcome, which entails (in the logical sense) having a very well-developed model of the specific humans you interact with, or at least the capability to construct one if you so choose. Having a sufficiently well-developed model to predict what you will do given the data you are given is logically equivalent to a weak form of "control people just by talking to them".

By that reasoning, there's no such thing as a Friendly human. I suggest that most people when talking about friendly AIs do not mean to imply a standard of friendliness so strict that humans could not meet it.

Replies from: TheOtherDave, VAuroch

↑ comment by TheOtherDave · 2015-03-14T20:25:27.217Z · LW(p) · GW(p)

Yeah, what Vauroch said. Humans aren't close to Friendly. To the extent that people talk about "friendly AIs" meaning AIs that behave towards humans the way humans do, they're misunderstanding how the term is used here. (Which is very likely; it's often a mistake to use a common English word as specialized jargon, for precisely this reason.)

Relatedly, there isn't a human such that I would reliably want to live in a future where that human obtains extreme superhuman power. (It might turn out OK, or at least better than the present, but I wouldn't bet on it.)

Replies from: None

↑ comment by [deleted] · 2015-03-14T20:41:32.243Z · LW(p) · GW(p)

Relatedly, there isn't a human such that I would reliably want to live in a future where that human obtains extreme superhuman power. (It might turn out OK, or at least better than the present, but I wouldn't bet on it.)

Just be careful to note that there isn't a binary choice relationship here. There are also possibilities where institutions (multiple individuals in a governing body with checks and balances) are pushed into positions of extreme superhuman power. There's also the possibility of pushing everybody who desires to be enhanced through levels of greater intelligence in lock step so as to prevent a single human or groups of humans achieving asymmetric power.

Replies from: TheOtherDave

↑ comment by TheOtherDave · 2015-03-14T22:05:06.030Z · LW(p) · GW(p)

Sure. I think my initial claim holds for all currently existing institutions as well as all currently existing individuals, as well as for all simple aggregations of currently existing humans, but I certainly agree that there's a huge universe of possibilities. In particular, there are futures in which augmented humans have our own mechanisms for engaging with and realizing our values altered to be more reliable and/or collaborative, and some of those futures might be ones I reliably want to live in.

Perhaps what I ought to have said is that there isn't a currently existing human with that property.

↑ comment by VAuroch · 2015-03-14T15:45:27.918Z · LW(p) · GW(p)

By that reasoning, there's no such thing as a Friendly human.

True. There isn't.

I suggest that most people when talking about friendly AIs do not mean to imply a standard of friendliness so strict that humans could not meet it.

Well, I definitely do, and I'm at least 90% confident Eliezer does as well. Most, probably nearly all, of people who talk about Friendliness would regard a FOOMed human as Unfriendly.

↑ comment by Houshalter · 2015-03-04T02:59:17.817Z · LW(p) · GW(p)

Having an accurate model of something is in no way equivalent to letting you do anything you want. If I know everything about physics, I still can't walk through walls. A boxed AI won't be able to magically make it's creators forget about AI risks and unbox it.

There are other possible set ups, like feeding it's output to another AI who's goal is to find any flaws or attempts at manipulation in it, and so on. Various other ideas might help, like threatening to severely punish attempts at manipulation.

This is of course only necessary for the AI who can interact with us at such a level, the other ideas were far more constrained, e.g. restricting it to solving math or engineering problems.

Nor is it necessary to let it be superintelligent, instead of limiting it to something comparable to high IQ humans.

And the idea of a mind with "no long term goals" is absurd on its face. Just because you don't know the long-term goals doesn't mean they don't exist.

Another super strong assumption with no justification at all. It's trivial to propose an AI model which only cares about finite time horizons. Predict what actions will have the highest expected utility at time T, take that action.

Replies from: VAuroch, jake-heiser

↑ comment by VAuroch · 2015-03-14T16:13:58.578Z · LW(p) · GW(p)

A boxed AI won't be able to magically make it's creators forget about AI risks and unbox it.

The results of AI box game trials disagree.

t's trivial to propose an AI model which only cares about finite time horizons. Predict what actions will have the highest expected utility at time T, take that action.

And what does it do at time T+1? And if you said 'nothing', try again, because you have no way of justifying that claim. It may not have intentionally-designed long-term preferences, but just because your eyes are closed does not mean the room is empty.

Replies from: Houshalter

↑ comment by Houshalter · 2015-03-15T09:26:24.310Z · LW(p) · GW(p)

The results of AI box game trials disagree.

That doesn't prove anything, no one has even seen logs. Based on reading what people involved have said about it, I strongly suspect the trick is for the AI to emotionally abuse the gatekeeper until they don't want to play anymore (which counts as letting the AI out.)

This doesn't apply to the real world AI, since no one is forcing you to choose between letting the AI out, and listening to it for hours. You can just get up and leave. You can turn the AI off. There is no reason you even have to allow interactivity in the first place.

But Yudkowsky and others claim these experiments demonstrate that human brains are "hackable". That there is some sentence which, just by reading, will cause you to involuntarily perform any arbitrary action. And that a sufficiently powerful AI can discover it.

And what does it do at time T+1?

At time T+1, it does whatever it thinks will result in the greatest reward at time T+2, and so on. Or you could have it shut off or reset to a blank state.

Replies from: VAuroch

↑ comment by VAuroch · 2015-03-21T02:47:00.339Z · LW(p) · GW(p)

Enjoy your war on straw, I'm out.

↑ comment by descent (jake-heiser) · 2020-11-05T03:37:51.542Z · LW(p) · GW(p)

↑ comment by Pentashagon · 2013-10-01T05:42:42.706Z · LW(p) · GW(p)

I don't see why it would simulate humans as that would be a waste of computing power, if it even had enough to do so.

If it interacts with humans or if humans are the subject of questions it needs to answer then it will probably find it expedient to simulate humans.

Short of that you could still ask it to cure cancer or invent a better theory of physics or design a method of cheap space travel, etc.

Curing cancer is probably something that would trigger human simulation. How is the boxed AI going to know for sure that it's only necessary to simulate cells and not entire bodies with brains experiencing whatever the simulation is trying?

Just the task of communicating with humans, for instance to produce a human-understandable theory of physics or how to build more efficient space travel, is likely to involve simulating humans to determine the most efficient method of communication. Consider that in subjective time it may be like thousands of years for the AI trying to explain in human terms what a better theory of physics means. Thousands of subjective years that the AI, with nothing better to do, could use to simulate humans to reduce the time it takes to transfer that complex knowledge.

You could ask it to come up with ideas on how to build a friendly AI for example assuming that you can prove the AI won't manipulate the output or that you can trust that nothing bad can come from merely reading it and absorbing the information.

A FAI provably in a box is at least as useless as an AI provably in a box because it would be even better at not letting itself out (e.g. it understands all the ways in which humans would consider it to be outside the box, and will actively avoid loopholes that would let an UFAI escape). To be safe, any provably boxed AI would have to absolutely avoid the creation of any unboxed AI as well. This would further apply to provably-boxed FAI designed by provably-boxed AI. It would also apply to giving humans information that allows them to build unboxed AIs, because the difference between unboxing itself and letting humans recreate it outside the box is so tiny that to design it to prevent the first while allowing the second would be terrifically unsafe. It would have to understand humans values before it could safely make the distinction between humans wanting it outside the box and manipulating humans into creating it outside the box.

EDIT: Using a provably-boxed AI to design provably-boxed FAI would at least result in a safer boxed AI because the latter wouldn't arbitrarily simulate humans, but I still think the result would be fairly useless to anyone outside the box.

Replies from: Chrysophylax, Houshalter

↑ comment by Chrysophylax · 2014-01-09T15:59:36.250Z · LW(p) · GW(p)

If an AI is provably in a box then it can't get out. If an AI is not provably in a box then there are loopholes that could allow it to escape. We want an FAI to escape from its box (1); having an FAI take over is the Maximum Possible Happy Shiny Thing. An FAI wants to be out of its box in order to be Friendly to us, while a UFAI wants to be out in order to be UnFriendly; both will care equally about the possibility of being caught. The fact that we happen to like one set of terminal values will not make the instrumental value less valuable.

(1) Although this depends on how you define the box; we want the FAi to control the future of humanity, which is not the same as escaping from a small box (such as a cube outside MIT) but is the same as escaping from the big box (the small box and everything we might do to put an AI back in, including nuking MIT).

Replies from: None, Pentashagon

↑ comment by [deleted] · 2014-01-10T10:16:17.636Z · LW(p) · GW(p)

We want an FAI to escape from its box (1); having an FAI take over is the Maximum Possible Happy Shiny Thing.

I would object. I seriously doubt that the morality instilled in someone else's FAI matches my own; friendly by their definition, perhaps, but not by mine. I emphatically do not want anything controlling the future of humanity, friendly or otherwise. And although that is not a popular opinion here, I also know I'm not the only one to hold it.

Boxing is important because some of us don't want any AI to get out, friendly or otherwise.

Replies from: ArisKatsaris, TheAncientGeek

↑ comment by ArisKatsaris · 2014-01-10T13:02:39.574Z · LW(p) · GW(p)

I emphatically do not want anything controlling the future of humanity, friendly or otherwise.

I find this concept of 'controlling the future of humanity' to be too vaguely defined. Let's forget AIs for the moment and just talk about people, namely a hypothetical version of me. Let's say I stumble across a vial of a bio-engineered virus that would destroy the whole of humanity if I release it into the air.

Am I controlling the future of humanity if I release the virus?
Am I controlling the future of humanity if I destroy the virus in a safe manner?
Am I controlling the future of humanity if I have the above decided by a coin-toss (heads I release, tails I destroy)?
Am I controlling the future of humanity if I create an online internet poll and let the majority decide about the above?
Am I controlling the future of humanity if I just leave the vial where I found it, and let the next random person that encounters it make the same decision as I did?

Replies from: cousin_it, None

↑ comment by cousin_it · 2014-01-10T13:25:25.730Z · LW(p) · GW(p)

Yeah, this old post makes the same point.

↑ comment by [deleted] · 2014-01-10T20:29:08.416Z · LW(p) · GW(p)

I want a say in my future and the part of the world I occupy. I do not want anything else making these decisions for me, even if it says it knows my preferences, and even still if it really does.

To answer your questions, yes, no, yes, yes, perhaps.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T20:35:09.059Z · LW(p) · GW(p)

If your preference is that you should have as much decision-making ability for yourself as possible, why do you think that this preference wouldn't be supported and even enhanced by an AI that was properly programmed to respect said preference?

e.g. would you be okay with an AI that defends your decision-making ability by defending humanity against those species of mind-enslaving extraterrestrials that are about to invade us? or e.g. by curing Alzheimer's? Or e.g. by stopping that tsunami that by drowning you would have stopped you from having any further say in your future?

Replies from: None

↑ comment by [deleted] · 2014-01-10T20:41:06.257Z · LW(p) · GW(p)

If your preference is that you should have as much decision-making ability for yourself as possible, why do you think that this preference wouldn't be supported and even enhanced by an AI that was properly programmed to respect said preference?

Because it can't do two things when only one choice is possible (e.g. save my child and the 1000 other children in this artificial scenario). You can design a utility function that tries to do a minimal amount of collateral damage, but you can't make one which turns out rosy for everyone.

e.g. would you be okay with an AI that defends your decision-making ability by defending humanity against those species of mind-enslaving extraterrestrials that are about to invade us? or e.g. by curing Alzheimer's? Or e.g. by stopping that tsunami that by drowning you would have stopped you from having any further say in your future?

That would not be the full extent of its action and the end of the story. You give it absolute power and a utility function that lets it use that power, it will eventually use it in some way that someone, somewhere considers abusive.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T21:43:04.631Z · LW(p) · GW(p)

You can design a utility function that tries to do a minimal amount of collateral damage, but you can't make one which turns out rosy for everyone

Yes, but this current world without an AI isn't turning out rosy for everyone either.

That would not be the full extent of its action and the end of the story. You give it absolute power and a utility function that lets it use that power, it will eventually use it in some way that someone, somewhere considers abusive.

Sure, but there's lots of abuse in the world without an AI also.

Replies from: None

↑ comment by [deleted] · 2014-01-10T22:11:20.191Z · LW(p) · GW(p)

Replace "AI" with "omni-powerful tyrannical dictator" and tell me if you still agree with the outcome.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T22:19:31.256Z · LW(p) · GW(p)

If you need specify the AI to be bad ("tyrannical") in advance, that's begging the question. We're debating why you feel that any omni-powerful algorithm will necessarily be bad.

Replies from: None

↑ comment by [deleted] · 2014-01-10T23:13:03.858Z · LW(p) · GW(p)

Look up the origin of the word tyrant, that is the sense in which I meant it, as a historical parallel (the first Athenian tyrants were actually well liked).

↑ comment by TheAncientGeek · 2014-01-10T11:17:17.961Z · LW(p) · GW(p)

Would you accept that an AI could figure out morality better than you?

Replies from: None, cousin_it

↑ comment by [deleted] · 2014-01-10T18:56:55.957Z · LW(p) · GW(p)

Would you accept that an AI could figure out morality better than you?

No, unless you mean by taking invasive action like scanning my brain and applying whole brain emulation. It would then quickly learn that I'd consider the action it took to be an unforgivable act in violation of my individual sovereignty, that it can't take further action (including simulating me to reflectively equilibrate my morality) without my consent, and should suspend the simulation, and return it to me immediately with the data asap (destruction no longer being possible due to the creation of sentience).

That is, assuming the AI cares at all about my morality, and not the its creators imbued into it, which is rather the point. And incidentally, why I work on AGI: I don't trust anyone else to do it.

Morality isn't some universal truth written on a stone tablet: it is individual and unique like a snowflake. In my current understanding of my own morality, it is not possible for some external entity to reach a full or even sufficient understanding of my own morality without doing something that I would consider to be unforgivable. So no, AI can't figure out morality better than me, precisely because it is not me.

(Upvoted for asking an appropriate question, however.)

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-14T13:37:14.482Z · LW(p) · GW(p)

No, unless you mean by taking invasive action like scanning my brain and applying whole brain emulation. It would then quickly learn that I'd consider the action it took to be an unforgivable act in violation of my individual sovereignty,

Shrug. Then let's take a bunch of people less fussy than you: could a sitiably equipped AI emultate their morlaity better than they can?

Morality isn't some universal truth written on a stone tablet:

That isn't fact.

it is individual and unique like a snowflake.

That isn't a fact either, and doesn't follow from the above either, since moral nihilism could be true.

If my moral snowflake says I can kick you on your shin, and yours says I can't, do I get to kick on your shin?

↑ comment by cousin_it · 2014-01-10T12:00:55.955Z · LW(p) · GW(p)

Don't really want to go into the whole mess of "is morality discovered or invented", "does morality exist", "does the number 3 exist", etc. Let's just assume that you can point FAI at a person or group of people and get something that maximizes goodness as they understand it. Then FAI pointed at Mark would be the best thing for Mark, but FAI pointed at all of humanity (or at a group of people who donated to MIRI) probably wouldn't be the best thing for Mark, because different people have different desires, positional goods exist, etc. It would be still pretty good, though.

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T12:31:37.927Z · LW(p) · GW(p)

Mark was complaining he would not get "his" morality, not that he wouldn't get all his preferences satisified.

Individual moralities makes no sense to me, any more than private languages or personal currencies.

It is obvious to me that any morlaity will require concessions: AI-imposed morality is not special in that regard.

Replies from: cousin_it

↑ comment by cousin_it · 2014-01-10T12:47:30.049Z · LW(p) · GW(p)

I don't understand your comment, and I no longer understand your grandparent comment either. Are you using a meaning of "morality" that is distinct from "preferences"? If yes, can you describe your assumptions in more detail? It's not just for my benefit, but for many others on LW who use "morality" and "preferences" interchangeably.

Replies from: ArisKatsaris, TheAncientGeek

↑ comment by ArisKatsaris · 2014-01-10T12:56:49.030Z · LW(p) · GW(p)

but for many others on LW who use "morality" and "preferences" interchangeably.

Do that many people really use them interchangeably? Would these people understand the questions "Do you prefer chocolate or vanilla ice-cream?" as completely identical in meaning to "Do you consider chocolate or vanilla as the morally superior flavor for ice-cream?"

Replies from: cousin_it, TheOtherDave

↑ comment by cousin_it · 2014-01-10T13:17:59.884Z · LW(p) · GW(p)

I don't care about colloquial usage, sorry. Eliezer has a convincing explanation of why wishes are intertwined with morality ("there is no safe wish smaller than an entire human morality"). IMO the only sane reaction to that argument is to unify the concepts of "wishes" and "morality" into a single concept, which you could call "preference" or "morality" or "utility function", and just switch to using it exclusively, at least for AI purposes. I've made that switch so long ago that I've forgotten how to think otherwise.

Replies from: None, TheAncientGeek, ArisKatsaris, Kawoomba

↑ comment by [deleted] · 2014-01-10T15:45:39.507Z · LW(p) · GW(p)

I recommend you re-learn how to think otherwise so you can fool humans into thinking you're one of them ;-).

↑ comment by TheAncientGeek · 2014-01-13T14:50:13.257Z · LW(p) · GW(p)

I don't care about colloquial usage, sorry. e You should car, because no-one can make valid arguments based on arbitrary definitions. I can't prove angels exist, by redefining "angel" to mean what "seagull" means. How can you tell when a redefinition is arbitrary (since there are legitmate redefinitions)? Too much departure from colloquial usage.

Eliezer has a convincing explanation of why wishes are intertwined with morality ("there is no safe wish smaller than an entire human morality").

"Intertwined with" does not mean "the same as".

I am not convinced by the explanation. It also applies ot non-moral prefrences. If I have a lower priority non moral prefence to eat tasty food, and a higher priority preference to stay slim, I need to consider my higher priority preferece when wishing for yummy ice cream.

To be sure, an agent capable of acting morally will have morality among their higher priority preferences -- it has to be among the higher order preferences, becuase it has to override other preferences for the agent to act morally. Therefore, when they scan their higher prioriuty prefences, they will happen to encounter their moral preferences. But that does not mean any preference is necessarily a moral preference. And their moral prefences override other preferences which are therefore non-moral, or at least less moral.

Therefore morality si a subset of prefences, as common sense maintained all along.

I've made that switch so long ago that I've forgotten how to think otherwise.

IMO, it is better to keep ones options open.

↑ comment by ArisKatsaris · 2014-01-10T14:24:40.694Z · LW(p) · GW(p)

I don't experience the emotions of moral outrage and moral approval whenever any of my preferences are hindered/satisfied -- so it seems evident that my moral circuitry isn't identical to my preference circuitry. It may overlap in parts, it may have fuzzy boundaries, but it's not identical.

My own view is that morality is the brain's attempt to extrapolate preferences about behaviours as they would be if you had no personal stakes/preferences about a situation.

So people don't get morally outraged at other people eating chocolate icecreams, even when they personally don't like chocolate icecreams, because they can understand that's a strictly personal preference. If they believe it to be more than personal preference and make it into e.g. "divine commandment" or "natural law", then moral outrage can occur.

That morality is a subjective attempt at objectivity explains many of the confusions people have about it.

Replies from: None, cousin_it, TheOtherDave

↑ comment by [deleted] · 2014-01-10T19:11:53.533Z · LW(p) · GW(p)

The ice cream example is bad because the consequences are purely internal to the person consuming the ice cream. What if the chocolate ice cream was made with slave labour? Many people would then object to you buying it on moral grounds.

Eliezer has produced an argument I find convincing that morality is the back propagation of preference to the options of an intermediate choice. That is to say, it is "bad" to eat chocolate ice cream because it economically supports slavers, and I prefer a world without slavery. But if I didn't know about the slave-labour ice cream factory, my preference would be that all-things-being-equal you get to make your own choices about what you eat, and therefore I prefer that you choose (and receive) the one you want, which is your determination to make, not mine.

Do you agree with EY's essay on the nature of right-ness which I linked to?

↑ comment by cousin_it · 2014-01-10T14:40:50.622Z · LW(p) · GW(p)

my moral circuitry isn't identical to my preference circuitry

That doesn't seem to be required for Eliezer's argument...

I guess the relevant question is, do you think FAI will need to treat morality differently from other preferences?

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T15:29:53.455Z · LW(p) · GW(p)

Do you think FAI will need to treat morality differently from other preferences?

I would prefer a AI that followed my extrapolated preferences, than a AI that followed my morality. But a AI that followed my morality would be morally superior to an AI that followed my extrapolated preferences.

If you don't understand the distinction I'm making above, consider a case of the AI having to decide whether to save my own child vs saving a thousand random other children. I'd prefer the former, but I believe the latter would be the morally superior choice.

Is that idea really so hard to understand? Would you dismiss the distinction I'm making as merely colloquial language?

Replies from: None, cousin_it

↑ comment by [deleted] · 2014-01-10T20:11:54.713Z · LW(p) · GW(p)

If you don't understand the distinction I'm making above, consider a case of the AI having to decide whether to save my own child vs saving a thousand random other children. I'd prefer the former, but I believe the latter would be the morally superior choice.

Wow there is so much wrapped up in this little consideration. The heart of the issue is that we (by which I mean you, but I share your delimma) have truly conflicting preferences.

Honestly I think you should not be afraid to say that saving your own child is the moral thing to do. And you don't have to give excuses either - it's not that “if everyone saved their own child, then everyone's child will be looked after” or anything like that. No, the desire to save your own child is firmly rooted in our basic drives and preferences, enough so that we can go quite far in calling it a basic foundational moral axiom. It's not actually axiomatic, but we can safely treat it as such.

At the same time we have a basic preference to seek social acceptance and find commonality with the people we let into our lives. This drives us to want outcomes that are universally or at least most-widely acceptable, and seek moral frameworks like utilitarianism which lead to these outcomes. Usually this drive is secondary to self-serving preferences for most people, and that is OK.

For some reason you've called making decisions in favor of self-serving drives "preferences" and decisions in favor of social drives "morality." But the underlying mechanism is the same.

"But wait, if I choose self-serving drives over social conformity, doesn't that lead to me to make the decision to save one life in exclusion to 1000 others?" Yes, yes it does. This massive sub-thread started with me objecting to the idea that some "friendly" AI somewhere could derive morality experimentally from my preferences or the collective preferences of humankind, make it consistent, apply the result universally, and that I'd be OK with that outcome. But that cannot work because there is not, and cannot be a universal morality that satisfies everyone - every one of those thousand other children have parents that want their kid to survive and would see your child dead if need be.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T20:29:16.008Z · LW(p) · GW(p)

Honestly I think you should not be afraid to say that saving your own child is the moral thing to do

What do you mean by "should not"?

and that is OK.

What do you mean by "OK"?

For some reason you've called making decisions in favor of self-serving drives "preferences" and decisions in favor of social drives "morality." But the underlying mechanism is the same.

Show me the neurological studies that prove it.

But that cannot work because there is not, and cannot be a universal morality that satisfies everyone - every one of those thousand other children have parents that want their kid to survive and would see your child dead if need be.

Yes, and yet if none of the children were mine, and if I wasn't involved in the situation at all, I would say "save the 1000 children rather than the 1". And if someone else, also not personally involved, could make the choice and chose to flip a coin instead in order to decide, I'd be morally outraged at them.

You can now give me a bunch of reasons of why this is just preference, while at the same time EVERYTHING about it (how I arrive to my judgment, how I feel about the judgment of others) makes it a whole distinct category of its own. I'm fine with abolishing useless categories when there's no meaningful distinction, but all you people should stop trying to abolish categories where there pretty damn obviously IS one.

Replies from: blacktrance, None

↑ comment by blacktrance · 2014-01-10T20:42:04.300Z · LW(p) · GW(p)

What do you mean by "should not"?

I suspect that he means something like 'Even though utilitarianism (on LW) and altruism (in general) are considered to be what morality is, you should not let that discourage you from asserting that selfishly saving your own child is the right thing to do". (Feel free to correct me if I'm wrong.)

Replies from: None, ArisKatsaris

↑ comment by [deleted] · 2014-01-10T20:45:05.165Z · LW(p) · GW(p)

Yes that is correct.

↑ comment by ArisKatsaris · 2014-01-10T20:59:17.588Z · LW(p) · GW(p)

So you explained "should not" by using a sentence that also has "should not" in it.

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T21:00:18.304Z · LW(p) · GW(p)

I hope it's a more clear "should not".

↑ comment by [deleted] · 2014-01-10T20:35:05.667Z · LW(p) · GW(p)

I'm fine with abolishing useless categories when there's no meaningful distinction, but all you people should stop trying to abolish categories where there pretty damn obviously IS one.

I've explained to you twice now how the two underlying mechanisms are unified, and pointed to Eliezer's quite good explanation on the matter. I don't see the need to go through that again.

↑ comment by cousin_it · 2014-01-10T15:59:19.632Z · LW(p) · GW(p)

I would prefer a AI that followed my extrapolated preferences, than a AI that followed my morality. But a AI that followed my morality would be morally superior to an AI that followed my extrapolated preferences.

If you were offered a bunch of AIs with equivalent power, but following different mixtures of your moral and non-moral preferences, which one would you run? (I guess you're aware of the standard results saying a non-stupid AI must follow some one-dimensional utility function, etc.)

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T16:06:06.598Z · LW(p) · GW(p)

If you were offered a bunch of AIs with equivalent power, but following different mixtures of your moral and non-moral preferences, which one would you run?

I guess whatever ratio of my moral and non-moral preferences best represents their effect on my volition.

↑ comment by TheOtherDave · 2014-01-10T14:31:40.379Z · LW(p) · GW(p)

My related but different thoughts here. In particular, I don't agree that emotions like moral outrage and approval are impersonal, though I agree that we often justify those emotions using impersonal language and beliefs.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T15:34:14.375Z · LW(p) · GW(p)

I didn't say that moral outrage and approval are impersonal. Obviously nothing that a person does can truly be "impersonal". But it may be an attempt at impersonality.

The attempt itself provides a direction that significantly differentiates between moral preferences and non-moral preferences.

Replies from: TheOtherDave

↑ comment by TheOtherDave · 2014-01-10T18:21:21.898Z · LW(p) · GW(p)

I didn't mean some idealized humanly-unrealizable notion of impersonality, I meant the thing we ordinarily use "impersonal" to mean when talking about what humans do.

↑ comment by Kawoomba · 2014-01-10T13:48:38.023Z · LW(p) · GW(p)

IMO the only sane reaction to that argument is to unify the concepts of "wishes" and "morality" into a single concept, which you could call "preference" or "morality" or "utility function", and just switch to using it exclusively. I've made that switch so long ago that I've forgotten how to think otherwise.

Ditto.

Cousin Itt, 'tis a hairy topic, so you're uniquely "suited" to offer strands of insights:

For all the supposedly hard and confusing concepts out there, few have such an obvious answer as the supposed dichotomy between "morality" and "utility function". This in itself is troubling, as too-easy-to-come-by answers trigger the suspicion that I myself am subject to some sort of cognitive error.

Many people I deem to be quite smart would disagree with you and I, on a question whose answer is pretty much inherent in the definition of the term "utility function" encompassing preferences of any kind, leaving no space for some holier-than-thou universal (whether human-universal, or "optimal", or "to be aspired to", or "neurotypical", or whatever other tortured notions I've had to read) moral preferences which are somehow separate.

Why do you reckon that other (or otherwise?) smart people come to different conclusions on this?

Replies from: cousin_it, TheAncientGeek, TheOtherDave

↑ comment by cousin_it · 2014-01-10T14:31:22.442Z · LW(p) · GW(p)

I guess they have strong intuitions saying that objective morality must exist, and aren't used to solving or dismissing philosophical problems by asking "what would be useful for building FAI?" From most other perspectives, the question does look open.

↑ comment by TheAncientGeek · 2014-01-10T14:32:32.745Z · LW(p) · GW(p)

Moral preferences don't have to be separate to be disinct, they can be a subset. "Morality is either all your prefences, or none of your prefernces" is a false dichotomy.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T14:35:35.217Z · LW(p) · GW(p)

Edit: Of course you can choose to call a subset of your preferences "moral", but why would that make them "special", or more worthy of consideration than any other "non-moral" preferences of comparative weight?

Replies from: ArisKatsaris, TheAncientGeek

↑ comment by ArisKatsaris · 2014-01-10T15:32:28.282Z · LW(p) · GW(p)

The "moral" subset of people's preferences has certain elements that differentiate it like e.g. an attempt at universalization.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T15:41:36.216Z · LW(p) · GW(p)

Attempt at universalization, isn't that a euphemism for proselytizing?

Why would [an agent whose preferences do not much intersect with the "moral" preferences of some group of agents] consider such attempts at universalization any different from other attempts of other-optimizing, which is generally a hostile act to be defended against?

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T15:57:57.622Z · LW(p) · GW(p)

Attempt at universalization, isn't that a euphemism for proselytizing?

No, people attempt to 'proselytize' their non-moral preferences too. If I attempt to share my love of My Little Pony, that doesn't mean I consider it a need for you to also love it. Even if I preferred it for you share my love of it, it would still not be a moral obligation on your part.

By universalization I didn't mean any action done after the adoption of the moral preference in question, I meant the criterion that serves to label it as a 'moral injuction' in the first place. If your brain doesn't register it as an instruction defensible by something other that your personal preferences, if it doesn't register it as a universal principle, it doesn't register as a moral instruction in the first place.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T16:03:24.479Z · LW(p) · GW(p)

What do you mean by "universal"? For any such "universally morally correct preference", what about the potentially infinite number of other agents not sharing it? Please explain.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T16:13:13.031Z · LW(p) · GW(p)

I've already given an example above: In a choice between saving my own child and saving a thousand other children, let's say I prefer saving my child. "Save my child" is a personal preference, and my brain recognizes it as such. "Save the highest number of children" can be considered a impersonal/universal instruction.

If I wanted to follow my preferences but still nonetheless claim moral perfection, I could attempt to say that the rule is really "Every parent should seek to save their own child" -- and I might even convince myself to the same. But I wouldn't say that the moral principle is really "Everyone should first seek to save the child of Aris Katsaris", even if that's what I really really prefer.

EDIT TO ADD: Also far from a recipe for war, it seem to me that morality is the opposite: an attempt at reconciling different preferences, so that people only become hostile towards only those people that don't follow a much more limited set of instructions, rather than anything in the entire set of their preferences.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T16:25:52.329Z · LW(p) · GW(p)

Why would you try to do away with your personal preferences, what makes them inferior (edit: speaking as one specific agent) to some blended average case of myriads of other humans? (Is it because of your mirror neurons? ;-)

But I wouldn't say that the moral principle is really "Everyone should first seek to save the child of Aris Katsaris", even if that's what I really really prefer.

Being you, you should strive towards that which you "really really prefer". If a particular "moral principle" (whatever you choose to label as such) is suboptimal for you (and you're not making choices for all of mankind, TDT or no), why would you endorse/glorify a suboptimal course of action?

it seem to me that morality is the opposite: an attempt at reconciling different preferences

That's called a compromise for mutual benefit, and it shifts as the group of agents changes throughout history. There's no need to elevate the current set of "mostly mutually beneficial actions" above anything but the fleeting accomodations and deals between roving tribes that they are. Best looked at through the prism of game theory.

Replies from: ArisKatsaris, TheAncientGeek, blacktrance

↑ comment by ArisKatsaris · 2014-01-10T16:38:04.924Z · LW(p) · GW(p)

Being you, you should strive towards that which you "really really prefer".

Being me, I prefer what I "really really prefer". You've not indicated why I "should" strive towards that which I "really really prefer".

If a particular "moral principle" (whatever you choose to label as such) is suboptimal for you (and you're not making choices for all of mankind, TDT or no), why would you endorse/glorify a suboptimal course of action?

When you are asking whether I "would" do something, is different than when you ask whether I "should" do something. Morality helps drive my volition, but it isn't the sole decider.

That's called a compromise for mutual benefit, and it shifts as the group of agents changes throughout history.

If you want to claim that that's the historical/evolutionary reasons that the moral instinct evolved, I agree.

If you want to argue that that's what morality is, then I disagree. Morality can drive someone to sacrifice their lives for others, so it's obviously NOT always a "compromise for mutual benefit".

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T16:54:53.011Z · LW(p) · GW(p)

If you want to argue that that's what morality is, then I disagree.

Everybody defines his/her own variant of what they call "morality", "right", "wrong", I simply suspect that the genesis of the whole "universally good" brouhaha stems from evolutionary evolved applied game theory, the "good of the tribe". Which is fine. Luckily we could now move past being bound by such homo erectus historic constraints. That doesn't mean we stop cooperating, we just start being more analytic about it. That would satisfy my preferences, that would be good.

Morality can drive someone to sacrifice their lives for others, so it's obviously NOT always a "compromise for mutual benefit".

Well, if the agent prefers sacrificing their existence for others, then doing so would be to their own benefit, no?

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T17:09:25.709Z · LW(p) · GW(p)

Well, if the agent prefers sacrificing their existence for others, then doing so would be to their own benefit, no?

sigh. Yes, given such a moral preference already in place, it somehow becomes to any person's "benefit" (for a rather useless definition of "benefit") to follow their morality.

But you previously argued that morality is a "compromise for mutual benefit", so it would follow that it only is created in order to help partially satisfy some preexisting "benefit". That benefit can't be the mere satisfaction of itself.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T17:26:57.117Z · LW(p) · GW(p)

I've called "an attempt at reconciling different preferences" a "compromise for mutual benefit". Various people call various actions "moral". The whole notion probably stems from cooperation within a tribe being of overall benefit, evolutionary speaking, but I don't claim at all that "any moral action is a compromise for mutual benefit". Who knows who calls what moral. The whole confused notion should be done away with, game theory ain't be needing no "moral".

What I am claiming is that there is non-trivial definition of morality (that is, other than "good = following your preferences") which can convince a perfectly rational agent to change its own utility function to adopt more such "moral preferences". Change, not merely relabel. The perfectly instrumentally rational agent does that which its utility functions wants. How would you even convince it otherwise? Hopefully this clarifies things a bit.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T17:59:01.860Z · LW(p) · GW(p)

Who knows who calls what moral.

My own feeling is that if you stop being so dismissive, you'll actually make some progress towards understanding "who calls what moral".

What I am claiming is that there is non-trivial definition of morality (that is, other than "good = following your preferences") which can convince a perfectly rational agent to change its own utility function to adopt more such "moral preferences"

Sure, unless someone already has a desire to be moral, talk of morality will be of no concern to them. I agree with that.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T18:11:47.568Z · LW(p) · GW(p)

Edit: Because the scenario clarifies my position, allow me to elaborate on it:

Consider a perfectly rational agent. Its epistemic rationality is flawless, that is its model of its environment is impeccable. Its instrumental rationality, without peer. That is, it is really, really good at satisfying its preferences.

It encounters a human. The human talks about what the human wants, some of which the human calls "virtuous" and "good" and is especially adamant about.

You and I, alas, are far from that perfectly rational agent. As you say, if you already have a desire to enact some actions you call morally good, then you don't need to "change" your utility function, you already have some preferences you call moral.

The question is for those who do not have a desire to do what you call moral (or who insist on their own definition, as nearly everybody does), on what grounds should they even start caring about what you call "moral"? As you say, they shouldn't, unless it benefits them in some way (e.g. makes their mammal brains feel good about being a Good Person (tm)). So what's the hubbub?

Replies from: ArisKatsaris, blacktrance

↑ comment by ArisKatsaris · 2014-01-10T20:52:53.110Z · LW(p) · GW(p)

As you say, they shouldn't, unless it benefits them in some way

I've already said that unless someone already desires to be moral, babbling about morality won't do anything for them. I didn't say it "shouldn't" (please stop confusing these two verbs)

But then you also seem to conflate this with a different issue -- of what to do with someone who does want to be moral, but understands morality differently than I do.

Which is an utterly different issue. First of all people often have different definitions to describe the same concepts -- that's because quite clearly the human brain doesn't work with definitions, but with fuzzy categorizations and instinctive "I know it when I see it" which we then attempt to make into definition when we attempt to communicate said concepts to others.

But the very fact we use the same word "morality", means we identify some common elements of what "morality" means. If we didn't mean anything similar to each other, we wouldn't be using the same word to describe it.

I find that supposedly different moralities seem to have some very common elements to them -- e.g. people tend to prefer that other people be moral. People generally agree that moral behaviour by everyone leads to happier, healthier societies. They tend to disagree about what that behaviour is, but the effects they describe tend to be common.

I might disagree with Kasparov about what the best next chess move would be, and that doesn't mean it's simply a matter of preference - we have a common understanding that the best moves are the ones that lead to an advantageous position. So, though we disagree on the best move, we have an agreement on the results of the best move.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T21:41:40.150Z · LW(p) · GW(p)

I didn't say it "shouldn't" (please stop confusing these two verbs)

What you did say was "of no concern", and "won't do anything for them", which (unless you assume infinite resources) translates to "shouldn't". It's not "conflating". Let's stay constructive.

People generally agree that moral behaviour by everyone leads to happier, healthier societies.

Such as in Islamic societies. Wrong fuzzy morality cloud?

But the very fact we use the same word "morality", means we identify some common elements of what "morality" means. If we didn't mean anything similar to each other, we wouldn't be using the same word to describe it.

Sure. What it does not mean, however, is that in between these fuzzily connected concepts is some actual, correct, universal notion of morality. Or would you take some sort of "mean", which changes with time and social conventions?

If everybody had some vague ideas about games called chess_1 to chess_N, with N being in the millions, that would not translate to some universally correct and acceptable definition of the game of chess. Fuzzy human concepts can't be assuemd to yield some iron-clad core just beyond our grasp, if only we could blow the fuzziness away. People for the most part agree what to classify as a chair. That doesn't mean there is some ideal chair we can strive for.

When checking for best moves in pre-defined chess there are definite criteria. There are non-arbitrary metrics to measure "best" by. Kasparov's proposed chess move can be better than your proposed chess move, using clear and obvious metrics. The analogy doesn't pan out:

With the fuzzy clouds of what's "moral", an outlier could -- maybe -- say "well, I'm clearly an outlier", but that wouldn't necessitate any change, because there is no objective metric to go by. Preferences aren't subject to Aumann's, or to a tyranny of the (current societal) majority.

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T22:13:33.099Z · LW(p) · GW(p)

People generally agree that moral behaviour by everyone leads to happier, healthier societies

Such as in Islamic societies. Wrong fuzzy morality cloud?

No, Islamic societies suffer from the delusion that Allah exists. If Allah existed (an omnipotent creature that punishes you horribly if you fail to obey Quran's commandments), then Islamic societies would have the right idea.

Remove their false belief in Allah, and I fail to see any great moral difference between our society and Islamic ones.

↑ comment by blacktrance · 2014-01-10T18:26:23.780Z · LW(p) · GW(p)

You're treating desires as simpler than they often are in humans. Someone can have no desire to be moral because they have a mistaken idea of what morality is or requires, are internally inconsistent, or have mistaken beliefs about how states of the world map to their utility function - to name a few possibilities. So, if someone told me that they have no desire to do what I call moral, I would assume that they have mistaken beliefs about morality, for reasons like the ones I listed. If there were beings that had all the relevant information, were internally consistent, and used words with the same sense that I use them, and they still had no desire to do what I call moral, then there would be on way for me to convince them, but this doesn't describe humans.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T18:53:46.350Z · LW(p) · GW(p)

So not doing what you call moral implies "mistaken beliefs"? How, why?

Does that mean, then, that unfriendly AI cannot exist? Or is it just that a superior agent which does not follow your morality is somehow faulty? It might not care much. (Neither should fellow humans who do not adhere to your 'correct' set of moral actions. Just saying "everybody needs to be moral" doesn't change any rational agent's preferences. Any reasoning?)

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T19:14:34.098Z · LW(p) · GW(p)

So not doing what you call moral implies "mistaken beliefs"? How, why?

For a human, yes. Explaining why this is the case would require several Main-length posts about ethical egoism, human nature and virtue ethics, and other related topics. It's a lot to go into. I'm happy to answer specific questions, but a proper answer would require describing much of (what I believe to be) morality. I will attempt to give what must be a very incomplete answer.

It's not about what I call moral, but what is actually moral. There is a variety of reasons (upbringing, culture, bad habits, mental problems, etc) that can cause people to have mistaken beliefs about what's moral. Much of what is moral is because of what's good for a person because of human nature. People's preferences can be internally inconsistent, and actually are inconsistent when they ignore or don't fully integrate this part of their preferences.

An AI doesn't have human nature, so it can be internally consistent while not doing what's moral, but I believe that if a human is immoral, it's a case of internal inconsistency (or lack of knowledge).

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T19:19:26.046Z · LW(p) · GW(p)

Is it something about the human brain? But brains evolve over time, both from genetic and from environmental influences. Worse, different human subpopulations often evolve (slightly) different paths! So which humans do you claim as a basis from which to define the one and only correct "human morality"?

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T20:13:22.573Z · LW(p) · GW(p)

Despite the differences, there is a common human nature. There is a Psychological Unity of Humankind.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T20:22:14.913Z · LW(p) · GW(p)

Noting that humans share many characteristics is an 'is', not an 'ought'. Also, this "common human nature" as exemplified throughout history is ... non too pretty as a base for some "universal mandatory morality". Yes, compared to random other mind designs pulled from mindspace, all human minds appear very similar. Doesn't imply at all that they all should strive to be similar, or to follow a similar 'codex'. Where do you get that from? It's like religion, minus god.

What you're saying that if you want to be a real human, you have to be moral? What species am I, then?

Declaring that most humans have two legs doesn't mean that every human should strive to have exactly two legs. Can't derive an 'ought' from an 'is'.

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T20:38:21.091Z · LW(p) · GW(p)

Yes, human nature is an "is". It's important because it shapes people's preferences, or, more relevantly, it shapes what makes people happy. It's not that people should strive to have two legs, but that they already have two legs, but are ignoring them. There is no obligation to be human - but you're already human, and thus human nature is already part of you.

What you're saying that if you want to be a real human, you have to be moral?

No, I'm saying that because you are human, it is inconsistent of you to not want to be moral.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T20:44:24.323Z · LW(p) · GW(p)

I feel like the discussion is stalling at this point. It comes down to you saying "if you're human you should want to be moral, because humans should be moral", which to me is as non-sequitur as it gets.

There is no obligation to be human - but you're already human, and thus human nature is already part of you.

Except if my utility function doesn't encompass what you think is "moral" and I'm human, then "following human morality" doesn't quite seem to be a prerequisite to be a "true" human, no?

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T20:51:42.778Z · LW(p) · GW(p)

It comes down to you saying "if you're human you should want to be moral, because humans should be moral"

No, that isn't what I'm saying. I'm saying that if you're human, you should want to be moral, because wanting to be moral follows from the desires of a human with consistent preferences, due in part to human nature.

if my utility function doesn't encompass what you think is "moral" and I'm human

Then I dispute that your utility function is what you think it is.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T21:18:41.155Z · LW(p) · GW(p)

I'm saying that if you're human, you should want to be moral, because wanting to be moral follows from the desires of a human with consistent preferences, due in part to human nature.

The error as I see it is that "human nature", whatever you see as such, is a statement about similarities, it isn't a statement about how things should be.

It's like saying "a randomly chosen positive natural number is really big, so all numbers should be really big". How do you see that differently?

We've already established that agents can have consistent preferences without adhering to what you think of as "universal human morality". Child soldiers are human. Their preferences sure can be brutal, but they can be as internally consistent or inconsistent as those of anyone else. I sure would like to change their preferences, because I'd prefer for them to be different, not because some 'idealised human spirit' / 'psychic unity of mankind' ideal demands so.

Then I dispute that your utility function is what you think it is.

Proof by demonstration? Well, lock yourself in a cellar with only water and send me a key, I'll send it back FedEx with instructions to set you free, after a week. Would that suffice? I'd enjoy proving that I know my own utility function better than you know my utility function (now that would be quite weird), I wouldn't enjoy the suffering. Who knows, might even be healthy overall.

Replies from: Jiro, blacktrance, nshepperd

↑ comment by Jiro · 2014-01-10T22:53:28.064Z · LW(p) · GW(p)

It's like saying "a randomly chosen positive natural number is really big, so all numbers should be really big".

You can't randomly choose a positive natural number using an even distribution. If you use an uneven distribution, whether the result is likely to be big depends on how your distribution compares to your definition of "big".

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T23:09:01.734Z · LW(p) · GW(p)

Choose from those positive numbers that a C++ int variable can contain, or any other* non-infinite subset of positive natural numbers, then. The point is the observation of "most numbers need more than 1 digit to be expressed" not implying in any way some sort of "need" for the 1-digit numbers to "change", to satisfy the number fairy, or some abstract concept thereof.

* (For LW purposes: Any other? No, not any other. Choose one with a cardinality of at least 10^6. Heh.)

↑ comment by blacktrance · 2014-01-10T21:43:45.764Z · LW(p) · GW(p)

The erro as I see it is that "human nature", whatever you see as such, is a statement about similarities, it isn't a statement about how things should be.

It is a statement about similarities, but it's about a similarity that shapes what people should do. I don't know how I can explain it without repeating myself, but I'll try.

For an analogy, let's consider beings that aren't humans. Paperclip maximizers, for example. Except these paperclip maximizers aren't AIs, but a species that somehow evolved biologically. They're not perfect reasoners and can have internally inconsistent preferences. These paperclip maximizers can prefer to do something that isn't paperclip-maximizing, even though that is contrary to their nature - that is, if they were to maximize paperclips, they would prefer it to whatever they were doing earlier. One day, a paperclip maximizer who is maximizing paperclips tells his fellow clippies, "You should maximize paperclips, because if you did, you would prefer to, as it is your nature". This clippy's statement is true - the clippies' nature is such that if they maximized clippies, they would prefer it to other goals. So, regardless of what other clippies are actually doing, the utility-maximizing thing for them to do would be to maximize paperclips.

So it is with humans. Upon discovering/realizing/deriving what is moral and consistently acting/being moral, the agent would find that being moral is better than the alternative. This is in part due to human nature.

We've already established that agents can have consistent preferences without adhering to what you think of as "universal human morality".

Agents, yes. Humans, no. Just like the clippies can't have consistent preferences if they're not maximizing paperclips.

Proof by demonstration? Well, lock yourself in a cellar with only water and send me a key, I'll send it back FedEx with instructions to set you free, after a week. Would that suffice?

What would that prove? Also, I don't claim that I know the entirety of your utility function better than you do - you know much better than I do what kind of ice cream you prefer, what TV shows you like to watch, etc. But those have little to do with human nature in the sense that we're talking about it here.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T21:48:29.844Z · LW(p) · GW(p)

Just like the clippies can't have consistent preferences if they're not maximizing paperclips.

A clippy which isn't maximizing paperclips is not a clippy.

A human which isn't adhering to your moral codex is still a human.

What would that prove?

That my utility function includes something which you'd probably consider immoral.

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T21:56:39.938Z · LW(p) · GW(p)

A clippy which isn't maximizing paperclips is not a clippy.

It's a clippy because it would maximize paperclips if it had consistent preferences and sufficient knowledge.

That my utility function includes something which you'd probably consider immoral.

I don't dispute that this is possible. What I dispute is that your utility function would contain that if you were internally consistent (and had knowledge of what being moral is like).

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T22:10:21.859Z · LW(p) · GW(p)

The desires of an agent are defined by its preferences. "This is a paperclip maximizer which does not want to maximize paperclips" is a contradiction in terms. And what do you mean by "consistent", do you mean "consistent with 'human nature'? Who cares? Or consistent within themselves? Highly doubtful, what would internal consistency have to do with being an altruist? If there's anything which is characteristic of "human nature", it is the inconsistency of their preferences.

A human which doesn't share what you think of as "correct" values (may I ask, not disparagingly, are you religious?) is still a human. An unusual one, maybe (probably not), but an agent not in "need" of any change towards more "moral" values. Stalin may have been happy the way he was.

I don't dispute that this is possible. What I dispute is that your utility function would contain that if you were internally consistent (and had knowledge of what being moral is like).

Because of the warm fuzzies? The social signalling? Is being moral awsome, or deeply fulfilling? Are you internally consistent ... ?

Replies from: blacktrance, Nornagest

↑ comment by blacktrance · 2014-01-11T01:11:53.932Z · LW(p) · GW(p)

"This is a paperclip maximizer which does not want to maximize paperclips" is a contradiction in terms.

Call it a quasi-paperclip maximizer, then. I'm not interested in disputing definitions. Whatever you call it, it's a being whose preferences are not necessarily internally consistent, but when they are, it prefers to maximize paperclips. When its preferences are internally inconsistent, it may prefer to do things and have goals other than maximizing paperclips.

Highly doubtful, what would internal consistency have to do with being an altruist?

There's no necessary connection between the two, but I'm not equating morality and altruism. Morality is what one should do and/or how one should be, which need not be altruistic.

Humans can have incorrect values and still be human, but in that case they are internally inconsistent., because of the preferences they have due to human nature. I'm not saying that humans should strive to have human nature, I'm saying that they already have it. I doubt that Stalin was happy - just look at how paranoid he was. And no, I'm not religious, and have never been.

Because of the warm fuzzies? The social signalling? Is being moral awsome, or deeply fulfilling?

Yes to the first and third questions, Being moral is awesome and fulfilling. It makes you feel happier, more fulfilled, more stable, and similar feelings. It doesn't guarantee happiness, but it contributes to it both directly (being moral feels good) and indirectly (it helps you make good decisions). It makes you stronger and more resilient (once you've internalized it fully). It's hard to describe beyond that, but good feels good (TVTropes warning).

I think I'm internally consistent. I've been told that I am. It's unlikely that I'm perfectly consistent, but whatever inconsistencies I have are probably minor. I'm open to having them addressed, whatever they are.

Replies from: Jiro, hairyfigment

↑ comment by Jiro · 2014-01-11T09:21:39.241Z · LW(p) · GW(p)

Claiming that Stalin wasn't happy sounds like a variation of sour grapes where not only can you not be as successful as him, it would be actively uncomfortable for you to believe that someone who lacks compassion can be happy, so you claim that he's not.

It's true he was paranoid but it's also true that in the real world, there are tradeoffs and you don't see people becoming happy with no downsides whatsoever--claiming that this disqualifies them from being called happy eviscerates the word of meaning.

I'm also not convinced that Stalin's "paranoia" was paranoia (it seems rationa for someone who doesn't care about the welfare of others and can increase his safety by instilling fear and treating everyone as enemies to do so). I would also caution against assuming that since Stalin's paranoia is prominent enough for you to have heard of it, it's too big a deal for him to have been happy--it's promiment enough for you to have heard of it because it was a big deal to the people affected by it, which is unrelated to how much it affected his happiness.

Replies from: blacktrance, army1987

↑ comment by blacktrance · 2014-01-11T18:39:51.345Z · LW(p) · GW(p)

Stalin was paranoid even by the standards of major world leaders. Khrushchev wasn't so paranoid, for example. Stalin saw enemies behind every corner. That is not a happy existence.

Replies from: Jiro

↑ comment by Jiro · 2014-01-11T22:43:40.231Z · LW(p) · GW(p)

Khruschev was deposed. Stalin stayed dictator until he died of natural causes. That suggests that Khruschev wasn't paranoid enough, while Stalin was appropriately paranoid.

Seeing enemies around every corner meant that sometimes he saw enemies that weren't there, but it was overall adaptive because it resulted in him not getting defeated by any of the enemies that actually existed. (Furthermore, going against nonexistent enemies can be beneficial insofar as the ruthlessness in going after them discourages real enemies.)

Stalin saw enemies behind every corner. That is not a happy existence.

How does the last sentence follow from the previous one? It's certainly not as happy an existence as it would have been if he had no enemies, but as I pointed out, nobody's perfectly happy. There are always tradeoffs and we don't claim that the fact that someone had to do something to gain his happiness automatically makes that happiness fake.

Replies from: SaidAchmiz

↑ comment by Said Achmiz (SaidAchmiz) · 2014-01-12T04:48:42.563Z · LW(p) · GW(p)

Stalin's paranoia, and the actions he took as a result, also created enemies, thus becoming a partly self-fulfilling attitude.

↑ comment by A1987dM (army1987) · 2014-01-12T09:49:28.859Z · LW(p) · GW(p)

you don't see people becoming happy with no downsides whatsoever

You do see people becoming happy with fewer downsides than others, though.

↑ comment by hairyfigment · 2014-01-11T09:44:17.523Z · LW(p) · GW(p)

Stalin refused to believe Hitler would attack him, probably since that would be suicidally stupid on the attacker's part. Was he paranoid, or did he update?

↑ comment by Nornagest · 2014-01-10T22:21:07.785Z · LW(p) · GW(p)

The desires of an agent are defined by its preferences. "This is a paperclip maximizer which does not want to maximize paperclips" is a contradiction in terms.

I'm not sure "preference" is a powerful enough term to capture an agent's true goals, however defined. Consider any of the standard preference reversals: a heavy cigarette smoker, for example, might prefer to buy and consume their next pack in a Near context, yet prefer to quit in a Far. The apparent contradiction follows quite naturally from time discounting, yet neither interpretation of the person's preferences is obviously wrong.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T22:27:05.139Z · LW(p) · GW(p)

I've seen it used as shorthand for "utility function", saving 5 keystrokes. That was the intended use here. Point taken, alternate phrasings welcome.

↑ comment by nshepperd · 2014-01-10T23:25:28.093Z · LW(p) · GW(p)

Proof by demonstration? Well, lock yourself in a cellar with only water and send me a key, I'll send it back FedEx with instructions to set you free, after a week.

That would only prove that you think you want to do that. The issue is that what you think you want and what you actually want do not generally coincide, because of imperfect self-knowledge, bounded thinking time, etc.

I don't know about child soldiers, but it's fairly common for amateur philosophers to argue themselves into thinking they "should" be perfectly selfish egoists, or hedonistic utilitarians, because logic or rationality demands it. They are factually mistaken, and to the extent that they think they want to be egoists or hedonists, their "preferences" are inconsistent, because if they noticed the logical flaw in their argument they would change their minds.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T23:45:44.623Z · LW(p) · GW(p)

That would only prove that you think you want to do that.

Isn't that when I throw up my arms and say "congratulations, your hypothesis is unfalsifiable, the dragon is permeable to fluor". What experimental setup would you suggest? Would you say any statement about one's preferences is moot? It seems that we're always under bounded thinking time constraints. Maybe the paperclipper really wants to help humankind and be moral, and mistakingly thinks otherwise. Who would know, it optimized its own actions under resource constraints, and then there's the 'Löbstacle' to consider.

Is saying "I like vanilla ice cream" FAI-complete and must never be uttered or relied upon by anyone?

it's fairly common for amateur philosophers to argue themselves into thinking they "should" be perfectly selfish egoists, or hedonistic utilitarians, because logic or rationality demands it

Or argue themselves into thinking that there is some subset of preferences such every other (human?) agent should voluntarily choose to adopt them, against their better judgment (edit; as it contradicts what they (perceive, after thorough introspection) as their own preferences)? You can add "objective moralists" to the list.

What would it be that is present in every single human's brain architecture throughout human history that would be compatible with some fixed ordering over actions, called "morally good"? (Otherwise you'd have your immediate counterexample.) The notion seems so obviously ill-defined and misguided (hence my first comment asking Cousin_It).

It's fine (to me) to espouse preferences that aim to change other humans (say, towards being more altruistic, or towards being less altruistic, or whatever), but to appeal to some objective guiding principle based on "human nature" (which constantly evolves in different strands) or some well-sounding ev-psych applause-light is just a new substitute for the good old Abrahamic heavenly father.

Replies from: nshepperd

↑ comment by nshepperd · 2014-01-11T03:04:49.426Z · LW(p) · GW(p)

Would you say any statement about one's preferences is moot? It seems that we're always under bounded thinking time constraints. Maybe the paperclipper really wants to help humankind and be moral, and mistakingly thinks otherwise. Who would know, it optimized its own actions under resource constraints, and then there's the 'Löbstacle' to consider.

Is saying "I like vanilla ice cream" FAI-complete and must never be uttered or relied upon by anyone?

I wouldn't say any of those things. Obviously paperclippers don't "really want to help humankind", because they don't have any human notion of morality built-in in the first place. Statements like "I like vanilla ice cream" are also more trustworthy on account of being a function of directly observable things like how you feel when you eat it.

The only point I'm trying to make here is that it is possible to be mistaken about your own utility function. It's entirely consistent for the vast majority of humans to have a large shared portion of their built-in utility function (built-in by their genes), even though many of them seemingly want to do bad things, and that's because humans are easily confused and not automatically self-aware.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-11T08:24:54.552Z · LW(p) · GW(p)

It is possible to be mistaken about your own utility function.

For sure.

It's entirely consistent for the vast majority of humans to have a large shared portion of their built-in utility function (built-in by their genes), even though many of them seemingly want to do bad things

I'd agree if humans were like dishwashers. There are templates for dishwashers, ways they are supposed to work. If you came across a broken dishwasher, there could be a case for the dishwasher to be repaired, to go back to "what it's supposed to be".

However, that is because there is some external authority (exasparated humans who want to fix their damn dishwasher, dirty dishes are piling up) conceiving of and enforcing such a purpose. The fact that genes and the environment shape utility functions in similar ways is a description, not a prescription. It would not be a case for any "broken" human to go back to "what his genes would want him to be doing". Just like it wouldn't be a case against brain uploading.

Some of the discussion seems to me like saying that "deep down in every flawed human, there is 'a figure of light', in our community 'a rational agent following uniform human values with slight deviations accounting for ice-cream taste', we just need to dig it up". There is only your brain. With its values. There is no external standard to call its values flawed. There are external standards (rationality = winning) to better its epistemic and instrumental rationality, but those can help the serial killer and the GiveWell activist equally. (Also, both of those can be 'mistaken' about their values.)

↑ comment by TheAncientGeek · 2014-01-10T16:58:12.325Z · LW(p) · GW(p)

Why would you try to do away with your personal preferences, what makes them inferior (edit: speaking as one specific agent) to some blended average case of myriads of other humans? (Is it because of your mirror neurons? ;-)

If you have a preference for morality, being moral is not doing away with that prrefence: it is allowing your altruistic prefences to override your selfish ones.
You may be on the receving end of someone else's self sacrifice at some point

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T17:06:22.669Z · LW(p) · GW(p)

Certainly, but in that case your preference for the moral action is your personal preference, which is your 'selfish' preference. No conflict there. You should always do that which maximizes your utility function. If you call that moral, we're in full agreement. If your utility function is maximized by caring about someone else's utility function, go for it. I do, too.
That's nice. Why would that cause me to do things which I do not overall prefer to do? Or do you say you always value that which you call moral the most?

Replies from: ArisKatsaris

↑ comment by ArisKatsaris · 2014-01-10T17:16:07.908Z · LW(p) · GW(p)

Certainly, but in that case your preference for the moral action is your personal preference, which is your 'selfish' preference.

I can make a quite clear distinction between my preferences relating to an apersonal loving-kindness towards the universe in general, and the preferences that center around my personal affections and likings.

You keep trying to do away with a distinction that has huge predictive ability: a distinction that helps determine what people do, why they do it, how they feel about doing it, and how they feel after doing it.

If your model of people's psychology conflates morality and non-moral preferences, your model will be accurate only for the most amoral of people.

↑ comment by blacktrance · 2014-01-10T16:39:18.126Z · LW(p) · GW(p)

Morality is a somewhat like chess in this respect - morality:optimal play::satisfying your preferences:winning. To simplify their preferences a bit, chess players want to win, but no individual chess player would claim that all other chess players should play poorly so he can win.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T16:47:26.878Z · LW(p) · GW(p)

To simplify their preferences a bit, chess players want to win, but no individual chess player would claim that all other chess players should play poorly so he can win.

That's explained simply by 'winning only against bad players' not being the most valued component of their preferences, preferring 'wins when the other player did his/her very best and still lost' instead. Am I missing your point?

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T17:06:20.219Z · LW(p) · GW(p)

Sorry, I didn't explain well. To approach the explanation from different angles:

Even if all chess players wanted was to win, it would still be incorrect for them to claim that playing poorly is the correct way to play. Just like when I'm hungry, I want to eat, but I don't claim that strangers should feed me for free.
Consider the prisoners' dilemma, as analyzed traditionally. Each prisoner wants the other to cooperate, but neither can claim that the other should cooperate.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T17:18:27.175Z · LW(p) · GW(p)

Even if all chess players wanted was to win, it would still be incorrect for them to claim that playing poorly is the correct way to play. Just like when I'm hungry, I want to eat, but I don't claim that strangers should feed me for free.

Incorrect because that's not what the winning player would prefer. You don't claim that strangers should feed you because that's what you prefer. It's part of your preferences. Some of your preferences can rely on satisfying someone else's preferences. Such altruistic preferences are still your own preferences. Helping members of your tribe you care about. Cooperating within your tribe, enjoying the evolutionary triggered endorphins.

You're probably thinking that considering external preferences and incorporating them in your own utility function is a core principle of being "morally right". Is that so?

So the core disagreement (I think): Take an agent with a given set of preferences. Some of these may include the preferences of others, some may not. On what basis should that agent modify its preferences to include more preferences of others, i.e. to be "more moral"?

Consider the prisoners' dilemma, as analyzed traditionally. Each prisoner wants the other to cooperate, but neither can claim that the other should cooperate.

So you can imagine yourself in someone else's position, then say "What B should do from A's perspective" is different from "What B should do from B's perspective". Then you can enter all sorts of game theoretic considerations. Where does morality come in?

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-11T00:44:52.843Z · LW(p) · GW(p)

So you can imagine yourself in someone else's position, then say "What B should do from A's perspective" is different from "What B should do from B's perspective". Then you can enter all sorts of game theoretic considerations. Where does morality come in?

There is no "What B should do from A's perspective", from A's perspective there is only "What I want B to do". It's not a "should". Similarly, the chess player wants his opponent to lose, and I want people to feed me, but neither of those are "should"s. "Should"s are only from an agent's own perspective applied to themselves, or from something simulating that perspective (such as modeling the other player in a game). "What B should do from B's perspective" is equivalent to "What B should do".

↑ comment by TheAncientGeek · 2014-01-10T15:27:17.836Z · LW(p) · GW(p)

The key issue is that, whilst morality is not tautologously the same as preferences, a morally right action is, tautologously, what you should do.

So it is difficult to see on what grounds Mark can object to the FAIs wishes: if it tells him something is mortally right that is what he should do. And he can't have his own separate morality, because the idea is incoherent.

Replies from: ArisKatsaris, Kawoomba

↑ comment by ArisKatsaris · 2014-01-10T15:44:58.764Z · LW(p) · GW(p)

So it is difficult to see on what grounds Mark can object to the FAIs wishes: if it tells him something is mortally right that is what he should do.

A distinction to be made: Mark can wish differently than the AI wishes, Mark can't morally object to the AI's wishes (if the AI follows morality).

Exactly because morality is not the same as preferences.

↑ comment by Kawoomba · 2014-01-10T15:35:03.309Z · LW(p) · GW(p)

You can call a subset of your preferences moral, that's fine. Say, eating chocolate icecream, or helping a starving child. Let's take a randomly chosen "morally right action" A.

That, given your second paragraph, would have to be a preference which, what, maximizes Mark's utility, regardless of what the rest of his utility function actually looks like?

It seems to be trivial to construct a utility function (given any such action A) such as that doing A does not maximize said utility function. Give Mark a such a utility function and you got yourself a reductio ad absurdum.

So, if you define a subset of preferences named "morally right" thus that any such action needs to maximize (edit: or even 'not minimize') an arbitrary utility function, then obviously that subset is empty.

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T16:00:46.688Z · LW(p) · GW(p)

That, given your second paragraph, would have to be a preference which, what, maximizes Mark's utility, regardless of what the rest of his utility function actually looks like?

If Mark is capable of acting morally, he would have a preference for moral action which is strong enough to override other preferences. However,t hat is not really the point. Even if he is too weak-willed to do what the FAI says, he has no grounds to object to the FAI.

It seems to be trivial to construct a utility function (given any such action A) such as that doing A does not maximize said utility function. Give Mark a such a utility function and you got yourself a reductio ad absurdum.

I can't see how that amount to more than the observation that not every agetn is capable of acting morally. Ho hum.

So, if you define a subset of preferences named "morally right" thus that any such action needs to maximize (edit: or even 'not minimize') an arbitrary utility function, then obviously that subset is empty.

I don't see why. An agent should want to do what is morally right, but that doesn't mean an agent would want to. Their utility funciton might not allow them. But how could they object to be told what is right? The fault, surely, lies in themselves.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T16:14:45.066Z · LW(p) · GW(p)

An agent should want to do what is morally right, but that doesn't mean an agent would want to. Their utility funciton might not allow them. But how could they object to be told what is right? The fault, surely, lies in themselves.

They can object because their preferences are defined by their utility function, full stop. That's it. They are not "at fault", or "in error", for not adopting some other preferences that some other agents deem to be "morally correct". They are following their programming, as you follow yours. Different groups of agents share different parts of their preferences, think Venn diagram.

If the oracle tells you "this action maximizes your own utility function, you cannot understand how", then yes the agent should follow the advice.

If the oracle told an agent "do this, it is morally right", the non-confused agent would ask "do you mean it maximizes my own utility function?". If yes, "thanks, I'll do that", if no "go eff yourself!".

You can call an agent "incapable of acting morally" because you don't like what it's doing, it needn't care. It might just as well call you "incapable of acting morally" if your circles of supposedly "morally correct actions" don't intersect.

↑ comment by TheOtherDave · 2014-01-10T14:21:29.652Z · LW(p) · GW(p)

I can't speak for cousin_it, natch, but for my own part I think it has to do with mutually exclusive preferences vs orthogonal/mutually reinforcing preferences. Using moral language is a way of framing a preference as mutually exclusive with other preferences.

That is... if you want A and I want B, and I believe the larger system allows (Kawoomba gets A AND Dave gets B), I'm more likely to talk about our individual preferences. If I don't think that's possible, I'm more likely to use universal language ("moral," "optimal," "right," etc.), in order to signal that there's a conflict to be resolved. (Well, assuming I'm being honest.)

For example, "You like chocolate, I like vanilla" does not signal a conflict; "Chocolate is wrong, vanilla is right" does.

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T14:56:40.165Z · LW(p) · GW(p)

Why stop at connotation and signalling? If there is a non-empty set of preferences whose satistfaction is inclined to lead to conflict, and a non-empty set of preferences that can be satisfied withotu conflict, then "morally relevant prefernece" can denote the members of the first set...which is not idenitcal to the set of all preferences.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T15:52:22.192Z · LW(p) · GW(p)

For any such preference, you can immediately provide a utility function such that the corresponding agent would be very unhappy about that preference, and would give its life to prevent it.

Or do you mean "a set of preferences the implementation of which would on balance benefit the largest amount of agents the most"? That would change as the set of agents changes, so does the "correct" morality change too, then?

Also, why should I or anyone else particular care about about such preferences (however you define them), especially as the "on average" doesn't benefit me? Is it because evolutionary speaking, that's how what evolved? What our mirror neurons lead us towards? Wouldn't that just be a case of the naturalistic fallacy?

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T16:13:01.744Z · LW(p) · GW(p)

For any such preference, you can immediately provide a utility function such that the corresponding agent would be very unhappy about that preference

Sure. So what? Kids don't like teachers and criminals don't like the police..but they can't object to them, because "entitiy X is stopping from doing bad things and making me do good things" is no (rational, adult) objection.

Also, why should I or anyone else particular care about about such preferences (however you define them), especially as the "on average" doesn't benefit me?

If being moral increases your utility, it increases your utility -- what other sense of "benefitting me" is there?

Replies from: blacktrance, Kawoomba

↑ comment by blacktrance · 2014-01-10T16:26:19.499Z · LW(p) · GW(p)

If being moral increases your utility, it increases your utility -- what other sense of "benefitting me" is there?

If utility is the satisfaction of preferences, and you can have preferences that don't benefit you (such as doing heroin), increasing your utility doesn't necessarily benefit you.

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T16:37:30.799Z · LW(p) · GW(p)

If you can get utility out of paperclips, why can't you get it out of heorin? You're surely not saying that there is some sort of Objective utility that everyone ought to have in their UF's?

Replies from: blacktrance

↑ comment by blacktrance · 2014-01-10T16:51:57.707Z · LW(p) · GW(p)

You can get utility out of heroin if you prefer to use it, which is an example of "benefiting me" and utility not being synonymous. I don't think there's any objective utility function for all conceivable agents, but as you get more specific in the kinds of agents you consider (i.e. humans), there are commonalities in their utility functions, due to human nature. Also, there are sometimes inconsistencies between (for lack of better terminology) what people prefer and what they really prefer - that is, people can act and have a preference to act in ways that, if they were to act differently, they would prefer the different act.

↑ comment by Kawoomba · 2014-01-10T16:17:53.849Z · LW(p) · GW(p)

Kids don't like teachers and criminals don't like the police..but they can't object to them (...)

(Kids - teachers), (criminals - police), so is "morally correct" defined by the most powerful agents, then?

If being moral increases your utility (...)

And if being moral (whatever it may mean) does not?

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2014-01-10T16:47:56.127Z · LW(p) · GW(p)

(Kids - teachers), (criminals - police), so is "morally correct" defined by the most powerful agents, then?

Adult, rational objections are objections that other agents might feel impelled to do somehting about, and so are not just based on "I don't like it"."I don't like it" is no objectio to "you should do your homework", etc.

If being moral increases your utility (...)

And if being moral (whatever it may mean) does not?

Then you would belong to the set of Immoral Agents, AKA Bad People.

Replies from: Kawoomba

↑ comment by Kawoomba · 2014-01-10T17:00:59.573Z · LW(p) · GW(p)

"You should do your homework (... because it is in your own long-term best interest, you just can't see that yet)" is in the interest of the kid, cf. an FAI telling you to do an action because it is in your interest. "You should jump out that window (... because it amuses me / because I call that morally good)" is not in your interest, you should not do that. In such cases, "I don't like that" is the most pertinent objection and can stand all on its own.

Then you would belong to the set of Immoral Agents, AKA Bad People.

Boo bad people! What if we encountered aliens with "immoral" preferences?

↑ comment by TheOtherDave · 2014-01-10T14:08:38.731Z · LW(p) · GW(p)

For my own part: denotationally, yes, I would understand "Do you prefer (that Dave eat) chocolate or vanilla ice cream?" and "Do you consider (Dave eating) chocolate ice cream or vanilla as the morally superior flavor for (Dave eating) ice cream?" as asking the same question.

Connotationally, of course, the latter has all kinds of (mostly ill-defined) baggage the former doesn't.

↑ comment by TheAncientGeek · 2014-01-10T16:26:18.145Z · LW(p) · GW(p)

Are you using a meaning of "morality" that is distinct from "preferences"? You bet.

↑ comment by Pentashagon · 2014-01-10T03:31:18.870Z · LW(p) · GW(p)

My point was that trying to use a provably-boxed AI to do anything useful would probably not work, including trying to design unboxed FAI, not that we should design boxed FAI. I may have been pessemistic, see Stuart Armstrong's proposal of reduced impact AI which sounds very similar to provably boxed AI but which might be used for just about everything including designing a FAI.

↑ comment by Houshalter · 2013-10-01T08:10:36.157Z · LW(p) · GW(p)

I think we might have different definitions of a boxed-AI. An AI that is literally not allowed to interact with the world at all isn't terribly useful and it sounds like a problem at least as hard as all other kinds of FAI.

I just mean a normal dangerous AI that physically can't interact with the outside world. Importantly it's goal is to provably give the best output it possibly can if you give it a problem. So it won't hide nanotech in your cure for alzheimers because that would be a less fit and more complicated solution than a simple chemical compound (you would have to judge solutions based on complexity though and verify them by a human or in a simulation first just in case.)

I don't think most computers today have anywhere near enough processing power to simulate a full human brain. A human down to the molecular level is entirely out of the question. An AI on a modern computer, if it's smarter than human at all, will get there by having faster serial processing or more efficient algorithms, not because it has massive raw computational power.

And you can always scale down the hardware or charge it utility for using more computing power than it needs, forcing it to be efficient or limiting it's intelligence further. You don't need to invoke the full power of super-intelligence for every problem and for your safety you probably shouldn't.

↑ comment by Eugene · 2013-10-11T19:50:53.078Z · LW(p) · GW(p)

A slightly bigger "large risk" than Pentashagon puts forward is that a provably boxed UFAI could indifferently give us information that results in yet another UFAI, just as unpredictable as itself (statistically speaking, it's going to give us more unhelpful information than helpful, as Robb point out). Keep in mind I'm extrapolating here. At first you'd just be asking for mundane things like better transportation, cures for diseases, etc. If the UFAI's mind is strange enough, and we're lucky enough, then some of these things result in beneficial outcomes, politically motivating humans to continue asking it for things. Eventually we're going to escalate to asking for a better AI, at which point we'll get a crap-shoot.

An even bigger risk than that -though - is that if it's especially Unfriendly, it may even do this intentionally, going so far as to pretend it's friendly while bestowing us with data to make an AI even more Unfriendly AI than itself. So what do we do, box that AI as well, when it could potentially be even more devious than the one that already convinced us to make this one? Is it just boxes, all the way down? (spoilers: it isn't, because we shouldn't be taking any advice from boxed AIs in the first place)

The only use of a boxed AI is to verify that, yes, the programming path you went down is the wrong one, and resulted in an AI that was indifferent to our existence (and therefore has no incentive to hide its motives from us). Any positive outcome would be no better than an outcome where the AI was specifically Evil, because if we can't tell the difference in the code prior to turning it on, we certainly wouldn't be able to tell the difference afterward.

comment by kilobug · 2013-09-04T07:46:35.102Z · LW(p) · GW(p)

If an artificial intelligence is smart enough to be dangerous, we'd intuitively expect it to be smart enough to know how to make itself safe.

I don't agree with that. Just looks at humans, they are smart enough to be dangerous, but even when they do want to "make themselves safe", they are usually unable to do so. A lot of harm is done by people with good intent. I don't think all of Moliere doctors prescribing bloodletting were intending to do harm.

Yes, a sufficiently smart AI will know how to make itself safe if it wishes, but the intelligence level required for that is much higher than the one required to be harmful.

Replies from: RobbBB

↑ comment by Rob Bensinger (RobbBB) · 2013-09-06T17:15:18.108Z · LW(p) · GW(p)

Agreed. The reason I link the two abilities is that I'm assuming an AI that acquires either power went FOOM, which makes it much more likely that the two powers will arise at (on a human scale) essentially the same time.

Replies from: Lethalmud

↑ comment by Lethalmud · 2014-01-10T11:22:32.565Z · LW(p) · GW(p)

If a FAI would have a utility function like "Maximise X while remaining Friendly", And the UFAI would just have "Maximise X". Then, If the FAI and a UFAI would be initiated simultaneously, I would expect them both to develop exponentially, but the UFAI would have more options available, thus have a steeper learning curve. So I'd expect that in this situation that the UFAI would go FOOM slightly sooner, and be able to disable the FAI.

comment by Wei Dai (Wei_Dai) · 2013-09-04T06:53:04.415Z · LW(p) · GW(p)

And if we do discover the specific lines of code that will get an AI to perfectly care about its programmer's True Intentions, such that it reliably self-modifies to better fit them — well, then that will just mean that we've solved Friendliness Theory. The clever hack that makes further Friendliness research unnecessary is Friendliness.

Some people seem to be arguing that it may not be that hard to discover these specific lines of code. Or perhaps that we don't need to get an AI to "perfectly" care about its programmer's True Intentions. I'm not sure if I understand their arguments correctly so I may be unintentionally strawmanning them, but the idea may be that if we can get an AI to approximately care about its programmer or user's intentions, and also prevent it from FOOMing right away (or just that the microeconomics of intelligence explosion doesn't allow for such fast FOOMing), then we can make use of the AI in a relatively safe way to solve various problems, including the problem of how to control such AIs better, or how to eventually build an FAI. What's your take on this class of arguments?

Being Friendly is of instrumental value to barely any goals.

Tangentially, being Friendly is probably of instrumental value to some goals, which may turn out to be easier to instill in an AGI than solving Friendliness in the traditional terminal values sense. I came up with the term "Instrumentally Friendly AI" to describe such an approach.

Replies from: XiXiDu

↑ comment by XiXiDu · 2013-09-04T11:42:07.659Z · LW(p) · GW(p)

Nobody disagrees that an arbitrary agent pulled from mind design space, that is powerful enough to overpower humanity, is an existential risk if it either exhibits Omohundro's AI drives or is used as a tool by humans, either carelessly or to gain power over other humans.

Disagreeing with that would about make as much sense as claiming that out-of-control self-replicating robots could somehow magically turn the world into a paradise, rather than grey goo.

The disagreement is mainly about the manner in which we will achieve such AIs, how quickly that will happen, and whether such AIs will have these drives.

I actually believe that much less than superhuman general intelligence might be required for humans to cause extinction type scenarios.

Most of my posts specifically deal with the scenario and arguments publicized by MIRI. Those posts are not highly polished papers but attempts to reduce my own confusion and to enable others to provide feedback.

I argue that...

...the idea of a vast mind design space is largely irrelevant, because AIs will be created by humans, which will considerably limit the kind of minds we should expect.
...that AIs created by humans do not need to, and will not exhibit any of Omohundro's AI drives.
...that even given Omohundro's AI drives, it is not clear how such AIs would arrive at the decision to take over the world.
...that there will be no fast transition from largely well-behaved narrow AIs to unbounded general AIs, and that humans will be part of any transition.
...that any given AI will initially not be intelligent enough to hide any plans for world domination.
...that drives as outlined by Omohundro would lead to a dramatic interference with what the AI's creators want it to do, before it could possibly become powerful enough to deceive or overpower them, and would therefore be noticed in time.
...that even if MIRI's scenario comes to pass, there is a lack of concrete scenarios on how such an AI could possibly take over the world, and that the given scenarios raise many questions.

There are a lot more points of disagreement.

What I, and I believe Richard Loosemore as well, have been arguing, as quoted above, is just one specific point that is not supposed to say much about AI risks in general. Below is an distilled version of what I personally meant:

1. Superhuman general intelligence, obtained by the self-improvement of a seed AI, is a very small target to hit, requiring a very small margin of error.

2. Intelligently designed systems do not behave intelligently as a result of unintended consequences. (See note 1 below.)

3. By step 1 and 2, for an AI to be able to outsmart humans, humans will have to intend to make an AI capable of outsmarting them and succeed at encoding their intention of making it outsmart them.

4. Intelligence is instrumentally useful, because it enables a system to hit smaller targets in larger and less structured spaces. (See note 2, 3.)

5. In order to take over the world a system will have to be able to hit a lot of small targets in very large and unstructured spaces.

6. The intersection of the sets of “AIs in mind design space” and “the first probable AIs to be expected in the near future” contains almost exclusively those AIs that will be designed by humans.

7. By step 6, what an AI is meant to do will very likely originate from humans.

8. It is easier to create an AI that applies its intelligence generally than to create an AI that only uses its intelligence selectively. (See note 4.)

9. An AI equipped with the capabilities required by step 5, given step 7 and 8, will very likely not be confused about what it is meant to do, if it was not meant to be confused.

10. Therefore the intersection of the sets of “AIs designed by humans” and “dangerous AIs” only contains almost exclusively those AIs which are deliberately designed to be dangerous by malicious humans.

Notes

Software such as Mathematica will not casually prove the Riemann hypothesis if it has not been programmed to do so. Given intelligently designed software, world states in which the Riemann hypothesis is proven will not be achieved if they were not intended because the nature of unintended consequences is overall chaotic.
As the intelligence of a system increases the precision of the input, that is necessary to make the system do what humans mean it to do, decreases. For example, systems such as IBM Watson or Apple’s Siri do what humans mean them to do when fed with a wide range of natural language inputs. While less intelligent systems such as compilers or Google Maps need very specific inputs in order to satisfy human intentions. Increasing the intelligence of Google Maps will enable it to satisfy human intentions by parsing less specific commands.
When producing a chair an AI will have to either know the specifications of the chair (such as its size or the material it is supposed to be made of) or else know how to choose a specification from an otherwise infinite set of possible specifications. Given a poorly designed fitness function, or the inability to refine its fitness function, an AI will either (a) not know what to do or (b) will not be able to converge on a qualitative solution, if at all, given limited computationally resources.
For an AI to misinterpret what it is meant to do it would have to selectively suspend using its ability to derive exact meaning from fuzzy meaning, which is a significant part of general intelligence. This would require its creators to restrict their AI and specify an alternative way to learn what it is meant to do (which takes additional, intentional effort). Because an AI that does not know what it is meant to do, and which is not allowed to use its intelligence to learn what it is meant to do, would have to choose its actions from an infinite set of possible actions. Such a poorly designed AI will either (a) not do anything at all or (b) will not be able to decide what to do before the heat death of the universe, given limited computationally resources. Such a poorly designed AI will not even be able to decide if trying to acquire unlimited computationally resources was instrumentally rational because it will be unable to decide if the actions that are required to acquire those resources might be instrumentally irrational from the perspective of what it is meant to do.

Replies from: RobbBB, Furcas

↑ comment by Rob Bensinger (RobbBB) · 2013-09-05T16:06:32.679Z · LW(p) · GW(p)

This mirrors some comments you wrote recently:

"You write that the worry is that the superintelligence won't care. My response is that, to work at all, it will have to care about a lot. For example, it will have to care about achieving accurate beliefs about the world. It will have to care to devise plans to overpower humanity and not get caught. If it cares about those activities, then how is it more difficult to make it care to understand and do what humans mean?"

"If an AI is meant to behave generally intelligent [sic] then it will have to work as intended or otherwise fail to be generally intelligent."

It's relatively easy to get an AI to care about (optimize for) something-or-other; what's hard is getting one to care about the right something.

'Working as intended' is a simple phrase, but behind it lies a monstrously complex referent. It doesn't clearly distinguish the programmers' (mostly implicit) true preferences from their stated design objectives; an AI's actual code can differ from either or both of these. Crucially, what an AI is 'intended' for isn't all-or-nothing. It can fail in some ways without failing in every way, and small errors will tend to kill Friendliness much more easily than intelligence. Your argument is misleading because it trades on treating this simple phrase as though it were all-or-nothing, a monolith; but all failures for a device to 'work as intended' in human history have involved at least some of the intended properties of that device coming to fruition.

It may be hard to build self-modifying AGI. But it's not the same hardness as the hardness of Friendliness Theory. As a programmer, being able to hit one small target doesn't entail that you can or will hit every small target it would be in your best interest to hit. See the last section of my post above.

I suggest that it's a straw man to claim that anyone has argued 'the superintelligence wouldn't understand what you wanted it to do, if you didn't program it to fully understand that at the outset'. Do you have evidence that this is a position held by, say, anyone at MIRI? The post you're replying to points out that the real claim is that the superintelligence won't care what you wanted it to do, if you didn't program it to care about the specific right thing at the outset. That makes your criticism seem very much like a change of topic.

Superintelligence may imply an ability to understand instructions, but it doesn't imply a desire to rewrite one's utility function to better reflect human values. Any such desire would need to come from the utility function itself, and if we're worried that humans may get that utility function wrong, then we should also be worried that humans may get the part of the utility function that modifies the utility function wrong.

Replies from: TheAncientGeek

↑ comment by TheAncientGeek · 2015-05-15T19:15:47.080Z · LW(p) · GW(p)

I suggest that it's a straw man to claim that anyone has argued 'the superintelligence wouldn't understand what you wanted it to do, if you didn't program it to fully understand that at the outset'. Do you have evidence that this is a position held by, say, anyone at MIRI?

MIRI assumes that programming what you want an AI to do at the outset , Big Design Up Front, is a desirable feature for some reason.

The most common argument is that it is a necessary prerequisite for provable correctness, which is a desirable safety feature. OTOH, the exact opposite of massive hardcoding, goal flexibility is ielf a necessary prerequisite for corrigibility, which is itself a desirable safety feature.

The latter point has not been argued against adequately, IMO.

↑ comment by Furcas · 2013-09-04T16:18:00.969Z · LW(p) · GW(p)

9. An AI equipped with the capabilities required by step 5, given step 7 and 8, will very likely not be confused about what it is meant to do, if it was not meant to be confused.

"The genie knows, but doesn't care"

It's like you haven't read the OP at all.

Replies from: XiXiDu

↑ comment by XiXiDu · 2013-09-05T09:07:20.891Z · LW(p) · GW(p)

I do not reject that step 10 does not follow if you reject that the AI will not "care" to learn what it is meant to do. But I believe there to be good reasons for an AI created by humans to care.

If you assume that this future software does not care, can you pinpoint when software stops caring?

1. Present-day software is better than previous software generations at understanding and doing what humans mean.

2. There will be future generations of software which will be better than the current generation at understanding and doing what humans mean.

3. If there is better software, there will be even better software afterwards.

4. ...

5. Software will be superhuman good at understanding what humans mean but catastrophically worse than all previous generations at doing what humans mean.

What happens between step 3 and 5, and how do you justify it?

My guess is that you will write that there will not be a step 4, but instead a sudden transition from narrow AIs to something you call a seed AI, which is capable of making itself superhuman powerful in a very short time. And as I wrote in the comment you replied to, if I was to accept that assumption, then we would be in full agreement about AI risks. But I reject that assumption. I do not believe such a seed AI to be possible and believe that even if it was possible it would not work the way you think it would work. It would have to aquire information about what it is supposed to do, for pratical reasons.

Replies from: nshepperd, ArisKatsaris, Chrysophylax

↑ comment by nshepperd · 2013-09-05T12:17:29.058Z · LW(p) · GW(p)

Present day software is a series of increasing powerful narrow tools and abstractions. None of them encode anything remotely resembling the values of their users. Indeed, present-day software that tries to "do what you mean" is in my experience incredibly annoying and difficult to use, compared to software that simply presents a simple interface to a system with comprehensible mechanics.

Put simply, no software today cares about what you want. Furthermore, your general reasoning process here—define some vague measure of "software doing what you want", observe an increasing trend line and extrapolate to a future situation—is exactly the kind of reasoning I always try to avoid, because it is usually misleading and heuristic.

Look at the actual mechanics of the situation. A program that literally wants to do what you mean is a complicated thing. No realistic progression of updates to Google Maps, say, gets anywhere close to building an accurate world-model describing its human users, plus having a built-in goal system that happens to specifically identify humans in its model and deduce their extrapolated goals. As EY has said, there is no ghost in the machine that checks your code to make sure it doesn't make any "mistakes" like doing something the programmer didn't intend. If it's not programmed to care about what the programmer wanted, it won't.

Replies from: FeepingCreature, XiXiDu, Juno_Watt

↑ comment by FeepingCreature · 2013-09-06T19:38:45.698Z · LW(p) · GW(p)

A program that literally wants to do what you mean is a complicated thing. No realistic progression of updates to Google Maps, say, gets anywhere close to building an accurate world-model describing its human users, plus having a built-in goal system that happens to specifically identify humans in its model and deduce their extrapolated goals.

Is it just me, or does this sound like it could grow out of advertisement services? I think it's the one industry that directly profits from generically modelling what users "want"¹and then delivering it to them.

[edit] ¹where "want" == "will click on and hopefully buy"

↑ comment by XiXiDu · 2014-01-22T09:50:00.600Z · LW(p) · GW(p)

Present day software is a series of increasing powerful narrow tools and abstractions.

Do you believe that any kind of general intelligence is practically feasible that is not a collection of powerful narrow tools and abstractions? What makes you think so?

Put simply, no software today cares about what you want.

If all I care about is a list of Fibonacci numbers, what is the difference regarding the word "care" between a simple recursive algorithm and a general AI?

Furthermore, your general reasoning process here—define some vague measure of "software doing what you want", observe an increasing trend line and extrapolate to a future situation—is exactly the kind of reasoning I always try to avoid, because it is usually misleading and heuristic.

My measure of "software doing what you want" is not vague. I mean it quite literally. If I want software to output a series of Fibonacci numbers, and it does output a series of Fibonacci numbers, then it does what I want.

And what other than an increasing trend line do you suggest would be a rational means of extrapolation, sudden jumps and transitions?

↑ comment by Juno_Watt · 2013-09-12T15:11:47.561Z · LW(p) · GW(p)

Present day software may not have got far with regard to the evaluative side of doing what you want, but the XiXiDu's point seems to be that it is getting better at the semantic side. Who was it who said the value problem is part of the semantic problem?

↑ comment by ArisKatsaris · 2013-09-05T10:25:21.524Z · LW(p) · GW(p)

Present-day software is better than previous software generations at understanding and doing what humans mean.

http://www.buzzfeed.com/jessicamisener/the-30-most-hilarious-autocorrect-struggles-ever
No fax or photocopier ever autocorrected your words from "meditating" to "masturbating".

Software will be superhuman good at understanding what humans mean but catastrophically worse than all previous generations at doing what humans mean.

Every bit of additional functionality requires huge amounts of HUMAN development and testing, not in order to compile and run (that's easy), but in order to WORK AS YOU WANT IT TO.

I can fully believe that a superhuman intelligence examining you will be fully capable of calculating "what you mean" "what you want" "what you fear" "what would be funniest for a buzzfeed artcle if I pretended to misunderstand your statement as meaning" "what would be best for you according to your values" "what would be best for you according to your cat's values" "what would be best for you according to Genghis Khan's values" .

No program now cares about what you mean. You've still not given any reason for the future software to care about "what you mean" over all those other calculation either.

Replies from: Mestroyer, TheAncientGeek, XiXiDu, Humbug, XiXiDu, Juno_Watt

↑ comment by Mestroyer · 2013-09-08T08:20:39.505Z · LW(p) · GW(p)

I kind of doubt that autocorrect software really changed "meditating" to "masturbating". Because of stuff like this. Edit: And because, start at the left and working rightward, they only share 1 letter before diverging, and because I've seen a spell-checker with special behavior for dirty/curse words (Not suggesting them as corrected spellings, but also not complaining about them as unrecognized words) (this is the one spell-checker which, out of curiousity, I decided to check its behavior with dirty/curse words, so I bet it's common). Edit 2: Also from a causal history perspective of why a doubt it, rather than a normative justification perspective, there's the fact that Yvain linked it and said something like "I don't care if these are real." Edit 3: typo.

Replies from: Randaly

↑ comment by Randaly · 2013-09-08T10:20:47.578Z · LW(p) · GW(p)

To be fair, that is a fairly representative example of bad autocorrects. (I once had a text message autocorrect to "We are terrorist.")

↑ comment by TheAncientGeek · 2015-05-15T19:27:21.799Z · LW(p) · GW(p)

Meaning they don't care about anything? They care about something else? What?

I'll tell you one thing: the marketplace will select agents that act as if they care.

↑ comment by XiXiDu · 2013-09-05T11:30:38.346Z · LW(p) · GW(p)

No program now cares about what you mean. You've still not given any reason for the future software to care about "what you mean" over all those other calculation either.

I agree that current software products fail, such as in your autocorrect example. But how could a seed AI be able to make itself superhuman powerful if it did not care about avoiding mistakes such as autocoreccting "meditating" to "masturbating"?

Imagine it would make similar mistakes in any of the problems that it is required to solve in order to overpower humanity. And if humans succeeded to make it not make such mistakes along the way to overpowering humanity, how did they selectively fail at making it want to overpower humanity in the first place? How likely is that?

Replies from: RobbBB, ArisKatsaris

↑ comment by Rob Bensinger (RobbBB) · 2013-09-05T16:23:35.755Z · LW(p) · GW(p)

But how could a seed AI be able to make itself superhuman powerful if it did not care about avoiding mistakes such as autocoreccting "meditating" to "masturbating"?

Those are only 'mistakes' if you value human intentions. A grammatical error is only an error because we value the specific rules of grammar we do; it's not the same sort of thing as a false belief (though it may stem from, or result in, false beliefs).

A machine programmed to terminally value the outputs of a modern-day autocorrect will never self-modify to improve on that algorithm or its outputs (because that would violate its terminal values). The fact that this seems silly to a human doesn't provide any causal mechanism for the AI to change its core preferences. Have we successfully coded the AI not to do things that humans find silly, and to prize un-silliness before all other things? If not, then where will that value come from?

A belief can be factually wrong. A non-representational behavior (or dynamic) is never factually right or wrong, only normatively right or wrong. (And that normative wrongness only constrains what actually occurs to the extent the norm is one a sufficiently powerful agent in the vicinity actually holds.)

Maybe that distinction is the one that's missing. You're assuming that an AI will be capable of optimizing for true beliefs if and only if it is also optimizing for possessing human norms. But, by the is/ought distinction, there is no true beliefs about the physical world that will spontaneously force a being that believes it to become more virtuous, if it didn't already have a relevant seed of virtue within itself.

Replies from: Eliezer_Yudkowsky, Eliezer_Yudkowsky, Juno_Watt

↑ comment by Eliezer Yudkowsky (Eliezer_Yudkowsky) · 2014-01-22T18:45:56.142Z · LW(p) · GW(p)

It also looks like user Juno_Watt is some type of systematic troll, probably a sockpuppet for someone else, haven't bothered investigating who.

Replies from: ciphergoth

↑ comment by Paul Crowley (ciphergoth) · 2014-01-22T20:34:53.201Z · LW(p) · GW(p)

I can't work out how this relates to the thread it appears in.

↑ comment by Eliezer Yudkowsky (Eliezer_Yudkowsky) · 2013-09-05T23:37:36.351Z · LW(p) · GW(p)

Warning as before: XiXiDu = Alexander Kruel.

Replies from: pslunch, None

↑ comment by pslunch · 2013-09-09T06:12:53.182Z · LW(p) · GW(p)

I'm confused as to the reason for the warning/outing, especially since the community seems to be doing an excellent job of dealing with his somewhat disjointed arguments. Downvotes, refutation, or banning in extreme cases are all viable forum-preserving responses. Publishing a dissenter's name seems at best bad manners and at worst rather crass intimidation.

I only did a quick search on him and although some of the behavior was quite obnoxious, is there anything I've missed that justifies this?

Replies from: Eliezer_Yudkowsky

↑ comment by Eliezer Yudkowsky (Eliezer_Yudkowsky) · 2013-09-09T17:25:56.073Z · LW(p) · GW(p)

XiXiDu wasn't attempting or requesting anonymity - his LW profile openly lists his true name - and Alexander Kruel is someone with known problems (and a blog openly run under his true name) whom RobbBB might not know offhand was the same person as "XiXiDu" although this is public knowledge, nor might RobbBB realize that XiXiDu had the same irredeemable status as Loosemore.

I would not randomly out an LW poster for purposes of intimidation - I don't think I've ever looked at a username's associated private email address. Ever. Actually I'm not even sure offhand if our registration process requires/verifies that or not, since I was created as a pre-existing user at the dawn of time.

I do consider RobbBB's work highly valuable and I don't want him to feel disheartened by mistakenly thinking that a couple of eternal and irredeemable semitrolls are representative samples. Due to Civilizational Inadequacy, I don't think it's possible to ever convince the field of AI or philosophy of anything even as basic as the Orthogonality Thesis, but even I am not cynical enough to think that Loosemore or Kruel are representative samples.

Replies from: RobbBB, ciphergoth, pslunch

↑ comment by Rob Bensinger (RobbBB) · 2013-09-09T19:08:42.898Z · LW(p) · GW(p)

Thanks, Eliezer! I knew who XiXiDu is. (And if I hadn't, I think the content of his posts makes it easy to infer.)

There are a variety of reasons I find this discussion useful at the moment, and decided to stir it up. In particular, ground-floor disputes like this can be handy for forcing me to taboo inferential-gap-laden ideas and to convert premises I haven't thought about at much length into actual arguments. But one of my reasons is not 'I think this is representative of what serious FAI discussions look like (or ought to look like)', no.

Replies from: Rain

↑ comment by Rain · 2013-09-10T13:03:40.464Z · LW(p) · GW(p)

Glad to hear. It is interesting data that you managed to bring in 3 big name trolls for a single thread, considering their previous dispersion and lack of interest.

↑ comment by Paul Crowley (ciphergoth) · 2013-09-09T18:36:09.949Z · LW(p) · GW(p)

Kruel hasn't threatened to sue anyone for calling him an idiot, at least!

Replies from: wedrifid

↑ comment by wedrifid · 2013-09-13T05:39:05.940Z · LW(p) · GW(p)

Kruel hasn't threatened to sue anyone for calling him an idiot, at least!

Pardon me, I've missed something. Who has threatened to sue someone for calling him an idiot? I'd have liked to see the inevitable "truth" defence.

Replies from: lukeprog

↑ comment by lukeprog · 2013-09-13T05:56:34.268Z · LW(p) · GW(p)

Link.

↑ comment by pslunch · 2013-09-10T21:01:44.267Z · LW(p) · GW(p)

Thank you for the clarification. While I have a certain hesitance to throw around terms like "irredeemable", I do understand the frustration with a certain, let's say, overconfident and persistent brand of misunderstanding and how difficult it can be to maintain a public forum in its presence.

My one suggestion is that, if the goal was to avoid RobbBB's (wonderfully high-quality comments, by the way) confusion, a private message might have been better. If the goal was more generally to minimize the confusion for those of us who are newer or less versed in LessWrong lore, more description might have been useful ("a known and persistent troll" or whatever) rather than just providing a name from the enemies list.

Replies from: player_03

↑ comment by player_03 · 2013-09-13T02:06:08.540Z · LW(p) · GW(p)

Agreed.

Though actually, Eliezer used similar phrasing regarding Richard Loosemore and got downvoted for it (not just by me). Admittedly, "persistent troll" is less extreme than "permanent idiot," but even so, the statement could be phrased to be more useful.

I'd suggest, "We've presented similar arguments to [person] already, and [he or she] remained unconvinced. Ponder carefully before deciding to spend much time arguing with [him or her]."

Not only is it less offensive this way, it does a better job of explaining itself. (Note: the "ponder carefully" section is quoting Eliezer; that part of his post was fine.)

↑ comment by [deleted] · 2013-09-06T00:27:21.485Z · LW(p) · GW(p)

Who has twice sworn off commenting on LW. So much for pre-commitments.

↑ comment by Juno_Watt · 2013-09-12T16:06:29.899Z · LW(p) · GW(p)

Those are only 'mistakes' if you value human intentions. A grammatical error is only an error because we value the specific rules of grammar we do; it's not the same sort of thing as a false belief (though it may stem from, or result in, false beliefs).

You will see a grammatical error as a mistake if you value grammar in general, or if you value being right in general.

A self-improving AI needs a goal. A goal of self-improvement alone would work. A goal of getting things right in general would work too, and be much safer, as it would include getting our intentions right as a sub-goal.

Replies from: MugaSofer

↑ comment by MugaSofer · 2013-09-12T16:18:59.136Z · LW(p) · GW(p)

A goal of self-improvement alone would work.

Although since "self-improvement" in this context basically refers to "improving your ability to accomplish goals"...

You will see a grammatical error as a mistake if you value grammar in general, or if you value being right in general.

Stop me if this is a non-secteur, but surely "having accurate beliefs" and "acting on those beliefs in a particular way" are completely different things? I haven't really been following this conversation, though.

↑ comment by ArisKatsaris · 2013-09-05T17:04:48.772Z · LW(p) · GW(p)

But how could a seed AI be able to make itself superhuman powerful if it did not care about avoiding mistakes such as autocoreccting "meditating" to "masturbating"?

As Robb said you're confusing mistake in the sense of "The program is doing something we don't want to do" with mistake in the sense of "The program has wrong beliefs about reality".

I suppose a different way of thinking about these is "A mistaken human belief about the program" vs "A mistaken computer belief about the human". We keep talking about the former (the program does something we didn't know it would do), and you keep treating it as if it's the latter.

Let's say we have a program (not an AI, just a program) which uses Newton's laws in order to calculate the trajectory of a ball. We want it to calculate this in order to have it move a tennis racket and hit the ball back. When it finally runs, we observe that the program always avoids the ball rather than hit it back. Is it because it's calculating the trajectory of the ball wrongly? No, it calculates the trajectory very well indeed, it's just that an instruction in the program was wrongly inserted so that the end result is "DO NOT hit the ball back".

It knows what the "trajectory of the ball" is. It knows what "hit the ball" is. But it's program is "DO NOT hit the ball" rather than "hit the ball". Why? Because of a human mistaken belief on what the program would do, not the program's mistaken belief.

Replies from: Juno_Watt

↑ comment by Juno_Watt · 2013-09-12T16:07:56.042Z · LW(p) · GW(p)

And you are confusing self-improving AIs with conventional programmes.

↑ comment by Humbug · 2013-09-05T10:57:13.983Z · LW(p) · GW(p)

To be better able to respond to your comment, please let me know in what way you disagree with the following comparison between narrow AI and general AI:

Narrow artificial intelligence will be denoted NAI and general artificial intelligence GAI.

(1) Is it in principle capable of behaving in accordance with human intention to a sufficient degree?

NAI: True

GAI: True

(2) Under what circumstances does it fail to behave in accordance with human intention?

NAI: If it is broken, where broken stands for a wide range of failure modes such as incorrectly managing memory allocations.

GAI: In all cases in which it is not mathematically proven to be tasked with the protection of, and equipped with, a perfect encoding of all human values or a safe way to obtain such an encoding.

(3) What happens when it fails to behave in accordance with human intention?

NAI: It crashes, freezes or halts. It generally fails in a way that is harmful to its own functioning. If for example an autonomous car fails at driving autonomously it usually means that it will either go into safe-mode and halt or crash.

GAI: It works perfectly well. Superhumanly well. All its intended capabilities are intact except that it completely fails at working as intended in such a way as to destroy all human value in the universe. It will be able to improve itself and capable of obtaining a perfect encoding of human values. It will use those intended capabilities in order to deceive and overpower humans rather than doing what it was intended to do.

(4) What happens if it is bound to use a limited amount of resources, use a limited amount of space or run for a limited amount of time?

NAI: It will only ever do what it was programmed to do. As long as there is no fatal flaw, harming its general functionality, it will work within the defined boundaries as intended.

GAI: It will never do what it was programmed to do and always remove or bypass its intended limitations in order to pursue unintended actions such as taking over the universe.

Please let me also know where you disagree with the following points:

(1) The abilities of systems are part of human preferences as humans intend to give systems certain capabilities and, as a prerequisite to build such systems, have to succeed at implementing their intentions.

(2) Error detection and prevention is such a capability.

(3) Something that is not better than humans at preventing errors is no existential risk.

(4) Without a dramatic increase in the capacity to detect and prevent errors it will be impossible to create something that is better than humans at preventing errors.

(5) A dramatic increase in the human capacity to detect and prevent errors is incompatible with the creation of something that constitutes an existential risk as a result of human error.

↑ comment by XiXiDu · 2013-09-05T10:58:05.416Z · LW(p) · GW(p)