📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

🚩OpenAI Safety Team "LOSES TRUST" in Sam Altman and gets disbanded. The "Treacherous Turn".

Wes Roth•18:25

Transcription

So, while OpenAI is doing an incredible job of announcing new products, revealing new capabilities, at the same time, there's some dark clouds brewing over AI safety at OpenAI. As we covered the other day, Ilia Sutskever leaves OpenAI; they decided to part ways, but no one's really talking too much about it. At the same time, Jan Leike leaves; says he resigns, and today he posts this. And it starts out like you would expect it to: he says it's been so fun and thank you, it's been a wild journey—kind of the standard stuff, boilerplate stuff that everybody says. But then he goes off script.

So, here he's saying all the stuff you normally say, wishing everybody the best, thanking people, and then the tone changes. He's saying, "Stepping away from this job has been one of the hardest things I have ever done because we urgently need to figure out how to steer and control AI systems much smarter than us. I joined because I thought OpenAI would be the best place in the world to do this research; however, I have been disagreeing with OpenAI leadership about the company's core priorities for quite some time, until we finally reached a breaking point. I believe much more of our bandwidth should be spent getting ready for the next generation of models on security, monitoring, preparedness, safety, adversarial robustness, super alignment, confidentiality, societal impact, and related topics." He's saying these problems are quite hard to get right, and he's concerned we aren't on a trajectory to get there. Over the past few months, his team has been sailing against the wind; sometimes they were struggling for compute, and it was getting harder and harder to get this crucial research done. Building smarter-than-human machines is an inherently dangerous endeavor; OpenAI is shouldering an enormous responsibility on behalf of all of humanity. But over the past years, safety culture and processes have taken a back seat to shiny products. We are long overdue in getting incredibly serious about the implications of AGI; we must prioritize preparing for them as best we can. Only then can we ensure AGI benefits all of humanity. OpenAI must become a safety-first AGI company. To all OpenAI employees, he wants to say, "Learn to feel the AGI; act with the gravitas appropriate for what you're building. I believe you can ship the cultural change that's needed; I'm counting on you; the world is counting on you, OpenAI."

Heart. So, first of all, whoa. AGI rolls around only once. Subscribe. This is very, very different to me than anything that kind of came before this. Before this, we had some hints and rumors and people talking in the background, but everyone was kind of tight-lipped. I'm sure there are contracts, non-disclosure agreements, various maybe vesting schedules that people don't want to lose; whatever the case, no one really talked about anything. Elon Musk sues OpenAI, as I saw, trying to get some of the documents to be read into the record, to be shown in front of a jury regarding things like Q-Star. What's happening behind the scenes? Here's a picture of Ilia Sutskever; he, of course, officially and finally parted ways with OpenAI. Here's Jan Leike. So this is all happening kind of in real time, and Vox.com just published today, "Why the OpenAI team in charge of safeguarding humanity imploded." Company insiders explain why safety-conscious employees are leaving, and many of them are more than we talked about. There's Helen Toner, the ex-board member that we believe is kind of responsible for that coup that happened in November, the firing of Sam Altman. Several AI safety researchers at OpenAI were fired for leaking information, for example, the leaked Q-Star details, or lack of details, but just the idea of that project existing—that was confirmed to be a real leak of a real project, but no further details were given.

Now, this gets much deeper, by the way, because we are just now getting some new information about kind of what's happening on the inside and why people are being kind of tight-lipped about it. But first, here's Sam Altman. So he's responding to Jan Leike's comment, the post that he just created five hours ago, saying that he's disagreeing with Sam Altman and the leadership, as he calls it, and kind of pretty clearly saying that he doesn't believe enough safety precautions are being taken, and the OpenAI employees should be very careful about how they're going to proceed with this. Sam responds, "I'm super appreciative of Jan Leike's contributions to OpenAI's alignment research and safety culture and very sad to see him leave. He's right, we have a lot more to do; we are committed to doing it. I'll have a longer post in the next couple, couple of days." We'll be on the lookout for that; hopefully we'll get more clarity into what specifically the issue is here, because Jan Leike, I mean, specifically said there's not enough compute—that was one of the complaints—but it feels like there's more going on, and I'd love to know exactly what.

Now, this was a red post from a while back, but before we take a look at it, here's the very important and kind of annoying thing to understand: the thing to understand is that right now we're living in the moment that AI politics are going mainstream; more and more people are taking sides, and it is just like politics—you have your little tribes and different sides of the issue yelling back and forth; no one sees eye to eye; there's less and less agreement; everyone's getting more and more polarized. So, as we read some of this stuff, keep in mind that we're beginning to get away from, you know, open-minded people discussing ideas and into this realm of highly politicized, polarizing arguments with more and more people joining in that may not be equipped to understand all the intricacies of, you know, AI alignment. As Andrew here says, "I don't think it's possible to explain the problems with alignment to the public without driving a bunch of people insane." I think this is very well said. The treacherous churn, in particular, is going to put a real bug in some people's ears. He's referring to this idea, a hypothetical event where an advanced AI system, which has been pretending to be aligned due to its relative weakness, turns on humanity once it achieves sufficient power, so much so that it can pursue its true objectives without risk. But I would agree with this response that I think most people won't even get that far and will be more preoccupied with pretty trivial things; most people will not have enough knowledge to argue coherently about this. But my point is, as we read some of the stuff, just keep in mind that a lot of this is just people's opinions; you don't have to think of them as right or wrong or even react to them in your way; just kind of think about where this whole thing is going.

So, here's that red post from a while back, saying, "Any company that makes AGI is going to want to feed it as many GPUs as money can buy while delaying having to announce AGI. They've now changed from a customer-facing company to a ninja throwing smoke bombs in order to throw people off the scent. They're going to want to release a bunch of amazing new products and make random cryptic statements to keep people guessing for as long as possible. Their actions will start to seem more and more chaotic and unnecessarily obtuse; customers will be happy but frustrated. They will start to release products that are unreasonably better than they should be, with unclear paths to their creation. There will be sudden breakdowns in staff loyalty and communications, firings, resignations, vague hints from people under NDAs." By the way, all those things we've seen: we've seen products that are unreasonably better than they should be. I'm thinking of Sora. I mean, technically it's not released, so maybe when it comes out we'll see that it was just not as good as we expected it, but as we've covered in this channel, the 3D sort of simulation of the physics of the fluid movements in some of those videos, they seem to be unreasonably better than they should be. Dr. Jim Fan from Nvidia has talked about it quite a bit. I mean, take a look at this Sora-produced Minecraft video. This isn't Minecraft; this isn't a 3D game; this isn't something running on an Nvidia graphics card; this is a text-to-video generation from Sora. This coffee cup with this pirate battleship that's going on in there, right? That's coffee swirling around in a cup; it's very difficult to produce. This is something that a lot of people spend a lot of time for video game development trying to create those fluid physics, and if Sora is released and we see that it easily generates it on the fly, certainly that would be surprising, I think. Then they're saying, "One day soon after, the military will suddenly take a large interest in OpenAI; the company will go quiet." Now this is a hypothetical scenario that they're talking about, but in November it was kind of surprising how quickly, I believe it was the Attorney General of New York Southern District, that was on the phone with Helen Toner and the other board members trying to settle the dispute—all the other very powerful people in the tech space who contributed to coming to the table and working things out. Next thing we know is the board of OpenAI is populated with people with close ties to the US government, people that have been very closely tied to the US government for, for, for decades. And as we'll see in a second, as this Vox article says, based on some leaks from inside the company, "Why OpenAI's safety team grew to distrust Sam Altman." Ilia Sutskever posted back in December 6th, right after that coup, he said, "I learned many lessons this past month; one such lesson is that the phrase 'the beatings will continue until morale improves' applies more often than it has any right to."

Now, before we go down that path, it is important to understand that there's an ideological, some would say movement, behind this pause AI, right? Here's a list of noteworthy people in tech space, a lot of well-known AI researchers and their P(Doom) values. So P(Doom), if you're not aware, is their estimation of what they think the chance of a catastrophic event due to AI is—something that would, for example, human extinction, right? Erase humanity, right? And we have various estimates from very, very low to some people—Elijah Yudkowsky notoriously, right?—saying greater than 99%. And here's Dr. Techlash; we've covered her before. So she's saying, "Remember how Jan Leike was a research associate at both Eliezer Yudkowsky's Machine Intelligence Research Institute and Nick Bostrom's Future of Humanity Institute? There's clearly an ideological influence at play here." And of course we see Jan Leike here, former alignment lead at OpenAI; his P(Doom) is 10 to 90%, so fairly wide range, I would say. And of course we've heard from Daniel Kahneman, also former OpenAI researcher, who also said some things that maybe weren't so positive for OpenAI's Safety Research team. Elon Musk is on here at 10 to 20% of catastrophic AI outcomes; Yann LeCun of Meta AI at less than 0.1%; Vitalik Buterin on here, Ethereum co-founder, also a person that funded a lot of these PAI efforts; he donated to some of the research teams behind AI safety efforts; he thinks it's 10%; Jeff Hinton is also at 10%; Lina Khan here, head of the FTC, she's at 15%; Dario Amodei at 10 to 25%; he is the CEO of Anthropic; Yoshua Bengio 20%; Emad Mostaque, who was supposed to take over as the CEO of OpenAI, he thinks it's between 5 and 50%. And here's why some of those former people at OpenAI are concerned. This is from Vox.com; I'll link a link down below. So they're saying here, "If you've been following this whole saga on social media, you might think OpenAI secretly made a huge technological breakthrough." The meme, "What did Ilia see?" speculates that Sutskever, the former chief scientist, left because he saw something horrifying, like an AI system that could destroy humanity. And certainly we've heard rumors like this, or at least rumors of OpenAI having something big, some big breakthrough that potentially unsettled some of the people there, including Ilia. But the real answer may have less to do with pessimism about technology and more to do with pessimism about humans, and one human in particular, Altman. According to sources familiar with the company, safety-minded employees have lost faith in him. It's a process of trust collapsing bit by bit, like dominoes falling one by one. A person with inside knowledge of the company told me, speaking on condition of anonymity, "Not many employees are willing to speak about this publicly." That's partly because OpenAI is known for getting its workers to sign offboarding agreements with non-disparagement provisions upon leaving. If you refuse to sign one, you give up your equity in the company, which means you potentially lose out on millions of dollars, and maybe billions if you think about how much equity in that company could be worth in the future. One former employee, however, refused to sign the offboarding agreement so that he could be free to criticize the company—Daniel Kahneman, who joined OpenAI in 2022. He said, "OpenAI is training ever more powerful AI systems with the goal of eventually surpassing human intelligence across the board. This could be the best thing that has ever happened to humanity, but it could also be the worst if we don't proceed with care. I joined with the substantial hope that OpenAI would rise to the occasion, behave more responsibly as they got closer to achieving AGI, but it slowly became clear to many of us that this would not happen, and that forced them to quit."

So, of course, a lot of this happened in November. Last November, Helen Toner and Ilia Sutskever, working together with the OpenAI board, tried to fire Altman. The reason they gave is that Altman was not consistently candid in his communications, and they really didn't say too much more than that. A lot of things happened; Microsoft invited all of OpenAI's top talent to Microsoft, effectively destroying OpenAI, but basically allowing them to continue to build under the Microsoft umbrella. And Altman, of course, came back more powerful than ever, has a more supportive board, and more power to run the company how he sees fit. "When you shoot at kings and miss, things tend to get awkward," which is well said, and certainly that's what happened with Ilia Sutskever, who finally officially left OpenAI and said he was heading off to pursue a project that is very personally meaningful to him. One thing they mention here is that it looks like Ilia has been remotely co-leading the super alignment team tasked with making sure a future AGI would be aligned with the goals of humanity rather than going rogue. I actually was not aware of this, so it seems like he basically worked remotely but on the same team, on the same objectives. Now this article kind of skews heavily against Sam Altman, right? So they're saying what happened in November revealed something about Sam Altman's character; his threat to hollow out OpenAI unless the board rehired him and his insistence on stacking the board with new members skewed in his favor showed a determination to hold on to power and avoid future checks on it. So I'm not sure if this is true; I don't think this is real because I don't think there were any threats to hollow out OpenAI unless the board rehired him. I believe Microsoft, being shrewd business people, offered to swallow up all of OpenAI's talent, which of course they would; it's a smart decision, but I wouldn't call that Sam Altman threatening to hollow out OpenAI. So keep that in mind as we look over this; this article is very much kind of leaning against Sam Altman. And there's a number of other examples of OpenAI safety researchers making various cryptic posts, like this one on the EA forum, saying that they resigned from OpenAI on February 15th, 2024. When asked why, they replied, "No comment." And the reason why they'd reply "no comment" is this: there's a very restrictive offboarding agreement that contains non-disclosure and non-disparagement provisions; former OpenAI employees are subject to it; it forbids them for the rest, rest of their lives from criticizing their former employer; even acknowledging that the NDA exists is a violation of it. If a departing employee declines to sign the document or if they violate it, they can lose all vested equity they earn during their time at the company, which is likely worth millions of dollars. And here's a piece from Wired saying OpenAI's long-term AI risk team has disbanded last year. OpenAI said that the team would receive 20% of its computing power, but now that team, the super alignment team, is no more; the company confirms.

Now I'd love to know what you think, but keep in mind that each side will always just tell their story. The people that are aligned with the PAI, a lot of them seemingly have connections with EA, Effective Altruism, and some of the actions that they did were not 100% on the up and up; there were some shenanigans going on there as well. They have a certain ideological lean, and they're pursuing that; a lot of the AI safety people seem to share those views. Here's Rune, another employee at OpenAI, saying, "Everyone constantly believes they deserve more GPUs; it's basically a necessary feature of being a researcher." That was in fact one of the big complaints that a lot of the AI researchers, including Jan Leike, he posted saying we didn't have enough compute; we didn't have enough GPUs; basically they didn't give us enough computer power to do our research, and therefore we quit. Could that have been the case? Could it just be a case of not getting enough resources and looking elsewhere to get those resources to pursue their research projects? Let me know what you think, but keep this in mind: we're going to have more and more discussions like this in the world, on TV, on Twitter, on Facebook, as more and more of the world's population gets dragged into this conversation. Get ready for some pretty wild takes. But whatever the case, my name is Wes Rth, and thank you for watching.