📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Elon Musk's approach to problem-solving | Lex Fridman Podcast

Lex Clips8:49

Transcription

Can you just speak to what it takes for a great engineering team? For you, the what I've saw in Memphis, the supercomputer cluster is just this intense drive towards simplifying the process, understanding the process, constantly improving it, constantly iterating it. Well, it's easy to say "simplify," and it's very difficult to do it.

Um, you know, I have this very basic, first basic first principles algorithm that I run kind of as like a mantra, which is: The first question, the requirements. Make the requirements, um, less dumb. The requirements are always dumb to some degree. So, if you want to sort off by reducing the number of requirements, um, and, um, no matter how smart the person who gave you those requirements, that's still dumb to some degree. Um, if you, you have to start there because otherwise, uh, you could get the perfect answer to the wrong question. So, so try to make the question the least wrong possible. That's what, um, "question the requirements" means.

And then the second thing is, try to delete the whatever the step is, the the part or the process step. Um, sounds very obvious, but, um, people often forget to do to to try deleting it entirely. And if you're not forced to put back at least 10% of what you delete, you're not deleting enough. Like, so, and it's, uh, somewhat illogically, people often, most of the time, um, feel as though they've succeeded if they've not been forced to put the put things back in. But actually, they haven't because they've been overly conservative and and have left things in there that shouldn't be. So.

And only the third thing is, try to optimize it or simplify it. Um, again, this sounds, these all sound, I think, very, very obvious when I say them, but, uh, the number of times I've made these mistakes is, uh, more than I care to remember. Um, that's why I have this mantra. So, in fact, I'd say the the most common mistake of smart engineers is to optimize a thing that should not exist. Right? So, so you, like, like you say, you run through the algorithm, yeah, and basically show, show up to a problem, uh, show up to the the the supercomputer cluster and see the process and ask, "Can this be deleted?" Yeah. First, try to delete it.

Um, yeah, yeah, that's not easy to do. No. And and actually, this is what what generally makes people uneasy is that you've got to delete at least some of the things that you delete, you will put back in. Yeah. But going back to sort of where our Olympic system can steer us wrong is that, um, we tend to remember, uh, with sometimes a jarring level of pain, uh, where we, where we deleted something that that we subsequently needed. Yeah. Um, and so people remember that one time they forgot to put in this thing three years ago, and that caused them trouble. Um, and so they overcorrect and then they put too much stuff in there and overcomplicate things. So you actually have to say, "No, we're deliberately going to delete more than we, we should," so that we're putting at least one in 10 things we're going to add back in. And and I've seen you suggest just that, that something should be deleted, and you can kind of see the the pain.

Oh, yeah, absolutely. Everybody feels a little bit of the pain. Absolutely. And and I tell them in advance, like, "Yeah, there's some of the things that we delete, we're going to put back in." And and that people get a little shook by that. Um, but it makes sense because if you, if you're so conservative as to never have to put anything back in, you obviously have a lot of stuff that isn't needed. So you got to overcorrect. This is, I would say, like a quot override to Olympic instinct, one of many that probably leads us astray. Yeah.

Um, there's like a step four as well, which is, um, any given thing can be sped up. How a fast you think it can be done, like whatever the speed, the speed is being done, it can be done faster. But but you shouldn't speed things up until it's until you try to delete it and optimize it, otherwise you're speeding up something that speeding up something that shouldn't exist as aod. Um, and then and then the the fifth thing is to to automate it. Yeah. And I've gone backwards so many times where I've automated something, sped it up, simplified it, and then deleted it, and I got tired of doing that. So that's why I've got this mantra.

That is a very effective five-step process. It's works great. Well, when you've already automated, deleting must be real painful. Yeah. It's like, it's like, "Wow, I really wasted a lot of effort there." Yeah. I mean, what you've done, uh, with the, with the cluster in Memphis is incredible, just in a handful of weeks. Yeah. It's not working yet, so I want to pop the champagne corks. Um, in fact, I, I have a, a call in a few hours with the Memphis team, um, because we're having some power fluctuation issues.

So, uh, yeah, it's like, kind of a, there's a, when you do synchronized training, when you, these computers that are training, uh, that where the training is synchronized to, you know, at the sort of millisecond level, uh, you, it's like having an orchestra, and then the the orchestra can go loud to silent very quickly, you know, at a sub-second level. And then the the electrical system kind of freaks out about that. Like, if you suddenly see giant shifts, 10, 20 megawats several times a second, uh, this is not what electrical systems are expecting to see. So that's one of the main things you have to figure out, the cooling, the power, the, uh, and then on the software, as you go up the stack, how to do the the distributed computer, all that, all that. Today's problem is dealing with with with with extreme power jitter. Power jitter. Yeah. It's a nice ring to that. So that's okay.

And you stayed up late into the night, as you often do there, last week? Last week, yeah. Yeah. We finally, finally got, got training going at, uh, oddly enough, roughly 4:20 a.m. uh, last Monday. Total coincidence. Yeah. I mean, maybe it was 4:22 or something. Yeah. Yeah. It's that universe again with the jokes. Ex, just love it.

I mean, I wonder if you could speak to the the fact that you, one of the things, uh, that you did when I was there is you went through all the steps of what everybody's doing, just to get a sense that you yourself understand it, and, uh, everybody understands it, so they can understand when something is dumb, or some something is inefficient, or that stuff. Can you speak to that?

Yeah, so I, like, like I try to do whatever the people at the front lines are doing, I try to do it at least a few times myself. So connecting fiber optic cables, diagnosing a PI connection, that tends to be the limiting factor for large training clusters is the cabling. There are so many cables. Um, because for for a coherent training system where you've got, um, RDMA, remote, so remote direct memory access, the the whole thing is like one giant brain. So if you've got, um, any to any connection, so it's the the any GPU can talk to any GPU out of 100,000. That is a that is a crazy cable layout. It looks pretty cool. Yeah. It's like it's like the human brain, but like at a scale that humans can visibly see. It is a brain. I mean, the human brain also has a massive amount of the brain tissue is the the cables. Yeah. So they get the gray matter, which is the compute, and then the white matter, which is cables. Big percentage of your brain is just cables. That's what it felt like walking around in the supercomputer center is like, we're walking around inside the brain. Will one day build a super intelligent, super super intelligent system.