📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

MIT Introduction to Deep Learning | 6.S191

Alexander Amini1:09:26

Transcription

Good afternoon, everyone, and uh, thank you for joining us today. My name is Alexander Amini, and together with Ava, we're going to be your instructors for the course this year. This is uh, MIT Introduction to Deep Learning, or 6.S191, is our official course title.

Now, we're super excited to welcome you to this class, and I think probably a good place to start is always, you know, asking ourselves, well, what is MIT Intro to Deep Learning? This is a one-week boot camp on everything deep learning, right? So it's both a very fun but also a very intense one week because we're going to cover a ton of material in just the next five days.

Now, this is our eighth year teaching this class, and the pace of the field, especially in the past couple of years, is really remarkable. And every time that we teach this class, it's getting more and more, I should say, interesting to introduce this lecture in particular, and how we introduced this lecture has really started to adapt and evolve over the many years. Now, many of you in the audience have probably even started to become almost a bit, I would say, desensitized to a lot of the progress of deep learning in the past couple of years because of this progress, how rapidly this progress is happening. So I think it's it's also important to not forget, you know, where we came from just a few years ago. So I want to show, you know, this this image right here just to start this off, and what better way to show you than for you to actually see the progress with your own eyes? Exactly one decade ago, this was the state of a state-of-the-art deep learning-based facial generation system. So this is not a real face; this was the state-of-the-art model that could generate faces, and this was the best that we could do.

Fast forward, you know, just a few years uh, down the pipeline, and progress in image generation had already started to advance tremendously. And here you can see, you know, a lot of more realism, photorealism in these types of images being created. And then, you know, fast forward another few years after that, and and these images start to come to life, right? They start to have temporal information; they start to have video and uh, you know, movement as well to those uh, to those images, right? And in fact, this this video that you see on the right is a video that we created in this class uh, some years ago. And for those of you who haven't seen it already, it's it's uh, you know, it's online, and people have seen it, but in case you haven't, I'll play just the first 10 seconds. I won't play the whole thing, just so you can see it as well.

"Hi everybody, and welcome to MIT 6S191, the official introductory course on deep learning taught here at MIT."

Now I won't play the whole thing, right? But you get the gist of this video. This this video was created five years ago; we made it as part of this class, and we we used it to actually introduce this class uh, back then. Now, when we did this in 2020, it got a lot of, even back then, right? Like, especially back then, maybe you guys aren't that impressed by it today, but back then this video was a huge deal, right? This was a huge jump in photorealism for uh, for for capabilities of deep learning models, and the clip went, you know, very uh, very rapidly a bit viral, and people commented a lot about the realism. But actually, one interesting thing that people didn't see—they saw the end result—but actually, at that point, what people didn't see was that for us to generate that clip, which was a two-minute clip—you only saw the first 10 seconds, but that clip was two minutes in total—and to generate that two-minute clip, it cost around two hours of professional audio data being recorded and captured uh, of the speaker, which was not us; it cost around 50 hours of professional high-def video data to build, you know, the face model; and it required around $155,000 US dollars of compute to generate that two-minute video. And all of that was just going to generate, you know, a predefined script, right? Something static; you couldn't talk to it; you couldn't interact with it; it was just a predefined script, static, not flexible at all. But uh, you know, still a tremendous amount of resources, both time, data, energy, and and financial resources as well.

And I wanted to start this class by basically asking this question of, you know, what would this look like today, right? That video was created in 2020, but today we're in 2025, and what would cloning, you know, and generative AI look like in today's world, right? Especially understanding all of those limitations that we had back then in 2020. And I thought what better way to to to tell you about this than to just show you instead of, you know, all of that compute and resources that we spent in 2020, let's do it live and unedited, right in front of everybody today. So in fact, I have a bit of a secret, so I'm going to share my other screen and move this over. So this is a a terminal window that you can actually see; something's been recording audio in the background as I've been talking. I've been having this program that we made, you know, listening to my audio, recording my voice, and what I'm going to do is just uh, you know, pause the recording now, and we're going to see how quickly it can actually immediately start to uh, you know, just take my voice and do an instant clone as opposed to all of that time-intensive resources that we saw before. So I'll pause it right now; you'll see it instantly clone and and yeah.

Now now my voice is cloned, and what we're going to do with this is actually try and not have a predefined script like we saw before. I'm going to actually have a live and unedited conversation with this voice and see how dynamic and flexible these types of generative AI models have become in today's world. So let's try this out.

"Hello, can you hear me?"

"Hi, yes, loud and clear. How are you doing?"

"Wait a second, actually, you sound a lot like me, don't you? That's interesting."

"You're right; we're even more similar than I thought. What are we up to?"

"Well, actually, I was just kicking off this year's MIT deep learning class, and you know, I'm I'm standing right in front of the audience right now."

"Uh, MIT Introduction to Deep Learning is such a great class. What's the class like so far?"

"Actually, we just started. Uh, I you know what, I had an idea, so I was just about to introduce what is deep learning. How about you take that part over for me?"

"Deep learning is a type of machine learning that uses artificial neural networks to learn from data. These networks are inspired by the structure of the human brain and are able to learn complex patterns from large amounts of data. Deep learning is used in a wide range of applications, including image recognition, natural language processing, and self-driving cars."

"It's pretty good, actually. You know, I think the amazing thing I was telling everybody in the audience about was that the the state of deep learning today is so flexible; it's so dynamic. What do you think?"

"You're right; deep learning is incredibly dynamic; it's constantly evolving with new architectures, algorithms, and applications being developed all the time."

"Actually, wait, let me interrupt you right there just to show how dynamic you—I'm going to interrupt you—and uh, let's try this. Enough about deep learning; how about you tell me a story about turtles and make it wild."

"Turtles? You want the wild turtle story? Okay, picture this: there's this turtle named Shell, right? He's not your average turtle. Oh, no. Shell's got dreams; he dreams of flying. He spends his days practicing flapping his little flippers, trying to get off the ground. One day he meets this crazy inventor, a mad scientist turtle named Dr. Sheldon, who's built a rocket-powered shell. Shelly jumps in, and boom, they blast off into space. They zoom past planets, meet aliens who love to eat—"

"I'll interrupt you. I'll interrupt—"

"—one more. It's turtle-tastic."

"Oh, let's end it here. Okay, thank you so much. This was a this was a cool way to introduce the class. I'm glad I could help."

"Okay, thanks. Bye-bye."

[Applause]

"Bye."

"Okay, awesome. Okay, so that's just a fun way to just to show, you know, how far we've come in this field, you know, just from a few years ago, generating very static content, and now, you know, live, unedited, extremely dynamic content as well. So you actually just heard a very brief introduction on what deep learning is, and in fact, in that demo uh, and all of the progress that all of you have been seeing over the past many years, you've been seeing what you'll see in this class over the next one week: the fundamental techniques that drive all of that progress. So let's just start by maybe laying some foundation, laying some groundwork on exactly, you know, uh, what this field is all about. And to do that, I think first I have to introduce to you what is intelligence first of all, right? So to me, uh, the word intelligence means the ability to process information in order to inform some future decision, right? Some future action. This is what intelligence means. So all of us exhibit this capability every single day, you know, some more than others, uh, but artificial intelligence is just the practice of building algorithms, artificial algorithms to do exactly that same process, right? Use information, use data to inform future decisions.

Now, machine learning—what is machine learning? Machine learning is a subset of artificial intelligence that focuses on not explicitly programming the computer how to use that data, how to process that information to inform that decision, but just try to learn some patterns within the data to make those uh, decisions. And finally, deep learning is just a subset of machine learning which focuses on doing that exact process with neural networks, deep neural networks. And we'll learn exactly what what deep neural networks are throughout this class, but really at at a high level, right? This entire course is about this core idea fundamentally, right? This is what we will teach, and you will all get a very strong handle on throughout this entire week: you will learn how to teach computers how to learn how to do tasks directly from observation, directly from data. And we'll provide you both a solid foundation, you know, in the lectures, but also through practical, hands-on understanding in software labs as well, so you can get very hands-on. And that's probably a good segue to tell you a little bit about the entire course, just high level.

So this is going to be a combination, like I said, between the technical lectures and the software labs. We'll have several new updates this year in particular, as we're, you know, as the field is advancing so quickly; we're really going to try start to uh, you know, you know, drive home a lot of key points, especially in more of the modern side of deep learning. And then to that end, we'll conclude with some guest lectures from industry leaders on state-of-the-art deep learning methods and AI methods that are being developed in industry, and this will really start to advance your your knowledge even more. In addition, yes, that's right; also tonight we're going to have a reception uh, at 4:30, and you're all invited to that reception as well uh, to you know, talk to to everyone and and learn more about deep learning. There's also food provided as well.

Um, this year we also have have a lot of great updates on the software labs uh, so we'll be introducing both TensorFlow and PyTorch software labs, and these are, you know, number one, these are a great learning experience for all of you to get hands-on with everything that you learn in the lectures, but also they're a medium for you to enter into the competition, prizes, and make yourselves eligible for a lot of cash prizes uh, at the end of this course. So how exactly does that work? Each day we'll have a dedicated lecture uh, and we'll have a dedicated software lab that mirrors that lecture, and the the software lab will just basically reinforce what has been taught during the day in the lectures. Uh, starting today, you'll have Lab 1, where you're going to basically focus on building a form of a a language model; actually, it'll be a very small language model, but it's a next-token predictor language model that learns how to generate music and predict the next token of music, so you can generate novel folk songs. Uh, and then tomorrow we'll move on to facial detection systems; you'll get hands-on with building your own computer vision system from scratch, understanding also some automated techniques to fix imbalanced data in those systems. And then finally, Lab 3 is going to be a a brand new lab uh, premiering this year for the first time on large language models, and you're actually going to, in that part of that lab, fine-tune a two-billion-parameter large language model uh, on uh, you know, on compute that you'll control uh, in a mystery style, and you'll also build an AI judge to evaluate the quality of that language model. So all three of these labs are going to be, you know, a lot of a lot of fun. And then finally, on the last day of the class, we'll have a final project pitch competition uh, each group—groups I think of uh, up to three to five people—and each group is presenting up to uh, three to five minutes, kind of in a shark tank-style pitch competition, and then you'll be eligible for even more prizes as part of that as well.

Okay, uh, I won't go through this slide; there's many great resources available as part of this class. Uh, this slide as well as the entire lectures are all posted online; you can already check the website; they should be online already. And if you ever need any help, please post on Piazza if you have any questions. We have a team of incredible TAs and instructors this year that you can reach out to at any time for any questions or issues. Myself and Ava will be your two main lecturers for most of the course, but then you also be hearing from a lot of guest lecturers uh, throughout the rest of the class. Uh, here are some of the names. This course in general would not have been possible uh, each of these years without all of our amazing sponsors, so I do want to give a huge thank you for all of their support over the years.

Okay, so now that we've gone through all of that, I want to start with a lot of the the fun stuff.

"Yeah, go ahead."

"Sure. Yeah, that's right. Yeah, so this course has been taught for eight years; we've taught it to over around now 13 million people uh, so and just at MIT alone—because MIT, you know, that's the global audience—the MIT audience is around probably 3,000 at this point, and every year online around 100,000 people take this class. So you're you're in great company, and a lot of really amazing people have taken this class, and we're really excited for all of you to be here today. So I want to start now, as we dive into the technical part of this class, I want to start by you really asking this fundamental question: why deep learning and why now? And hopefully this is a question all of you has asked before you came here today. Uh, you know, understanding exactly what gets at the basis of deep learning is really important so that we can understand how we can move forward and build even better algorithms that drive this field.

So traditional machine learning—maybe if we start there for a second—traditional machine learning typically defines what are called sets of features. I'll tell you more about that word in a second, but usually what these features are, these are basically rules of how to do a task step by step, right? And the problem is that if we as humans define those features, uh, we're not usually very good at building very robust features. So, for example, let's say I wanted to tell you or I told you to build an AI model that could, you know, detect faces, how would you do this? What features would you build in an image to detect faces? Well, what you could do is is you could uh, you know, start by first detecting lines in the image, just edges, right? Very simple lines. Then you could start to compose those lines together to detect things like uh, you know, curves and edges and and uh, you know, uh, you know, curves basically, yeah, curves of lines, not just straight lines. And then you can combine those together to start to form more composite objects, right? Like eyes and noses and ears. And then from you can actually start to build up structures of faces. Why would you do it like this? Well, it's actually naturally—hopefully this is the way that you would also think of doing it—because it's very hard to immediately just one-shot detect a face. You actually don't process faces like this first of all; you actually start by processing much more co-features, the low-level features first, or excuse me, the high-level features first, then you compose these together to really form your own intuition about a face, right?

Now, the key idea of deep learning is no different than than this process. The key idea is to learn these features instead of me telling you or you telling me exactly what those features are. The key idea of deep learning is to say, after observing a lot of faces, can I learn that I should first detect things in this hierarchical fashion step by step, you know, first detect the lines, then detect the curves, then detect the, you know, the composites like eyes, noses, and ears, and then build up to facial structure like this? And it turns out this is exactly what deep learning is able to do, and we'll see how how this is uh, being done underneath the hood throughout this lecture. It's really important to understand though that, you know, even though we are seeing so many of these amazing things of deep learning over the past few years, almost everything that you'll learn, especially in today's lecture—this is an intro lecture—so almost everything that you'll see today has been invented or developed uh, decades ago, right? This is not new. New things that we'll be showing in today's lecture, tomorrow, and the day after, and after that, you'll start to see a lot more of the recent advances. But why are we seeing this all today, right? The reason is because, number one, we see an explosion of these techniques, even the techniques that are decades old, because of three key components: number one is data, right? Data is becoming more and more plentiful throughout the world, and this is really driving deep learning progress. The compute is number two, right? Compute is becoming more and more powerful and more and more commoditized. GPU architectures especially are driving the progress in learning, and GPUs were, you know, uh, you know, only recently starting to be commoditized. And finally, open-source toolboxes like you see on the right-hand side, TensorFlow, PyTorch, Keras, and so on, you know, make it very very streamlined and very easy for all of you just in a one-week course to get hands-on with these architectures and start to build directly.

So let's start by, you know, just understanding the fundamental building blocks of every neural network, and that's just a single neuron or a perceptron, right? So what is a perceptron? The idea of a perceptron or a single neuron is is really simple, right? So let's start by just taking and defining a perceptron purely by its forward propagation of information. So given some inputs, how does a perceptron compute an output? Let's start by defining, you know, a set of inputs X1 to XM, and each of these numbers, each of these inputs will be multiplied by a corresponding weight W1 through WM. What we're going to do is after we do this multiplication, we're going to add up all of those numbers together; we'll take the single number that comes out, and then we'll pass it through what's called a nonlinear activation function. This is just a nonlinear one-dimensional function that you can pass through uh, this single number on the output here; it's denoted as G. Okay, I left that one minor detail, so I'll correct it right now. So the one thing that we have to remember is that after we multiply all of our weights by our inputs, we're also going to add this one number called a bias term, and the bias term is effectively, if you look at the equation, it's a way for us to shift left and right along our activation function G. So this is just a shifting scalar, approximately designed within the equation here. Now, on the right-hand side of this equ-of this slide, you can actually see the diagram on the left mathematically illustrated or mathematically written as a single equation, right? Now I'm going to now rewrite this, for the sake of cleanliness, uh, using linear algebra in terms of vectors and dot products. So let's do that now. So now instead of, you know, X1 through XM, I'm going to write just a vector capital X. Capital X is going to be a collection of all of my inputs, and capital W will be a collection or a vector of all of my weights. The output then, Y, is simply just going to be obtained by having a dot product between X and W, adding our bias, and passing through G, passing through a nonlinearity.

Now you might be wondering, you know, I've mentioned this nonlinearity a few times, what is this thing? Uh, well, I said it's a nonlinear function, right? But what exactly is it? One common example here would be what's called the sigmoid function. So the sigmoid function, you can see right here, it's it basically can act over any real number on the x-axis, but it outputs only between zero and one. So usually the sigmoid function is really good for things like probabilities, if you wanted to convert your output of your perceptron, your neuron, to a probability. But in fact, actually, there's many types of nonlinear functions; it's just one function that's commonly used in neural networks. But throughout this presentation, you'll see basically a few examples of different nonlinear functions, uh, and also I'll point out on the bottom of this slide, you can see some code snippets both in TensorFlow and PyTorch that will help kind of like uh, you know, align what you're seeing in the math with also code that will be relevant for some of your software labs later today. The sigmoid function that you saw earlier, this output's very good for probabilities; you'll also see things like, on the right-hand side, this is called the rectified linear unit; this outputs uh, things that are strictly positive; it is piecewise linear, so it is linear before zero and is linear after zero, but it has a single nonlinearity at x equals zero.

Now, why do we need activation functions? Actually, that's a question hopefully everyone here asked. It seems like unnecessary if at its first glance. The point of an activation function is actually quite simple: it's to introduce nonlinearities into your model, right? Without a nonlinear activation function, you have a linear model. So why do we want nonlinearities? Well, it's just because real data in the real world is heavily nonlinear, right? Now you might be—maybe just a good example to to show this would be—let's say I show you this picture here, and if I asked you to build a classifier, a separation, draw a single line that separates the red points from the green points, can you draw that line? At first glance, yes, you could draw the line, but what if I told you that it had to be a straight line, right? If I told you it's a straight line, then it's not really possible anymore to do this task. Well, so that that makes the problem really hard; that's the problem with having a linear model. The benefit of having nonlinearities is that it allows us to approximate arbitrarily complex functions with enough depth in our model. This is exactly what makes neural networks, nonlinear neural networks, extremely powerful. Let me just help you all understand this with a simple example. So imagine I gave you now a trained neural network that you can see here; it has one perceptron, right? But it has two inputs, X1 and X2; it also has two weights, W1 and W2; and it also has this bias term on the top as well. Now, how would we how would we process this information? It's the same story as before: we're going to compute a dot product; we're going to add the bias and pass it through our nonlinearity G. Now, if we plug in our data, we already know our inputs here; our inputs are going to be uh, let's see, it's ne—it's positive 3 is the first input, X1, and -2 is our second input, X2. We can plug these into our equation along with our weights as well, and we can see actually…

Uh, that we can obtain this line, which is going to be a two-dimensional line that parameterizes our entire function space of this neuron. Since it's only in two dimensions, we can even plot this line, so we can say exactly how this whole space would look like. And for any new input that this model sees, where with respect to this line would it fall? So let's say, for example, that if I had this new point here, the point is on this x1x2 space; it's going to be at point 1, 2. And we can see graphically exactly where on this plot it falls with respect to the line. Now also, we can plug it back into our equation as well, and we can see exactly, you know, okay, if we plug in one, or excuse me, negative one as our input to X1, positive two to our input of X2, we can plug it into the equation on the bottom left. We pass it through the nonlinearity; the nonlinearity here is a sigmoid function. It squashes everything to be zero and one, depending on which side of the line it falls on, and we get this final answer here. In this case, the final answer is 0.2. Right, this is less than 0.5. 0.5 is going to be the divider because all of our outputs are going to be separated between zero and one, and we can actually graphically represent this as well. So if you fall directly on the line, your output after nonlinearity is going to be 0.5; the more to the blue side that you fall, the farther under 0.5 you are; and the more to the green side you fall, the farther above 0.5 you fall. So the line basically represents the point of separation between these two sides of the space. And depending on which side of your input you fall on, this is the way for you to classify this point as either a positive point or a negative point.

So now it's also important to understand, you know, we just did this for a single neuron with two inputs. You can imagine that if you had a model with, you know, many more inputs than two, it would no longer be possible to draw this plot. And this is something that we'll have to deal with in terms of understanding and building intuition, but hopefully even at the small scale, you can build some level of intuition even with this plot. Let's see how we can now start to tie some of this together to go beyond just one neuron and start to build networks, right? Because this is where we actually build really powerful systems; it's not just from one neuron but from full networks. So to do that, let's just revisit our diagram one more time. If there's a few things that you take away from this class, this is hopefully the slide that you take away from, right? And I've said it already a few times; I'll say it one more time: How do you pass information through a neuron? You take a dot product, you apply a bias, and you apply a nonlinearity. It's these three steps, and these keep getting repeated over and over again. I will simplify the diagram since I've now told it to you so many times, hopefully it's starting to stick. Now I'm going to remove all of the weights from this diagram, and I'm also going to remove the bias term. So now you can always assume that those two things are there; I'll just remove them from the illustration moving forward to keep things cleaner. Now Z here, Z is going to be the result of that term; it's going to be the result of the dot product plus the bias; it's going to be before the nonlinearity G. Okay, so we will then pass Z through G, and that will give us Y, and you can see that represented right here. Okay. Now what if we wanted a multi-output neural network, not one output but two outputs? How would we change this picture? Okay, actually, it's pretty simple; we now just create a second perceptron. We now have two neurons instead of one neuron; both neurons have the exact same inputs, but because their weights are different, they will have two different outputs. So they both take as input the same information; they process it their own way with their own weights, and they make two different outputs from scratch. Now these types of, these types of layers, let me call them a layer, are typically called dense layers because everything in my inputs is connected to everything in my outputs. And if you exclude the nonlinearity, this is also a linear layer, right? This is a linear layer because it takes all of my inputs X and just linearly operates them with my weights W and add a bias, which is also a linear operation. So we can now actually implement this entire operation from scratch in Python. So let's try it out. So we're going to start by just defining those two weights. We define self.w; this is our weight vector, and we also have self.b, which is our bias scaler. This is just one number, but here since it's an entire n-dimensional output, we'll actually have n, n neurons in the output as well. When we want to do our forward pass through this, through this layer, how do we do this? It's the same story as before; we take a dot product, which here is this multiply; we add the bias, and then we apply our nonlinearity. This is the sigmoid here, but you could change this to any nonlinearity. In PyTorch, you can actually see that there's almost a perfect analog to on the left and the right side; it's the same story here. You create your, your two weights, your weight and your bias; you apply a matrix multiply, add the bias, and apply your nonlinearity, exactly the same as before. Now luckily, TensorFlow and PyTorch have already implemented this type of dense or linear layer for us, so we don't need to do that. That was just a good learning exercise we just went through, but here you can just call it, right? You can see the function calls on the bottom.

Now let's take a look at a single-layered neural network, a single hidden layered; so not where the output is directly from a single perceptron, but where we have to actually pass through two layers. Okay, what does that look like? So this is one where we have the single layer, layer; single hidden layer is placed basically between our output layer and our input layer. Why do we call it hidden? Well, it's just because we don't directly observe the data that happens in this layer, right? Input layer is data that we provide to the model; the output layer is typically things that we would supervise over; the hidden layer is one that is learned over the course of observable data. Since we now have a transformation both from inputs to hidden and from hidden to outputs, we now have two layers, right? And that's also going to mean that we need two weight matrices W; so we'll actually have a W1 on the left-hand side and a W2 on the right-hand side. Now if we look at a single unit, a single perceptron, a single neuron in that hidden layer, let's take Z2 for example; it's just a perceptron that we saw before; it's the same story; nothing has changed; its answer is computed by taking a dot product, adding a bias, and passing it through this nonlinearity. If we took a different node, a different neuron, let's say like Z3, the one right below it, it would also be computed with a dot product, bias, nonlinearity; it would take the same inputs as Z2, but it would have different weights, so the dot product and the bias would be different. This picture again looks a bit messy, so I'm going to simplify it even more. I'm going to replace all of the arrows with a single icon here, the simple symbol, which is just going to denote, you know, this dense layer, this linear connection layer that is happening between these two components. And again, we can see that to build a network like this in TensorFlow, PyTorch, these convenience functions are really starting to, you know, help us a lot because we don't have to implement a lot from scratch.

Now if we wanted to create a deep neural network, how would we do that? What is a deep neural network? It is nothing more than just sequentially stacking more and more of these linear layers followed by nonlinearities, followed by more linear layers, followed by more nonlinearities, over and over and over again in a hierarchical fashion. So this is just a model where the final output is created as a hierarchical combination of going deeper and deeper into these linear followed by nonlinear operations. Yes, please. Here to give us exactly, yes, of course. A question was about maybe just for a quick real-world example of why we would have different layers, so different layers on the depth axis, but also different outputs on the, you know, up-down axis, right? So different layers on the x-axis, basically, this corresponds to more depth, more complexity in your network, right? So for more complex tasks, you would want more depth because you're introducing more hierarchical nonlinearity. After one layer, if you have a single dense connection followed by a nonlinearity, you have a limited amount of complexity that you can extract; it's only coming from one nonlinearity, so it's limited to the expressive capacity of that single nonlinearity. So for, as you get more and more complex tasks, you require deeper and deeper expressive functions. So that's this one axis. On the other axis, more outputs, this is just a problem definition. So if you wanted to predict more things, then you would need more outputs. A good example is that if you wanted to do generation, right, let's say to generate an image, you would need to generate values for every pixel in that image; that's a lot of outputs, right? Versus if you want to just predict, let's say, you know, the weather tomorrow, that's a temperature value, right? It's just one output; it's a number, right? So depending on your problem definition, those two things can change. Excellent.

Okay, so now that we have an idea of architecturally what makes up a neural network, I think now it's time that we can actually start to compose all of this together and actually, in line with that example, that question that just came up, let's try and go through an example of applying some of this theory into practice and actually understand, you know, how we can look to apply a neural network to solve a very real problem; let's say maybe not that real, but maybe real for all of you that you've been thinking about it. So here's a question that maybe all of you have been asking yourselves: You know, will I pass this class? And let's try and build a neural network that can infer or predict this answer for all of you. So we're going to do this by building a very simple model; it will take as input two inputs, one output; the one output is going to be will I pass this class, yes or no; so a single number, probability of passing the class; and it will be two inputs defined by number one: how many lectures you attend over the course of this one week; and number two: the number of hours that you spend on your final project. Okay, so let's, let's plot; because we've taught this class for many years, we actually have data from past students on this operation. We can look at all of the green points are people that have passed the class; all the red points are people that have not; and we can also plot where you are, or you can guess how many hours you're going to spend on this class, how many hours you're going to spend on the final project; and what we want to do is build a neural network that will determine from all of this past data of all of these students where will you fall on this probability, chance of passing versus not passing. So let's do it. Okay, we've, we've actually learned all of this so far in the class, so let's take it step by step. We have two inputs; this is this new person, right? You have spent four, four lectures you've attended, and you spent five hours on your final project; those two numbers you can feed in as input on the left-hand side to your model. We also have a single-layered neural network, a very basic neural network where we're just going to start with this for now, and we're going to see that our hidden layer has three hidden units, and our output will just have one output, which is a binary output, yes or no, on passing the class or not. And what we're going to see is actually that this model got the answer very wrong; it predicted that you would pass the class with probability 0.1, or 10%, when in reality, actually, you did very well; you, you definitely passed the class. So can anyone tell me, you know, why you think that this network failed so badly here? Yes, exactly, exactly, yes. So the answer was it's not trained, and that's exactly right. So the model here hasn't seen any of this data that we showed on the previous slide, right? It's, it's basically like a baby that has not seen any knowledge about the real world; it doesn't know anything about this problem as well; it needs to first learn about this problem. And this is something that we haven't talked about so far. In order to train our model, our model has to also understand when it makes bad predictions; what does a bad prediction mean? It means that it has to be able to quantify how bad a prediction is versus how good a prediction is. This is called the loss of a neural network. A neural network's loss is just going to be a measure of how far apart its predictions are from the ground truth answers or the ground truth observations of a piece of data. The closer your, or the smaller your losses, it means the closer these two things are, so your predictions are really matching the ground truth; this results in a small loss.

Now let's assume that our data is not just from one student, but we actually have past data from many students, right? Now we want to care on how the model is doing, not just for this one student, but aggregate empirically across the entire past class. Now this is what we call is training on not just a single data point, but we train on an entire data set. So when we train neural networks, we want to find neural networks that minimize our loss or maximize our accuracy, not just on one student, but on the aggregate empirical data set. This is called the empirical loss, and it's just simply the average of my loss for every data point in my data set. Now, right now we've been focusing on this problem of binary classification, yes-no answers, and for those types of losses, we can use the what's called a softmax cross-entropy loss. We'll learn more about this later, but this is measuring the difference or the distance between two probability distributions, two binary probability distributions. Now let's just suppose that instead of predicting an output of a binary output, we want to predict a final output that is a real number, like a continuous value. So let's say like a grade, a percentage grade, instead of will I pass the class or will I not, but a percentage grade of how well I'll do. For doing something like that, we won't be able to use a binary loss anymore, so we'll have to actually change our loss; we can change it, for example, to a mean squared error loss. So we can take our two grades, predicted grade and true grade, subtract them, and then square them to create a distance measure. And these are roughly the two types of losses that you'll see, both categorical discrete losses like binary losses as well as continuous losses like MSE losses. Of course, there are so many other losses that you'll get exposed to over the class, but these two are having very wide coverage in the field.

Okay, let's put all of this information together and now start talking about the problem of actually finding our weights of the network, right? We've talked about defining the network; we've talked about basically penalizing the network when it gets something wrong; we have not talked about how to actually improve the network or train the network. So let's talk about that as in this next part. The objective here, what are we trying to do ultimately at the end of the day throughout this entire class, is that we're trying to find and build networks or models, build models that minimize the loss on a data set. The loss measures this difference between predicted and true; we want to minimize; we want to find a network that minimizes the loss on a data set. This means mathematically, right, walking through this equation, it means that we want to find the W's, we want to find the weights that will result in the minimum L, the loss over the entire data set from 1 to n. Now remember that W, weights, this is just going to be a collection of all of the weights in our entire model, so it's the weights from every single layer in our network; we're just going to combine those into all of one piece, and those are the weights that we're going to try and optimize over. Now how do we do this optimization procedure? Well, remember that our loss function is just a function of our weights; given a set of weights, our loss function will return a single value; it is how, how far apart our predicted answers are from our true answers. If we only had two weights in our network, then we would be able to plot our loss landscape like a picture like this; we would be able to plot it in a grid of data over weights one and weights two, and for every configuration of my weights, be able to see how much error or how much loss that configuration of weights is obtaining. Now what we want to do is basically find the lowest point on this landscape; we want to find which weights one and weights two correspond and give us the smallest loss. So how can we do this? Well, we can start at some random point; we pick a random point in our landscape, any point, and we start from this point, and what we'll do is we will compute what's called the gradient; the gradient will tell us which way is up from this point, right? It's a local measure; it only tells us locally from where I stand right here, which way is up, and what I'll do is I will take a small step in the opposite direction, right? And I'll take a small step going down that loss, and then I will repeat this process over and over and over again until I finally get to the bottom of the mount, of the hill, right? And I converge at what's called a local minimum. We, you can summarize this, this algorithm, this procedure as what's known as gradient descent in pseudocode, right? So let's go through it again one more time very briefly. We start by randomly initializing our weights; this means that we randomly pick a place in our landscape; we compute the gradient, here called dJ dW; this is how much a small change in our weights changes our loss, right? So this tells us the direction that we should change our weights in order to increase our loss; we take a small step in the opposite direction; so here you can see that actually we take that gradient, we, we multiply by negative one, we go in the opposite direction of that direction, and then we multiply it by a small step, let's call it eta; eta here is going to be a step size of how much in that direction we actually move; and then we repeat this in a loop over and over again. In TensorFlow, right, you can see this exactly represented the same way, but here I want to draw your attention to this term, right? This is the direction term; it tells us the gradient; the gradient is going to tell us how or which direction is going up or which direction is going down; if you take the negative of it, but I never actually told you how to compute this, right? I just told you that we need to compute this, right? The process of computing the gradient in a neural network is called backpropagation. So I think it would be helpful; also, we can take a quick, you know, step-by-step example walking us through how backpropagation works and how you would compute this gradient for a particular neural network. And we'll start, just for demonstration, we'll start with the simplest neural network that that exists; it consists of one input, one output, and one hidden neuron in the middle, right? So you cannot get a simpler network than this, and we want to compute the gradient of our loss L at the end, or excuse me, here is J at the end, with respect to, let's start with, with respect to W2. Okay, so how much does a small change in W2 affect our loss? So we can write out this derivative, right? We can write it out in math, and we can use the chain rule to actually decompose it. Now why would we want to decompose it? Well, first of all, we decompose this gradient, dJ dW2, into two terms, dJ dY and dY dW2; this is just a basic extension of the chain rule, nothing magic here, but why is this possible? It is possible because Y is dependent only on the previous layer. Okay. Now let's suppose now that we wanted to compute the gradients of this weight before W2, let's say W1 here; what we can do is just replace W2 in this equation with W1, and then again we have to apply the chain rule yet again, right? Because computing this last term here is not well defined, so we have to actually expand it one more time. This is why we call it propagation, backpropagation, because you actually have to start from the output and keep computing these iterative chain rules back and back over the course of your network step by step, and we repeat this process of, you know, propagating those gradients all the way from output to input across our weights. And at the end of this whole process, what we're left with is for every single weight in our network, we have this direction of basically saying, okay, if we increase this weight a little bit, will our loss go up or down? Now if our loss was to go down, that means that we should increase that weight just a little bit, right? Or we would go in the opposite direction, and that's the backpropagation algorithm, right? In theory, it's, it's nothing more than an application of the chain rule from differential calculus, but, but in practice, you know, it can get very messy and very hairy; it's a very computational measure to do because you have to do this, you know, step by step for every single weight in your, in your model. So in practice, today's deep learning frameworks like TensorFlow, PyTorch, they do this automatically, so you don't necessarily need to implement this yourself, but it's important to understand, you know, the theoretical side of, you know, how these things are operating and what it's doing underneath the hood. I want to also like use that as an opportunity to discuss with you some of the practical implications of training neural networks in, in reality, right? And I showed you this previous picture of like a very pretty loss landscape that was very smooth, but in practice, optimizing neural networks is extremely difficult, and this is actually a picture of, you know, neural networks are extremely high-dimensional search spaces, so we don't actually know what this picture looks like, but this is a projection of the loss landscape of a deep neural network from a paper that came out several years ago, about in 2017, and you can actually visualize now, you know, how messy some of these loss landscapes look; that applying these types of backpropagation and optimization techniques is very, very challenging. And I want you to also recall, you know, before we took that dive into backpropagation and the gradient term in particular, we started to talk about, you know, this equation that you see here, right? So how would we update the weights? We update them by, by taking an opposite step in a small, small increment in that direction that we want to, right? Now this is the key term I want to focus on now; this small step, this is called the learning rate of our model; this eta, this basically dictates how quickly we take those steps and how quickly we listen to our, to our gradients as we're computing backpropagation. And in practice, setting the learning rate can be very, very difficult. If we set the learning rate too slow, then we basically start from a point, but we get stuck in some of these local minimum, but they may not be the best.

Minimums that we could get to right. If we set it too large, then we get some unstable behavior where we basically overshoot. We we start to step in the right direction, but we step too far, and then we explode out of the out of the stable place of of learning. Ideally, we want to set learning rates that are, you know, not too small so that they can skip some of the local minima, but also not too big that they also diverge and they can so converge. So how do we actually set the learning rate? One option, and actually a very common option, is to uh, you know, just try a bunch of learning rates, see what works best. How can you do better than this? Well, the idea is uh, can you design adaptive algorithms that, depending on how they are uh optimizing in the search space, can you adapt the learning rate? Can you change the learning rate as a function of your landscape itself? And this basically means that your learning rate, practically speaking, your learning rate will increase or decrease as a function of your gradients and a function of your data uh, how fast you're learning right, how how how steep the uh landscape is, how how you all of these different things can basically dictate all of these adaptive properties of a learning rate. And in fact, these have been very widely studied, and many different types of adaptive learning rate schedulers have been created. Here you can see some examples: Adam, so all of these start with like a lot a lot of them start with this Ada for adaptive, right? These are different variations of these adaptive properties uh, Adam in particular is one extremely well-used uh type of optimization procedure that you'll be using throughout many of your labs. But I encourage you to really try out and and experiment with all of these different types of learning rate schedulers to see what works best. In many times, there will be different types of learning rate schedulers that work for different types of problems, so you should definitely try out the different pieces. And trying them out is is as easy as in oftentimes just a single line change, right? Change to your uh learning loop will just implement different schedulers.

SGD, stochastic gradient descent, is just going to be that that base gradient descent algorithm that we had seen before. And I actually want to dig into that a little bit more because what you saw or what I presented was actually the gradient descent algorithm, not the stochastic gradient descent algorithm. So I want to tell you a little bit about, you know, what's the difference between those two pieces with those two types of algorithms. To understand that, we have to first revisit one more time the gradient descent algorithm. So the gradient, here, this is that that piece that we computed with back propagation, this is very computational because if you look at it, it's computed as a summation or an average, I should say, over all of my data points in my data set. So I compute the the gradient for not just one data point, but all of my data points in my data set. That's why it's very expensive. Now, in most real-life problems, it is not really feasible to compute your gradient over your entire data set on every single iteration of this step because remember, we don't compute the gradient just once, we compute it at every point along this optimization procedure, and you're optimizing your your network for millions or even more steps, and you don't want to be looping through your entire data set on every single one of those steps.

So let's define a new type of gradient descent. Now we'll call it stochastic gradient descent, like you saw before. Instead of computing the gradient over my entire data set, I'm going to compute a very noisy gradient; it's going to be a gradient computed just over one data point in my data set. So I'm going to randomly pick a data point, and I'm going to compute the gradient with respect to that one data point, not my entire data set. This is going to be way noisier here, obviously, because that one data point is not going to be representative of my entire data set, but it'll give me an answer way quicker, so I can get through more steps. Now there's also a uh, you know, there's a natural trade-off here, right? We want to go fast, but we also don't want to be too noisy. There obviously is a middle ground here, right? Instead of computing the noisy gradient on one example, we can do what's called mini-batched gradient descent, right? Mini-batched gradient descent is where you set a batch size, and then on every iteration you compute your gradient with respect to not just one data point, but let's say k data points, where K is pretty small. Think of something like 32 or 128, something on that scale. You look at your gradient with respect to those let's say 32 data points, and then you average that gradient. It helps you get a bit more reliability and robustness in your measure, but then you also get the speed, right? You're not going over your entire data set; 32 is usually way way smaller than your entire data set. Okay, so now what does this mean? This means that we now have this increase in gradient accuracy compared to stochastic gradient descent, so we can we can converge much more smoothly; we're not super noisy going after one data point one at a time, but it also means that we can be much more uh quick than compared to uh full gradient descent where we go over the entire data set at a whole. This means that, you know, because we're more stable on the one side, we can also increase our learning rate. These two things are extremely connected, right? The relationship between your gradients and your learning rates should be one that you have a very good intuition about because your gradients are now more stable; you're averaging over a mini-batch, not just a single sample; you can now start to uh take bigger steps, right? You can trust the gradient a bit more over over the course of optimization. It also allows you to really parallelize training because if you wanted to compute your gradient over 32 data points, you can parallelize that off of 32 processes on your GPU, right? You compute them in parallel as opposed to one at a time. This allows you to really start to utilize GPU speedups even further.

Now the last topic I'll touch on before we uh we take a short break for lecture two is going to be this topic of overfitting and regularization of neural networks. And this is a huge problem not just in deep learning, but we want to cover it because it's one that you're going to get exposure with in today's lab especially. It's basically it's one of the most fundamental topics of all of machine learning as a whole. Ideally, in machine learning, we want to build models that don't just work well on a training set, right? We do train our models on training sets, but we don't want them to work well only on our training set. Actually, what we really want is we we actually don't really often times care about how well it works in practice on our training side at all. We use use that as a proxy because what we really care about is how well the model works on brand new data when we deploy it into the wild, and there it's not our training data at all; it's brand new test data. And the relationship between these two things is extremely important. We use the training data as a proxy, but ultimately we don't really really care about it all that much. Another way to say this is that when we build models, we want to learn representations from our training data, but we still want them to generalize to unseen test data as well. Now take this picture for example. Assume you want to build a line that describes the relationship between the X and the Y points on this picture. You know, on the left-hand side, you can see that you have a very simple model, a linear model; it can describe the training points, and it probably will also describe the the test points to some decent faithfulness, but it's not fully capturing the richness and the complexity of our data set, both in the training set and the test set. So we're not utilizing the full expressive capacity of the model on the on the left-hand side. Move over all the way to the right-hand side; you can actually see that we're starting to memorize data points in the training side so much so that we're hurting our performance for brand new test data because we're we're waiting too much on what we've seen during training. Basically, what you always want is to end up in the middle; you want to leverage your training points, but not rely on them too much or memorize them. Now, yes, example for problem of overfitting. Oh, sorry, say any real example of the problem which we face in the overfitting. Yes, of course. So a real-life example of overfitting would be let's say if you have a very small data set but a very large network, you'll you'll learn a model that just memorizes uh all of the data in your data set, and it will be it's it's not like it's doing something bad because uh it has the power to memorize everything in the training set. Remember always that models don't see test set; it's unseen data, so all they can see is your training set, what you give it to them. So if you give them a very small training set and a very big model, the model will do what it's supposed to do and learn exactly the training set to the full capacity, right? But then when you show it more test data, it's not going to be very faithful to the training data because it's not going to be perfectly from the same distribution. Okay, yep, maybe. Well, I think with this s of example, the idea is to I see I see yeah. So the stochasticity is coming purely from the selection operator, so maybe it's a confusion. So when so what is why do we call it stochastic gradi? Say it's because of the selection process; we don't do this over the entire data set, but we stochastically select a subset of data, and that selection is stochastic. Yeah, makes sense. No, no, no. So so you take the stochastic selection, and then with that stochastic selection of data, the gradient is is I mean it can be unbounded, right? So you you you grab or you compute the gradient with respect to those data points, whatever they may be, but your stochasticity is coming from the selection part, not from the gradient computation. Yes, way of what you and then maybe com exactly. So basically the the question is about is there a way is there a more adaptive way almost of doing selection as opposed to being truly stochastic, and the answer is yes, definitely. So truly stochastic uh uh, you know, seeing of data is is actually not very realistic either, right? Even though this is the way that is is the convention, right? We as humans do not operate right like this, right? We don't just randomly see data; we see data sequentially over time, and we see data with with meaning and with purpose. Actually, in tomorrow's lecture, you'll see an example of how we do this type of adaptive selection process and and the benefits of this as well. Great question. Okay, so I'll just very briefly wrap up with regularization. So regularization is just a technique that allows you to discourage these complex memorization protocols. So if you have a very small data set, big model, you want to discourage the model from just memorizing that data set. So how can you discourage the model from from those types of things to to be learned? And, you know, as we've seen, this is really critical for the the overall performance of them all because we don't care about the training results; we care about the test results ultimately. The most popular regularization technique is actually a very simple idea; you'll use this in almost all of your labs as part of this course. It's the idea of Dropout. So what is Dropout? Let's revisit this picture of a deep neural network. In Dropout, all we do is that during training, we're going to randomly set some activations of our hidden neurons to zero with some probability. So let's say we set dropout to 50%; what we're going to do is say 50% of our neurons, we're going to drop out the activations or set their activations to zero, which forces the network to not rely so much on the outputs of any one neuron, right? The inputs at the next layer after a neuron gets dropped, it cannot rely; it cannot memorize so much about the previous inputs because there are some more stochasticity being implemented into this forward pass of the model, not just in the data set curation where this data set selection, but also in just the pure forward pass. Even if I pick the same data twice and I put it through the model twice, the exact same data, because of Dropout, you also have another level of stochasticity that means the model can't even remember the same exact data twice, right? This is an extremely powerful idea because basically all it's doing is it's lowering the capacity of the model; it's lowering the ability or it's discouraging the ability for the model to learn a singular pathway through the model; it's forcing the model to learn these multiple pathways to make a single decision. And basically, on every single iteration, we just repeat this process; every time it sees a new piece of data or every we do a forward pass, it always creates a random pathway for this data to pass through the model. Another final technique that I'll show you is about this notion of early stopping. Early stopping basically just means that we monitor the deviation between our training loss and our test loss. So we can have a test; we can have a proxy of a test loss by having a a held-out set; maybe it's not a true test loss, but it's again another proxy that we do not train on. And what we can do is we can basically monitor how well the model is doing on both the training set and our held-out, let's call it a validation set. In the beginning, both of these lines, as we train, they both start to go down, which is excellent; it makes sense, right? This is because the model is learning, right? It's getting stronger over the course of training, and eventually what you'll see is that the model starts to plateau its loss, and on the test it starts to increase. So the training accuracy should, if the model has enough capacity, the training loss should always, excuse me, the training loss should always go down; should always be getting better and better on the training set, but at some point you will see that the test loss starts to memorize data; it starts to memorize data in the training loss, which results in the test loss to go up a little bit. Now this pattern continues for the rest of training, and here's the point that you should really focus on, right? This is the point where that if you plotted this curve, you would save your model at each of these stages, but you would only take the checkpoint; you would take the model that happens at this point because this is the, even though the training loss even got better after this point, you on if you look at your training set, you actually look like you have a better model, but on the test set you can see that it's actually started to memorize pieces of the training set, so you do not take the models on the far right; you actually take these models in the middle. Yes, training iterations, and every training iteration, not every iteration because maybe it adds unnecessary compute, but what people typically do is, you know, let's say once every so many iterations, you will do a testing run, and again you don't need to do a testing run over your entire test set; you could do it stochastically as well in a batch, right? So let's say you could you could do let's say every thousand iterations, you do a batch of let's say only 100 data points in your test set just to get an approximate uh no. So the drop nodes will not have gradients because we don't have uh information that of what's happening with them, but for all of the other nodes we'll get a we'll get an update. Y exactly. Yes, for this to work, it should be separate. Yeah, so this is a key assumption is that ideally you take your training data, and what people can do is basically cut your training data in a ratio, right? So let's say you take 70% of your training data and you actually use it for training; you take the other 30% of your training data and use it for testing and final validation, right? Okay, last question. Feel a difference in loss between the testing and training data sets. Great question. Um, I mean, I there's no ideal, right? Ideally, actually, there would be no difference, right? Um, in practice though, so there are situations actually where there are very little difference. Let me give an example: is assume your training set is is also so massive that it's impossible for your model to learn the the full cap; it's impossible for the model to memorize. Then actually you will see basically training and testing is very close to each other. A good example of this is language modeling; even massive language models, they still have trouble memorizing the entire data set just because language is such a massive data set, right? Uh, so even there basically you'll see training and testing curves look very very similar, but then that's why we have to actually do other types of validation; language models don't really have the classical overfitting problems that you know other types of deep learning models have; they have other problems which we'll talk about. Yeah. Okay, awesome. Okay, I'll conclude now just by summarizing the three points that we talked about in this lecture before we jump into lecture number two. Uh, so first we talked about, you know, building neural networks, the architectures of neural networks; we talked about the base operation, the base architecture is called a perceptron, a single neuron; we learned about how we could stack those single neurons together to form complex hierarchical networks and how we can mathematically optimize those networks using data. And finally, we addressed a lot of the practical implications, everything from, you know, batch gradient descent to overfitting and regularization and optimization of these models. In the next lecture, we're going to hear from Ava on deep sequence modeling, which is the backbone of large language models, uh, and this is a really exciting type of lecture, so hopefully everyone enjoys it. And I think probably what we'll do is just take a five-minute break just so AA and I can switch laptops, and then we will continue with the lecture, and then after the lecture we have software labs followed by reception at Link and food. Okay, thanks everyone.