Transcription
All right, well, um, sorry we're a little bit late, but let's get started. Um, so welcome to the first lecture on the lecture series called "Future of AI: Foundation Models and Generative AI." This is actually the second year we, we hold this class. And so I started working on this course before the recent hype and breakthroughs of ChatGPT, um, and I really felt that we're starting to see kind of a new approach to AI in the community that was really going to change things for real. And I think we started to see that right now. Um, and really, what I want to accomplish in this lecture series is to give you an understanding of why this is happening right now, what's the underlying kind of change in perspective, and also going beyond just kind of the tip of the iceberg, which is ChatGPT. So I'm going to give you a deep but non-technical introduction to these subjects. And, and, uh, last year when I gave this course, we were excited about, uh, text-to-video and text-image models, right? Guess we still are. Uh, we were excited about, uh, super-human robotics, self-driving cars, and AI applied now to other domains as well, like genomics, etc. And, um, of course, a lot of these breakthroughs, even if they're in very different domains, they come back to this underlying technological achievement of foundation models, generative AI. Um, we're going to dive into, last year, we were excited about ChatGPT, right? So we could ask it to write an engaging introduction to an AI lecture. If we ask the updated GPT-4, it does it, produces more text, but maybe does a better job as well. Uh, last year we asked it to produce an engaging artwork of artificial intelligence, uh, and now we're asking the newest version. Uh, it also might be doing a better, better job. It looks more involved, at least. Art is subjective, but, uh, I think it's better. So if we were excited about these things, uh, in, you know, early 2023, what's happened, right? What, uh, what's happened during this year? Well, uh, a lot of things. Of course, there's been a tremendous hype. So there's been a lot of, you know, money pouring into these, uh, areas. We've had companies that only, you know, few weeks or months old, reaching a $2 billion valuation, which is a team of five people. We've had excitement about autonomous agents. We're going to talk about during this course as well, like GPT Engineer, that's able to plan and, and even act in a more human-like way in terms of intelligence. Nvidia, that provides all the different GPUs, right, that these models need, have reached a, a huge valuation, like the $1 trillion club. We've seen, uh, sweeping, uh, regulatory, uh, kind of acts and, and, uh, uh, initiatives, right, both from the White House and the European Union, for example. There's been a lot of drama in the AI space, right? OpenAI, for example, the CEO and the company behind ChatGPT, the CEO was outed and then came back in. So maybe the transparency problem in AI doesn't only apply to the models, but to the structures and companies behind them. And also, you know, we're seeing kind of some hype and some winners, uh, in terms of the this new AI technologies, but also now some companies are actually losing, uh, users and usage, right? Stack Overflow, for example, people saying that kind of is being killed by an AI that's training its own data, which is kind of ironic. Of course, one of the big questions that remain is, have you reached artificial general intelligence yet? Uh, some people say we have. I think, uh, there's still quite a long way to go, but we're going to try to also explore a little bit, you know, what, what can we actually mean with AGI and how could we potentially reach it, given the technology that we have right now, and can we give some kind of very, order of magnitude estimation to when we'll get there? So, uh, I'm Richard. I was, uh, born in the land of Albania, which is Sweden. And, uh, before MIT, I was at Stanford, uh, for seven years, and I did research as well on AI and, and this stuff. Also started a company that, that does, uh, foundation models in generative AI in commercial settings. I hope to bring in a little bit of that, those perspectives as well. Uh, for the last four years, almost now, I think, yeah, almost four years, I've been at MIT, uh, where I do research on, on self-supervised learning and, and financial models, and that good stuff. Okay, so, uh, quickly on this course schedule. So today, we'll give an introduction and kind of a history of AI from a high-level perspective, as well as what's going on and giving some kind of intuition. And, and then in the next lecture, we'll dive much more into details about how these different algorithms work and how we arrive at these models that we use. H, after that, we'll do an in-depth analysis of ChatGPT, like a case study. Then we'll do a similar case study on image generation and Stable Diffusion, right? And these four first lectures will be very similar to last year's offering. But then we're adding on January 23rd, we're going to talk about kind of emerging foundation models, right? It's going to be a combination of them existing out there, probably not one single model to rule them all. So we're going to talk about that, especially how that looks in, uh, industry and corporate settings. Uh, so Professor Manolis, uh, will come as well. He's an expert on, uh, biology and genomics. And then also Artem working, who's, uh, an MBA from MIT. He will talk about autonomous agents. And then we'll have a fun, kind of, or it's going to be a fun, hopefully fun lecture on AI and ethics, which is, of course, perhaps a little bit more fuzzy, but we'll also bring in regulations, what kind of, what's happening in terms of the institutions regulating AI. H, and after that, we also have a panel with Manolis and Artem. Uh, so should be fun. All right, so what will cover? Well, we'll cover all the buzzwords and neural network, supervised learning, representation, unsupervised learning, reinforcement learning, generative AI, foundation models, self-supervised learning. And we'll try to put together with a lot of applications, a lot of intuition, because it really, you know, should be non-technical. And I think as well, I want to try to explain things in, in simple but true, you know, deep ways. And I think if you're not able to explain something in a simple way, you're actually not doing a good job explaining it. So that's what we're going to try to do. At least today, we're going to give you a short, succinct answer to, what is the secret sauce behind foundation models and generative AI? And then when we've done that, we're gonna ask, how's the world structured? Because how we think the world is structured, uh, really influences how we learn in the world. So we're going to kind of explore that from a more philosophical perspective and see how that actually leads us to, uh, foundation models and generative AI. And then we'll at the end, we'll cover two applications of how we can use this in, in, in both research and in business. Again, right, we're going to try to use intuition, examples, and examples from both sciences and business. And hopefully, you'll, you'll understand why the hype actually is real. I said it's last year, but it is real. And maybe we understand what's actually just hype and what's the, the kind of more foundational aspect of it, right? What matters? Okay, so I think in trying to, uh, understand and, and, uh, you know, have a, in understanding of AI, it's good to use ourselves as a reference frame, right? We're human beings. And I think one of the core questions that we can ask is, well, at some point, right, you're born, you're a blank slate baby, fairly useless. You grow up, you interact with the world and different stakeholders in the world, and you acquire knowledge, and you become a fairly useful, knowledgeable adult, right? In terms of AI and learning, right, one of the key questions is, what happens in between you being a blank slate, fairly useless baby, to you being a knowledgeable, useful adult? It's one of the key questions, and we're going to use ourselves as a reference frame here. So let's consider a few candidates that might be responsible for giving us most of the knowledge that we have about the world. Is it our parents? They, they give birth to us, they raise us, they give us a lot of our core values. So they're definitely good candidates to consider. Is it perhaps our DNA and our genes? They're responsible for giving us a lot of our physical characteristics, also most some of our personal characteristics. But maybe it's, you know, nurture versus nature is very flexible. People can come very different based on where they grow up. And also maybe DNA works on a more generational scale, sometimes again, very delayed feedback loop here, not a lot of learning happening in a short-term scale. Maybe it's, uh, academia, right? Maybe it's teachers and professors and, and the educational system. It's supposed to be to educate us, it makes us make a use, make us useful, learn different skills, and also know how the learn how the world works, right? So that's also a good candidate. And lastly, maybe it's more kind of our immediate environment, and we, maybe what's more important is our goals that we want to be loved, we want to be happy, we want to be successful. And by optimizing those objectives and goals, we learn about the environment in the process. You know, we learn how the environment and different components and environments help or doesn't help us reach our goals, and therefore we learn about it. Okay, so we can split this apart a little bit more and high-level perspective, see how they correspond to different disciplines within AI, right? So this parent, child, student, teacher, leadership, that's supervised learning. It's supervision. So here, there's a human expert, like a teacher or a parent, that puts the world in order for you. It structures it, it labels it, and puts it in a condensed format so you can learn from it. And on the other hand, this more delayed feedback, like kind of evolutionary, and you're interacting with the environment, optimizing some goals, and those goals are what we're mostly focusing on, right? That's reinforcement learning. And as this delayed gratification and you're optimizing this goal, uh, in an environment, but it turns out actually that none of these approaches are responsible for giving you most of the knowledge that you have about the world. In fact, you can thank yourself for that, because most things you know, you learn by yourself. So how is possible? Well, um, it's possible by defining meaning by the company it keeps. So and learning from observing the world. So let's take the, the example of a dog. You actually don't learn what a dog is from your parents telling you or for your emotions guiding you. You learn what a dog is by observing dogs in different contexts and correlating and contrasting dogs with other concepts. So a dog is something that's walked by an owner with a leash. It's something that has an antagonistic relationship with cats. It's something that chases free space when those are present. And, and this is this kind of relational context that allows to understand what a dog is. And as you learn what dogs are by correlating, contrasting dogs with other concepts like cats, for example, you intern learn what cats are, right? This is where you get this relational understanding of different concepts. And, and even language can be learned this way, because the word or name dog would be uttered more in contexts where dogs appear. And if you think about it, if a, you know, a little child points at a dog and asks his parents, like, hey, what's that? Like, all the parent is really doing is giving some label. And the, and the child having the ability to generalize that label to all dogs shows how has a very, very robust understanding of an entity already, a concept of a dog, just not the name yet, perhaps. But that entity concept is got, is gotten by just observing dogs. So, for example, let's take the, uh, cat and mouse here. So I think intuitively for ourselves, when we hear cat and mouse, there's a strong correlation, and we can understand a cat in terms of mouse and vice versa, right? A cat is something that chases mice, and a mouse is something that's being chased typically by a cat. But when it comes to, when it comes to a mouse and a dog, I think that the relation is less strong and less obvious, right? What's the relationship between a dog and a mouse? It's not completely clear in our kind of contextual understanding of the world. So let's take this observation and, uh, fit it into this foundation generative AI models that also have this similar understanding of meaning, right? So a lot of this, I mean, basically all the, the images you see in this lecture is produced by generative AI. So let's ask, uh, the generative AI text-to-image to generate a cat playing with a mouse. Right, this makes sense based on how it's trained. It understands and generates a mouse, you know, that's, or a cat playing with a mouse in some sense. But if you ask it instead to generate a dog playing with a mouse, it gets confused because it doesn't really make as much sense in, in a contextual relation understanding that this model has of the world. So it generates a, you know, a mouse-looking dog playing with a computer mouse. So this shows that like when the context is not clear, meaning is not clear. And this is fundamental to how this model has started to learn about the world. And it also corresponds to how we think about it, probably. When I said cat and mouse to you, like you immediately collapsed to a meaning space where there was obvious what a mouse was. You didn't even think about a computer mouse when I said cat mouse because it just contextually made meaning clear for you. Okay, so basically, more relations leads to a better understanding of meaning. You understanding what love is, for example, helps understand what a dog is, because an owner loves his dog. So you know, why not train a huge model with a ton of parameters that's able to compress a lot of these different relations on as much data as possible to learn as many relations as possible, right? So you get the most precise understanding of all of these concepts involved. And then you use this model basically anywhere where any of these concepts appear, right? And that's exactly what a foundation model is based upon. And somehow also, you know, your brain is a prime example of some kind of foundation model. You get a world, you know, model and knowledge model about the world from just observing it and building this relational structures. And this is also what's, you know, what's behind this current AI evolution that we're seeing. This is the building block that is, uh, fueling all of the different breakthroughs that we're seeing right now. Okay, so this might sound kind of intuitive and make sense, but why did it take such a long time to get here? Well, this is where I think is quite useful to go on a little bit of a philosophical digression and, and I mean, provide a little bit of my, uh, uh, big picture on how this happened, and also it's shared with other, other researchers, some of them. So, um, I think this kind of rests a little bit on two kind of opposing perspectives about how the world is structured and how we learn in the world. So I'm going to try to, uh, get a sense of by giving some examples and contrast some stuff. So on one hand, we have learning versus designing. Where designing is this, you know, like a clockwork. We know how every piece works, we put them together. Every piece has like a, a perfect role to play. We know what's going on, and it's built with a specific blueprint and purpose in mind, right? That's how we design things. And on the other hand, learning is something that we perhaps not fully understand, and we're not completely conscious about it. It just happens. We just get better at something as we're exposed to the phenomena more and more and more. And there's really no, you know, end goal in mind before we start learning, right? We just adapt, uh, and it's very, very flexible and intuitive. Um, okay, similarly, we have chaos versus order. Where order is kind of what exists on the scale of planets or atoms, where things behave according to beautiful, simple rules, like math, physics, and beautiful theory, right? What chaos is, perhaps, you know, the more unpredictable reality that we find ourselves as human beings, right? The animal world, the human world, where things are unpredictable and, and chaotic. And one of the core questions here, as well, is that, well, you know, the chaos that we experience in our everyday life, is that like, is there some simple order behind all of that? If you just find that order, will all of our experience make sense? Or is there perhaps a limit to the order of the world, and how much equations actually can explain, and do we have to deal somehow with this chaotic world in another way? Similarly, we have this perspective of bottom-up and top-down, right? In a, top-down organization, there is a boss or some, you know, top person that's able to come up with, uh, a nice, uh, framework for how things should be done from the top, just by analyzing data, and then push that through throughout the organization to everybody involved somehow. H, on the other hand, a bottom-up organization really then, it's really necessary to have a lot of people at the bottom that interact, surprise, with customers and products and, and deals with all these different, you know, chaos that happens in all the particulars. And there's no real simple top-down decisions that can be made to make your business really work. You need to account for all the particulars, and it really matters how you engage with a customer on a, you know, personal level. It's not enough for the boss to come up with some 10 simple rules for to solve these things. You need to deal with this chaos. H, and also a lot of the wisdom in a, in a bottom-up organization comes from really listening, uh, to the people that are closest to the end consumers. Okay, uh, so I think at least in the Western world, we have had for quite a long time this, uh, designing, ordered, top-down perspective. And I think this is kind of due to the ancient Greeks, right? So Socrates, for example, had this allegory of the cave, where, uh, basically human beings are have very imperfect senses and a very kind of imperfect, uh, understanding of how the world really works and what's going on. And they gave this, this allegory of the cave, where human beings are actually here on the left-hand side, looking, you know, at the reflections of the cave. So basically, our experience of the world is so untrue and imperfect, so we don't even get to experience the world firsthand. We get to experience gods walking with depictions of real objects, right? And we don't get to see that even. We get to see the reflection of those depictions through, you know, a fire on the cave wall. That's how distorted our view of reality is. And for Socrates, like, well, we have to accept that we have very distorted, imperfect experience of the world. We have to always strive to understand the real, true world, the beautiful world of the gods, which exists, you know, outside of the world we get to experience. So I think like for, for Socrates, a dog, for example, like there is a true perspective of a perfect dog in this godly world that we should try to understand. And all the variation that we're seeing in the real world are just some kind of imperfection from our sense, and we should strive to understand the true dog in the, in this perfect, absolute world. So also, I think it makes sense because the Greeks at this time, they were discovering mathematics. And in math, for example, there exists a perfect circle that obeys very, very simple equations. But every time you take a perfect circle and try to recreate it in the real world, it's always off, it's always imperfect. So like this, this kind of correspondence somehow influenced Socrates' thinking and also makes sense in this kind of mathematical perspective. And I think this has been extremely, uh, fruitful for us. It's led to the kind of golden era of design. We've had a, a scientific revolution and industrial revolution, right? We've had modern math, physics, and modern medicine. And we even went to the moon with this design way, you know, top-down, ordered way of thinking. So it's been extremely, extremely good for us. But, you know, assuming that this top-down, ordered, designed way of thinking is the be-all of our existence, how come we're not better at it? Right? We've had billions of years of evolution here on this Earth. How come we're not more like a computer or calculator? Right? If, if we can just try to find the simple mathematics and order behind the world, and we'll be able to perfectly exist in it. And I think this is, um, kind of a strong indication that there actually is a limit to what order can explain, because, and we're not in, like, we're not intrinsically very good at logic and math because we don't live in a top-down ordered world. We live actually in a bottom-up world of chaos where math is not that useful. Instead, in a bottom-up world of chaos, what's useful is intuition, flexibility, and speed, the things that we actually are good at, right? So in this order versus chaos perspective, when it comes to our everyday interaction as human beings, uh, you know, besides the scale of planets or atoms, actually the world is chaotic, and we cannot escape that fact. We have to deal with it. There won't be some just simple equations that explain everything for us that we can rely on. We have to deal with all this chaos that is somehow unavoidable. So what can we do? Right? You just give up? So if, is there an instrument that can help us to contain and navigate all of this chaotic world that we find ourselves in? Well, if we had billions of years of evolution and we didn't create a computer calculator, or nature didn't, then what did it do? It created a brain, which is our best tool of navigating and making sense and learning in a world of chaos. And then the neural network in artificial intelligence is just our best attempt of replicating the brain inside of a computer, right? So it's very, very flexible and adaptive. It consists of a ton of kind of neurons or parameters, H, and it's very, very simple computations, but done in a hierarchical scale. Uh, and also it's extremely slow to train, but very, very adaptable and flexible and fast to execute when actually learned something. Okay, so we now have accepted the world as chaotic, and we have a tool that's able to, uh, still function and navigate in a chaotic world. How do we, how do we use this? Well, the thing is that these neural networks still exist inside of a computer, right? And a computer only speaks the language of code and math. H, so still, there is a divide now where we have the real chaotic world, and we have a computer side of machine. So we somehow have to go through the world of order and math to tell the brain inside of the computer what to to optimize for or what to focus on, right? To use this brain for, we have to somewhat describe this in a more exact way for it to be useful. So we still have, we can't still completely ignore the world of order. So how can we define such objectives rigorously? H, and that's what we're going to try to go through right now. So first thing we can do is to say that, well, we understand how the works, work, world works, right? We understand how things work, and we have a lot of knowledge about the world. So why don't we just impart that knowledge onto the computer? Why don't we just structure the world in a way that makes sense for the computer? We label all of it, and then we can feed that information to a computer so it can start learning from that, right? That's, that's our first attempt, and that's supervised learning, where we structure the world, we label it, and then computers can learn from us. Um, and, you know, pretty immediately, we run into some problems. First off, of course, this scales with human labor, human experts, and labels. That's why, you know, you have these outsourcing centers, people label data constantly. And can everything even be labeled? Like, do, are we that kind of self-conscious about, like, are we that conscious about how things work? Like love, for example, can we label concepts like love? It's probably going to be quite hard. And we maybe overestimate our understanding of how things work and how we can, uh, isolate concepts. And, uh, really, you know, maybe things again, like love or other things are not very labelable. And maybe things are not as categorical or like distinct. Maybe the world is actually more continuous in a sense, and only in the limit of unlimited number of labels do you actually start to understand the world is really structured, right? Maybe as, you know, maybe have a set of labels and you learn from them, but you will always find these, uh, points in between where labels don't really explain what's going on. For example, here, right, you know, is this a, is this a dog or a cat? The left-hand side, I mean, somehow at some points, things get more and more close, and we, we reach a limit to how much we can label the world. Um, and then we just need more and more labels actually to make sense of it. So because of this, somehow, uh, we see that the supervised learning doesn't by itself generalize well enough. It just doesn't work enough. It's, it's too expensive in terms of having human, human people, expert label data. It doesn't generalize really well to, uh, you know, the diverse setting of the real world. Okay, so let's say now that we try to rely on ourselves defining things, structuring the world, and we do almost the opposite direction. So we say that as human beings, we have goals, we have desires that we want to optimize, or, you know, that we do optimize, and we hope that if you just impart those goals and desires onto the computer, it will learn about the world in the process, right? Focus on the end, like where you want to end up, and the computer has to figure out how to get there. So this is reinforcement learning. But I think, you know, if a good, uh, analogy for a blank slate computer, it's like a blank slate baby. I think it's very, very hard for us to even understand what it means to be a complete blank slate. I think a lot of our knowledge is so intuitive, we just take it for granted. So what does it actually mean for something to be a complete blank slate, like having no understanding of how the world works at a starting point? Um, so let's say, you know, you, you know, you're this blank slate baby and no understanding how the world works, and you want, you want to optimize certain goals, like maybe you want to optimize success or, uh, becoming rich, right? First of, there's a huge delay in the feedback, right? You do something, you won't immediately know if it's actually helping you or not. You have to wait a long time before you get a signal, right? So it's very, very, you know, difficult to know what's working and not. You need to keep at it for a long time before you, you, you see that it makes a difference, right? But even if you take something that's perhaps a little bit more immediate, like trying to become less hungry, it's still some delay in terms of maybe minutes or something. But let's say you, you try to reduce your hunger, right? But you have no understanding of how the world works. You just randomly pick a concept that you're observing, and you start to explore how that concept affects your ability to become less hungry, right? To become full. Maybe that's, you just start focusing on the moon and how the moon and the characters around the moon affects your ability to become, uh, less hungry, right? You're going to spend so much time, uh, exploring complete nonsense that has basically no relation to your ability to become less hungry, and you're going to die out of hunger way before you make any progress whatsoever. So I think that's kind of, that's kind of hard to even understand. Let's take the, the example of a car that's a complete black slate. You want, you know, you want the car to learn how to drive to your home using reinforcement learning. I mean, again, maybe it starts exploring how the, you know, moon affects the ability to reach home, and it will make no progress. But even if it starts focusing on things that are actually relevant, but just share coincidence, and it starts focusing on other drivers or other human beings in traffic, right? Still, we cannot afford the car to hit like a million human beings and, and crash a million times before it actually reaches home, get some signal, and starts making some progress, right? This is too expensive and too risky in real life. Uh, if you take, you know, imagine putting a real baby in the driver's seat of a car. I mean, how many billions of years, if this baby would live forever, how many billions of years would it take for this baby to just by coincidence reach home? I mean, it's going to take forever. And then when it does that, you say like, hey, good baby, awesome work, you know, here's your signal, now do it again. I mean, this is somehow how reinforcement learning works when there's just a delayed feedback and this blank slate. And this, I mean, this works better in, in chess or something where there's an idealized set of rules and the state space is much smaller. But this kind of just blows up in, in real life. So what we need is a basic model of how the world works that we get from just observing the world, because it's the only thing we can really afford, right? Reinforcement learning, as we just said, is too dangerous, you'll die before you make any understanding of how the world works, and so it's too risky and too expensive and too slow. And supervised learning is also too expensive because it relies on human experts to label the world, and at the end of it, it doesn't work because you can't label the world. Okay, again, so what comes to rescue is this, you know, breakthrough behind foundation models called self-supervised learning, where you learn from self-supervision by just observing the world, and we define meaning by the company it keeps, which allows us to learn from just observing. Okay, so quick recap. Uh, we talked about chaos versus order, and, and we conclude that actually the world is chaotic, we cannot ignore that, we have to deal with that. The brain is the best tool that we have to learn and compress and navigate in a chaotic world. But we still have to, uh, define how this brain should interact with the world. Supervised learning didn't work because it's too expensive and doesn't generalize, like labeling the world. Reinforcement learning is too dangerous and too slow. So that's why we end up with, uh, learning from observation and self-supervised learning. Okay, so let's dive into some specific use cases of this. So, uh, here's a kind of a collection of different, uh, self-supervised learning algorithms and how they learn from data, from just observing the data. We'll cover all of them in, in subsequent lectures, uh, but what they all rely on is learning from data itself, right? There's no need for human experts in the loop. So it scales extremely well to unlimited amount of data, and then they use these very, you know, broad capabilities in a wide range of of different tasks. And in this lecture, we'll just talk quickly about predicting the future and positive pairs. Okay, so this idea of learning by predicting the future based on the past relies on this idea that in order to predict the future, like we need to understand the past. So let's take an example of language model or learning, you know, how learning from text data. So this is good because we have an unlimited, basically amount of text data from the internet. So we can just download a sequence of text and then we can remove the last word and try to predict the last word based on previous words. That is extremely simple to define, and then we can let the huge, huge brain of a neural network start doing this task. And let's now assume that this neural network or computer has become really, really good at predicting the next word based on previous words. What does that actually imply? Does it, does what does it learn? Does it learn grammar? Well, in order to generate and predict grammatically correct sentences, it has to understand grammar, of course. I mean, does it have to understand the meaning of words? I mean, if the sentence that we give it is, you know, "the dog is," it needs to understand what adjectives describe the dog. So also need to understand the, the meaning of words. Does it have to understand the difference between an informal social media post or a formal news article? Well, you know, if it's right now being fed a, an informal Facebook post and it wants to complete that accurately, it needs to understand what kind of language that corresponds to, and vice versa, right? And similarly, like if people in the, if the way people write changes based on their political beliefs, they also start to need to pick up their political beliefs based on how they write things, right? Which only gets kind of scary, but a really, really powerful model can start picking up those things that are so implicit because it, it's optimized to do this, this fairly simple task. And again, if I ask it, you know, "What's the capital of Stockholm? Question mark," right? I give that sentence to the, uh, model, and it has to complete it accurately and optimally, it has to give the correct answer, which is Stockholm. So it also becomes very knowledgeable about, about the world in a bigger sense. And this is the core, you know, approach behind ChatGPT and why it works well and why it's so flexible and capable, right? And how such a simple objective can lead to extremely, you know, broad sense of intelligence. And, uh, again, you know, similar things, uh, applies to real-life examples like frames in real life or in a movie or something, right? If a model is able to say that, okay, it sees a human being with a leash and a dog, and in the next frame it sees a frisbee, if it's able to predict that and say, well, probably the human being will throw the frisbee and the dog will run to to catch it, right? Being able to do that, you know, prediction to combining these objects shows that this model understands how these objects relate to each other. So it's extremely, extremely flexible and powerful. Okay, another approach that's very popular in vision is it's called positive pair contrastive learning. Here we learn, you know, here again, we can just say we say that we think that objects that appear in the same image are more related than objects that appear in different images. So then we can just download a ton of images from online, and we can just kind of randomly crop the images and push the crops from the same image close together and far away from crops of other images. And if we do this, we're going to see like, well, okay, so a human being, this on the left-hand side, right? The crop here with the human being, a leash, a dog, is going to push a dog kind of close, meaning to a human being with a leash and a dog and a frisbee being pushed together somewhere because they're more related. Well, also the really, really cool thing here is that it kind of rests on this assumption as well, that that you don't even have to appear in the same context to be understood. It's like people that have similar friends or similar people. So here again, like, okay, the frisbee and the human being in the leash don't appear in the same image, but they appear together with similar objects like a dog. So someone as well, it starts, it will start to capture the relationship between these, which can be kind of more abstract and, and more isolated, like more far away in some sense, but it still captures that in a, in a very robust sense. Okay, so let's apply this. So we're going to talk from one example in science and another in business. Okay, let's say we want to apply this new paradigm to genomics. It's a good setting because we have an abundance of DNA base pairs, like this basically this tech sequence that we get from people from sequencing people's genomes. So what we've done, and what we could do, is that we could use our own human intelligence and look through this data and see like, well, there seems to be these recurring sequences, like genes, for example, that recur, and we can look at these evolutionary trees, how these things are being passed on, etc., to build up a structure and start building features around parts of the genome, and then use that for, for example, protein structure prediction, like start understanding how the genes work. Or we can just kind of rely on, on self-supervised learning and just try to predict the next DNA base pair letter based on previous ones. So let's say we train a huge model on a ton of DNA data to just predict the next DNA base pair based on previous ones. Like, what does it learn in this process implicitly? So if it's able to do this really, really well, like again, does it learn the meaning of genes? And genes, yes, so genes are just recurring sequences. So if it's, if it's wants to be able to like complete this, uh, the genetic sequence really, really well, it needs to understand, like, identify, hey, I'm inside of this gene right now, and that's why I need to generate these new letters, right? Does it need to understand as well implicitly if it's a, if it's kind of completing the genome of a dog versus a cat? Yes, if they differ, it needs to, because it needs to, is it going to change how it, how it thinks about predicting the next, uh, uh, base pair? Uh, and it turns out there actually creates really, really good features that maybe it takes this awesome time for human beings to should understand exactly what they encode. But if you then use these features and, and you train another model to predict protein structure, right, what kind of protein structures DNA will lead to, it works really, really well. Okay, another example, uh, in the business case that I was involved in in my startup here. Um, let's say you come to a retail company, and they want to understand their consumers, and maybe do, you know, have a better assortment and give recommendations, etc. So typically, what they do is that they, uh, come in, take some consultants to start trying to define some user profiles, right? If it's, uh, young single people or, or families with kids, and they try to build that up and kind of label people and, and consumers, and they do a lot of questionnaires and stuff where they ask customers like, hey, who are you? You know, what do you prefer, etc., to build up some kind of knowledge structure. So first off, right, there's a problem because every, every human being is unique. So every time you, you enforce some kind of user profiles, you typically lose a lot of actual understanding and performance of your models because they, uh, rely on very, very coarse, uh, information and, and, and structure. And also, it's actually very, very hard to ask people what they want and what, and how they work, because they don't are not completely self-aware of what makes them make a certain decision and what they would prefer. And, and it takes also a lot of, lot of work, like a lot of manual work to map up all these people, ask them questions and questionnaires to start to build up this understanding of your consumers. So what we can do instead is start to kind of look at the behavior of the customers and let the customers do the work for you by just acting in your channels, right? And, and tracking some of their interaction points. So, uh, here, for example, right, we have a, on the top, somebody buying a wine bottle, a cheese, and some chocolate. And if the model is kind of good at understanding predicting, for example, the next step of a consumer, it can say like, and, you know, in the bottom, we have a soda and a candy. It can perhaps start predicting like, okay, the top person is an adult or something that has that wants to relax on a Friday night, and the, the bottom is a, a child equivalent version. And then if he looks at some behavior of a, of a family coming into the store, it can kind of recommend both, maybe, because if it's a family, you know, it has some, both of these features. And this, like the capabilities of these models, when you train them on enough data, becomes extremely sophisticated. And it's very, very cheap to track data rather than building up all yourself from, from human, you know, human work. And also, right, some of what this eventually gets to is a more deep understanding of your consumers, like your products and the customers, right? How to interact. And if you're in retail, that's basically all of your business. It's understanding your products and your customers and making the best combination. So then it starts building up this understanding, it can use it in a, in a broader sense across the company for business intelligence, etc. So, so by building a more deep and general intelligence, make it more applicable, uh, broadly in the company, instead of, instead of approaching every problem in isolation, it doesn't scale that well. Okay, so summarize, we, uh, we started with a short, uh, answer to like, what's the core fundamental change in perspective behind foundation models and generative AI, right? Learning from observation. We, uh, started to ask like, how's the world structured? Chaos versus order, top-down versus bottom-up, design versus learning. And we said there is a limit to order and design, and neural networks are a way to compressing and dealing with this chaos. But as we define these different objectives, we still end up with self-supervised learning and learning from observation as somehow the, at least we know now, a viable option forward. We have to learn from unlabeled data directly and learn from observing rather than interacting, because that can be too dangerous and expensive. Uh, and then we also covered two applications, uh, one in science and one in business. Okay, and I want to leave you with this picture. I think, um, right, the meaning is relational. And also maybe ask yourself, like, do you have implicit self-supervised learning algorithms going on in your own head? Right, this is how you learn. If I would take this clicker, for example, and I would just drop it, and it would float in there, I mean, you probably be upset or even, you know, surprised and upset this is happening, like almost subconsciously, because you probably have this mental model in your head that's always trying to predict what's going to happen next based on previous action, right? And by doing that and slightly adjusting when you observe the world, you implicitly learn about the whole world in the process, right? You have these algorithms going on in your own head. Okay, next lecture, we will, uh, make this much, much more concrete. We'll go through different algorithms, and that should be very, very exciting. We also have a, a website, futureof.ai.mit.edu, where you can get all the, the updates and other good stuff. So thank you super much, and if you have any questions, feel free to [Applause] ask. Well, the intuition on the first two models that you said, but in the case like, model didn't really get intuition behind it. You got the intuition of the two first, uh, predict in the future and, uh, positive pairs. Did you understand the intuition there a little bit? But you didn't understand self-supervised learning. That's exact. Self-supervised learning. These are those are two examples of self-supervised learning. So I mean, actually, the first time we offered this course, it was called "Foundation Models and Self-Supervised Learning." So self-supervised learning is how you train, uh, foundation models and generative AI, right? Foundation models, generative AI is the output. So those, those two examples are self-supervised learning. So self-supervised learning is a family of algorithms that give you generative AI and foundation models. Does that make sense? So actually, I think maybe, yeah, pointing it out is important because self-supervised learning is the approach and how we train these models, right? And when we train them, we call them foundation models. So ChatGPT is trained using self-supervised learning. So it gets a little bit technical. I mean, if you had like Yann LeCun, for example, he said like, oh, just call it self-supervised learning. Why call it foundation model, generative AI? But then it's like, oh, let's call it foundation model because that's more, sounds better than what it is. And there's some definition from Stanford that they use. But then the media is like, well, generative AI sounds cooler, stuff like that. So there are terms as well that are different based on, based on the context. But all those, those two examples you understand are self-supervised learning, that's examples of it. Thank you.