Transcription
Yes, >> a brief introduction to elephantology in two volumes. There is artificial intelligence, just synthesize a voice and there you go. >> This very well shows the real state of affairs in artificial intelligence technologies. You can generate fakes, you can detect fakes. The paradox is that in the era of such rapid technological progress, human qualities are precisely what become most in demand. Well, this is some kind of synthesis, yes, of science and art. [music] [applause] With us today on the podcast is Sergey Markov. Sergey Markov is a researcher in the field of machine learning and artificial intelligence. I would combine it all into artificial intelligence. Head of a team of researchers at Sber. He is the author of one of the strongest chess programs of its time. You correct me, Sergey. If, by chance, it was a long time ago, >> it was a long time ago, but one way or another, in its time, it was one of the strongest, author of the book "Hunting for Sheep". >> And Sergey, well, the floor is yours, perhaps you will supplement my introduction. >> Oh, I should start with a disclaimer right away, because I am, of course, not a real scientist, да, because I also, I remember, I performed at Scientists Against Myths, and they told me there: "Sergey, what are you doing, you are not a real scientist." Well, yes, I have some kind of embarrassing Hirsch index in general, probably, да, around five. So. But, uh, for the last many years, I have mainly been involved in leading teams of researchers and, to a greater extent, even teams of engineers who build a bridge between research and the creation of applied products and services. That is, it is applied science. And, um, of course, lately there have been more bold, so to speak, hypotheses that we are testing, but in general, I have been involved in artificial intelligence technologies for over 20 years. I am also involved in science popularization. And my book, which I worked on for 6 years, is one of the products of this activity. That's it. Besides this, I give many popular science lectures. That's it. And, well, I read a lot, let's say, scientific literature, much more than I write. And, well, since 2016, I have been actively engaged in lecturing. >> Let's show the book. I, well, it so happened that I have it. >> Let's, yes, and we will also raffle it off at the end of our show. >> So that everyone understands what 6 years means. This is this, this is this book, a two-volume set. It's very beautiful, it's brand new. >> Yes, >> a brief introduction to elephantology in two volumes. That's it. Uh, well, in general, >> the book is heavy. Heavy in weight, not heavy to perceive. Sergey, are you a fan of Blade Runner? >> I'm not a fan of anything in general, probably. Fanaticism in general is very, you know, contrary to my nature, probably. That's it. But in general, of course, I love cyberpunk classics, I love Philip K. Dick, but, by the way, the title of this book is a reference to two works at once. The second one, perhaps less noticeable, is a reference to Murakami's "Hunting for Sheep". That's it. So, perhaps, this is to some extent an indication of the combination of both physical and lyrical components, да, because in this book, in fact, quite a lot is written about the people who created modern artificial intelligence technologies and about their creativity, about their relationship with the surrounding world, starting from economic to political. That's why, uh, well, this is some kind of synthesis, да. of science and art. >> What prompted you to write this book? How did you come to it? >> Annoyance, probably. Annoyance at what is written in the yellow press about artificial intelligence. And, in fact, the last straw was a lecture by a science popularizer from the field of artificial intelligence. I won't mention his last name, just in case, the person is actively alive and active, but he gave a lecture in which he touched upon my field, chess programming. and the work of, well, in general, a person whom I knew well from forums, with whom we corresponded a lot, and, I was aware of his work, да. And, in general, in a very popular lecture, everything that could be distorted was distorted. And, uh, I got angry and decided that things couldn't go on like this. We need researchers to talk about their work to a wider audience, because otherwise, well, people will draw conclusions about our industry based on such poor popularization. That's it. Therefore, this was probably the last straw why I started doing popularization. That's it. And how do you generally feel? We will talk a lot today about artificial intelligence, including LLMs and all related terms. How do you generally feel about the terms that are introduced? You once said and mentioned that even in the book, I think it was presented about the term artificial general intelligence. It's quite commercial, one might say. >> Well, I would say that it is not so much commercial as conceptual and philosophical. In general, Mark Gubrud in his 1997 article, dedicated, by the way, to nanotechnology and international security, he first proposed this term there. Why was it introduced? It was introduced to denote a certain pole in technologies, да. All existing artificial intelligence systems at that time, they undoubtedly gravitated towards the pole of narrow artificial intelligence, да, or applied, weak, да, there are many terms to denote this edge, да, but it was necessary to somehow denote the other edge, да, absolutely hypothetical at that time, да, that is, systems that could solve an indefinitely wide range of intellectual tasks, similar to how the human mind does it. And, uh, this term was introduced, which, in general, well, later took shape in business literature, you know, in a very simple formula: these are systems that can solve any intellectual tasks accessible to humans. That's it. But, uh, it is very important to understand that the time when this term was proposed, it was, well, from a technological, not a temporal, but a technological point of view, very far from creating, well, even distant prototypes of such systems. That's it. And as we began to move in this direction, да, and large language models appeared, fundamental models in general, да, then a number of questions arose regarding this definition, because it was underdefined, да, because, well, okay, uh, they can solve any tasks accessible to humans. Which humans? Humans are different in general. There are experts, there are non-experts. Who are we orienting ourselves towards? Are we orienting ourselves towards a task that, well, maybe there is a task that only one person in the world can solve. Should we include it in the benchmark, да, so to speak, or not? That's it. The second question. What does it mean to solve a task? да, if we are talking about finding the roots of a quadratic equation, then everything is clear, да? But if it's a generative task, да, let's say, I don't know, write a sales text about colored plungers, да, and the machine will write such a text, how good is it? Can we count this solution and say that success has been achieved? да. What it means to solve successfully is not entirely clear, да? The next point is, within what resources should it be solved? Because, well, it's very easy to write a chess program that will make perfect moves. Well, you just need to write a program that iterates through all possible moves, да, and then chooses the move with the maximum evaluation. Well, and it will absolutely make perfect moves. The only problem is that it takes billions of years to find this move, да? That's it. Uh, and should we admit, да, that such a machine that solves chess in this way, да, has achieved success, да, it can play chess, да, but such a program could have been written in the 1940s, >> and say: "Everything, the task of playing chess is solved." This program will beat any human without time constraints. да. That's it, and so on, да? So, as soon as we try to turn this definition into a concrete definition of how to evaluate whether a created system is HI or not, we run into problems. That's it. And, naturally, what began to happen? Well, what usually happens? Different research teams rushed to redefine this definition. That's it. And, uh, the culmination of this, I believe, was this famous leak from the correspondence between OpenAI and Microsoft, да, where they agreed that AGI is a system that can earn 500 billion dollars. да, that's it, uh, it always reminds me. No, it immediately reminded me. Remember in The Little Prince there is such a conversation about how adults are not impressionable. And, uh, how do you explain to an adult that a house is beautiful? да? You need to say that this house costs a million dollars, да? Or, well, in Exupéry, the amount was, of course, not in dollars, but that's not the point, да? And, uh, then the adults will say: "Oh, yes, that's a wonderful house." That's it. But it's both funny and sad, in fact. That's it. And I believe that this should not be done, because, well, this is already a purely marketing story, because it is clear that, well, the company that first announces that it has created AGI, да, then its statement, its statement will be weighty enough, it will, its shares will skyrocket, да, it will reap a rich monetary harvest. That's it. And therefore, of course, okay, well, in the coming years you will hear many announcements about the creation of AGI, да? But it will be AI defined in its own way, да? >> Yes. They defined it themselves. >> For the research sphere, this is an absolutely unproductive story, because, well, like, some people, well, have come up with their own definition, seemingly of a generally accepted term, and are now trying to use it for marketing purposes. That's why we prefer to say that we have the concept of practical AI, да? That's it, we are not redefining the terms AGI that everyone uses, but introducing, well, our own term, да, to denote the point we are striving for, where we are going, да, and, uh, we strive here to give a definition from which a clear definition of, benchmark, whatever, can be made. >> Have we gone through this stage of the AI revolution, or is it too early to talk about it? >> Well, you know, revolution is what? Revolution is some kind of serious qualitative leap, да, in the development of anything. That's it. And in this sense, well, like, everything that is happening lately, да, and what can we say, everything that is happening in the field of artificial intelligence, is a constant revolution, да? So, if we talk about, uh, well, >> well, if we define the two terms evolution and revolution, in this context, >> well, as always, problems with metrics begin, where, what, what jump is considered sharp. >> No, listen, well, I think that the twenty-second year, да, November twenty-second year, when ChatGPT was released, it was a revolution in some sense, because it changed not so much the research part as the perception of all these models and the general understanding of what such models are and what they can do in the whole world. This is, uh, the question of how much this revolution is technological in nature, and how much, well, uh, a revolution in public consciousness, да, and in my opinion, the latter, because, well, like, before ChatGPT, thank God, there was GPT, GPT2, GPT3, да. That's it. And everyone watched this progress. All specialists saw how these models were becoming smarter, how they were generating more and more meaningful texts, how they were becoming more and more universal, да, regarding GPT3, for example, the concept of prompting, да, appeared, then Gvern Brun launched a huge site with a lot of examples, there was a lot of enthusiasm and also, well, >> but when they released a service that everyone in the world could touch, да, so the main revolution was precisely in this, that, well, roughly speaking, if they had launched the same thing, but based on GPT3 a little earlier, or I don't know, given >> the guys from Google to launch their own, да, for example, or the guys from Facebook their own, what was it called, Bard, да? That's it, the revolution would have been the same, most likely, here's the factor that worked. It's the same Bren, he said, on the third anniversary, I think, or second anniversary, I don't remember, of GPT3, that no major company would dare to release a service based on these models, because, well, he explained it somehow, I don't remember, but >> the stock would collapse, he said that, >> well, intuitively it was a dangerous undertaking, because everyone remembered, say, 2014 with Microsoft and the Tay bot, да, which became racist, having learned from texts from Twitter. да, it's not surprising that OpenAI acted as such a bold player here, because, well, what was OpenAI at the time of the launch of ChatGPT? Well, of course, a cool laboratory, one of the best in the world, but a laboratory, that is, well, how much was it worth, how many people were there. Uh-huh. >> But a technological revolution, well, there were still important steps forward. You can't say that there isn't some technological step forward, because the creators of chatbots before OpenAI, they were very focused on the voice domain. That is, for them, a chatbot was always like a cheat-chat, да, that is, well, to talk about something briefly, да, there, well, like to support small talk. That's it. But what can you do with a chatbot to which you formulate the conditions of an intellectual task, and it will solve it, да, this is fully. This had unclear product prospects. That is, it's difficult to put it into a smart speaker, да, because, well, what will you listen to, да, two pages of text that will be told to you somewhere, that is, where, on what surface, in what form to launch it. A product like, well, like modern chats with generative models, did not exist then. And it was generally unclear, they say, whether people would buy this or not. That's it. And the brilliant, of course, insight was that let's still train the dialogue model not just to chat, but to solve some tasks in dialogue mode. >> That's it. And the second thing, well, they, by the way, grabbed onto it at first and said: "Oh, like, here's the secret sauce of RLHF." да. That's it. And here, the fact that shortly before the launch of ChatGPT, InstructGPT came out, and everyone started looking for, what, what has OpenAI published recently, and stumbled upon this work on InstructGPT and, and saw that there RLHF started talking about, да, they made a breakthrough, they connected with NLP, да, which had not been successful for a long time, да, that is, there were many works, many attempts, but they, it seems, managed. That's it. And they came up with, you know, this RLHF, that we will now fine-tune the model, predict human preferences. But in fact, now, after several years, we understand perfectly well that, well, this is, of course, a nice-to-have thing, but we now have a lot of models that are smarter than GPT3 and in which there is no RLHF. That's it. So, this is an important point, да? That is, it's not like, when ChatGPT appeared, they thought that it was impossible to make, so to speak, the same thing without RLHF. In fact, you can. That's it. RLHF is a very capricious thing. And this was, I don't know, one could even jokingly say, sabotage, да, from OpenAI, because a lot of people then spent time on this RLHF, and it's very capricious. That is, not just anyone can do it for you. You need really cool specialists to be able to launch this RLHF in some form, да? And the most important thing is that later it turned out that, in general, it's not necessary to do it, да, you can do something much simpler and the effect will not be much worse, and somewhere even better. That's why, but it's obvious that much more important technical innovations in the last, say, 15 years have been the appearance of transformers. Uh-huh. >> Why was it quite important? >> I remember how translators changed then. It was just some kind of leap-like change that, да, in general, just >> thanks to transformers, well, a lot has appeared, да, so many tasks that were previously solved poorly, and generative ones in the first place, and they started to be solved well. Why? Because transformers allowed for more efficient use of parallel computing, because, well, like, classic RNN, or even LSTM, да, they, well, like, to generate the next token, да, you need to process the previous one, да, always. And you can't calculate it normally. You need to calculate element by element, да, this loss. That's it. But here you get a model that processes all elements of the sequence in parallel. And therefore, you can now build a giant cluster and run a giant model on it to train, and it will turn out to be computationally efficient. It will turn out to be efficient. And, uh, so this model, it, like, made it possible to use, you know, distributed computing clusters of GPUs to create really large generative networks. That's it. What before, да, well, before deep learning as such, да, and in principle, well, like, there is the term revolution of deep learning. go, >> AlexNet in 2012. That's it. Uh, but with AlexNet, the story was like this. Well, like, it's not the first convolutional network, it's not the first model trained on GPU. It's not even the first model that achieved superhuman level in some task, because, I don't know, a year before that, someone from Schmidhuber's graduate students, as always, Dan Chereshan, I think, he trained a network to recognize road signs in a competition. True, the image size was small, I think, 32x32 pixels, if my memory serves me right, but a superhuman level was achieved there. Well, and it's clear that before that, on MNIST, it was, so to speak, achieved, indistinguishable from humans, at least. That's it. And, um, well, like, what exactly was the innovation? And the innovation was that, well, like, a mature pipeline appeared, да, that is, like, a mature technology appeared that united all this, да? Like, train a convolutional network on GPU, да, there, correctly, correctly initialize, take a lot of data, again, ImageNet. appeared. And look, we are now surpassing all other, confidently surpassing any other image recognition methods with a large amount of inductive bias embedded in them. да, we just now, look, we made a deep network. But behind the appearance of this stack were the efforts of many scientists, many researchers. And, uh, modern ML engineers, they often don't even know and don't understand what the secret was, why. Well, like, come on, recurrent networks or convolutional ones, they were conceptually invented a very long time ago, да? That is, well, in the sense that you shouldn't think that they appeared in 1986, да, when, say, DNN or CNN by LeCun appeared, because the convolution operation itself, well, damn, well, Neocognitron, Fukushima, да, all sorts of locally connected networks by Rosenblatt, the idea is roughly the same, described back in the principles of neurodynamics in 1962. That's it. So, these are all very old ideas, but they, uh, such networks began to be trained. Why? Because they learned to use GPUs for training. This is an additional boost. And in principle, machines became faster, but also, you know, processors appeared that can perform vector-matrix operations well and quickly. A lot of data appeared. This was a powerful limiting factor. In the seventies, well, or in the sixties, when Rosenblatt conducted his experiments. Well, there weren't that many digitized pictures and texts. Where from? да. That is, until digital photography appeared, until OCR appeared, until OCR, well, like, personal computers appeared, social networks appeared, the internet itself appeared. Well, that is, everything that led to an increase in the amount of digitized data, there were many steps. That's it. Then, in terms of models, it's not enough to make a convolutional network, well, or a deep one, you need to know how to train it. To know how to train it, it turns out that it's not just about backpropagation as such, but also about correctly initializing weights. This moment is often missed. And why, I don't know, in the early 2000s Hinton struggled with these Deep Belief Networks, да, and trained them layer by layer, да, in the spirit of how Ivakhnenko did it in the sixties. Why? да, because it, well, it didn't work. Uh, in a deep network, gradients faded very quickly. And so, what was solved? When correct initialization methods appeared, when the norms of random initialization in different layers became different, it turned out that this was the secret sauce that allowed us to move away from these complex, layer-by-layer stories, да, and immediately train a large network, simply by initializing it correctly. That's it. Again, modern gradient methods also took shape. Well, that is, many such stars aligned here, да, both in terms of mathematical models, and in terms of data, and in terms of computations, so that all this together resulted in the appearance of this powerful technology. That's it. And what can definitely be said is that, well, you can tie it to AlexNet or not, it's a matter of taste, да. In the last, say, decade and a half, many tasks that were a tough nut to crack for good old artificial intelligence, да, have been solved. That is, in machine translation, for a long time, well, in general, there was stagnation, да? Well, okay, here's our topic on rules, here's our model of meaning, text, we're doing something there, да, very quickly our ontologies swell, да, people can't control this complexity. There are some internal contradictions in this ontology, and that's it, the complexity threshold that people can't overcome. >> And only large, conjunctive models could overcome it, well, when they grew, да, when they became large. That threshold. Otherwise, well, and many, many other tasks, generative ones, please, да, transformers appeared, a breakthrough occurred in the field of generative artificial intelligence. That's it. And, well, I hope this is not the last such serious breakthrough. >> Ah, well, now, the current one, да, what kind of breakthrough can be considered the appearance of reasoning? Correct? Can it be attributed to some kind of breakthrough? The very idea of reasoning and, say, Chain of Thought, it's also not new. And it, in fact, appeared even in the pre-transformer era, strangely enough, да? That is, even with those old language models on LSTMs, the idea was there. Uh, and moreover, the story with reasoning, it, and in general with chains of reasoning, it is itself a crutch. Why? Because, look, if you look at a modern transformer, it actually consists of two parts. Well, the first is connectionist, да, this model that generates the probability distribution of tokens for you. And the second is sampling, which, for a moment, is a symbolic mechanism of a set of rules, да? Well, that is, like, now we have nucleus sampling, this and that, these are non-differentiable parts of the model. That is, in fact, a transformer is a neuro-symbolic model, when it's there, and when it's there, will neuro-symbolic models appear? Well, we are using them. That is, a transformer is a neuro-symbolic model, the tokenizer in it is symbolic, and the sampler in it is also symbolic. What's the problem? A regular, vanilla transformer is not universal. It's a feed-forward network. It has a finite number of layers, as you understand. And this implies the following: that for generating the next token, we spend a fixed amount of computation. That's it, in flops, да? What does this, in turn, imply? Well, that there are some tasks for which, in order to find an answer, you need to spend more computation. Let's ask our language model to solve a task with a number of cities. Well, a larger, larger number, да? And give the answer in one token, да? It's obvious that this won't work. It won't work. Why? Because when a certain size of this task is reached, the amount of computation performed by our transformer will be less than the required amount to solve this, say, exponential task. Well, I took such an extreme example on purpose, but we must understand that many purely applied tasks can also be, well, so to speak, computationally limited, да, in a sense for this transformer, for generating one token. And there is no Turing completeness, of course, here, да, there is no recurrence. No recurrence, no Turing completeness. Well, that is, and how to organize a cycle? No way. The model is limited. What to do? Well, you can add recurrence, да? That's it, literally half a year, I think, after the vanilla transformer, да, a paper was published on the universal transformer, да, let's make the central part a certain number of attention blocks.
through which activations we will drive cycles, yes, and this, in principle, solves the problem. There. But in reality, no, because we don't know how to train recurrent networks well. There. And therefore, the idea arose: "Let's also come up with some kind of crutch." And what will the crutch consist of? We will ask the model to solve the problem in parts, yes? That is, if the model can solve the problem of dividing the task into subtasks, and each of these subtasks will turn out to be within the required computational limit. And thus, voilà. Good. In our pre-training corpora, we have some examples of solving problems step-by-step, yes, I don't know, like school textbooks, in which some example is solved step-by-step, yes, or a logic textbook, in which some. As soon as we start to deviate from something that resembles chains of reasoning for typical problems, the performance of such models, of course, drops. And what do they try to do? They try to somehow boost this chain of thought by means of all sorts of synthetics, by creating specialized sets, yes, and then cross our fingers and hope that some knowledge transfer will occur, yes? If we learn to build chains of reasoning in mathematics and logic, then we will be able to build the same chains of reasoning elsewhere, in other areas, in other symbolic systems, and so on. It works to some extent, yes, but not as well as we would like. And the problem is that here there is no, well, like a full-fledged opportunity to fully fine-tune it, because there is a non-differentiable part, yes? That is, the sampler is non-differentiable. Everyone understands this problem, they try to solve it in different ways, but so far, well, there is no good solution. But from the consumer's point of view, even what exists, what we can do in some arithmetic, logical problems to generate some chains of reasoning, to apply basic logic to analyze some situation, it is very spectacular. Well, that is, like an organizer, look, it is also important that it generates a chain in a human-understandable language, yes, which is the solution to some problem. And in this sense, it is like explainable AI for the poor. But in practice, it doesn't add much to the metrics, objectively speaking, well, well, look, it's one percentage point on the main benchmarks. Well, this is, in principle, again, a convenient way to do what is now called test compute, yes, to spend compute after training the model for improvement, yes, there is, in principle, a broader approach like scaffolding. And within scaffolding, building chains, trees, graphs, whatever, reasoning is an important thing. And soon, I think that Next Bing will also add visual modality. Soon our models will not only build reasoning, but also draw on a board, yes, something, some visual representations in the process of reasoning, they will draw the architecture of some solution. Well, yes, yes, well, it exists, it's called Board of Thought, I think. There. In short, it looks cool. From an applied point of view, it's a way to spend a lot more compute to slightly improve metrics. There. In some areas, it gives a good boost. Well, like, if you need to do symbolic reasoning somewhere, yes, there, well, yes, it's really useful and good. In general, well, I wouldn't overestimate it. I think it's still largely a crutch. Until a good way to train this thing appears. That's why everyone is chasing all sorts of Q-stars. Let's generate trajectories for tasks where we know the solution, then use these trajectories for some, I don't know, contrastive learning. For now, we are at a stage of active, extensive technological improvement. And many ideas work. You just said that, well, and there are again, yes, such definitions, winter and spring in artificial intelligence. You once said that now is the summer of artificial intelligence. And it is precisely this summer not only in terms of investment, yes, but also in terms of technology and people's enthusiasm for research in this area. The concept of winter and spring has always amused me. But, well, firstly, what kind of vicious annual cycle do you have, where there is winter, then spring, then winter again? Have you ever seen anything like that? This doesn't happen. There. And secondly, well, it's obviously far-fetched. That is, when you start working with historical material and try to understand, well, there is a simplified picture that they will always tell you in bad popular science. It's like this: Frank Rosenblatt made his perceptron, and then he started getting a lot of money from the government. And then Minsky and Papert came, and they wrote the book Perceptrons and where they destroyed, you know, all these perceptrons. And then Rosenblatt's money was taken away. And Rosenblatt drowned himself out of grief. And the winter of artificial intelligence began. Nothing was done in the seventies. Yes, this is nonsense from beginning to end. And every statement does not correspond to reality, because the reduction in research budgets for Rosenblatt occurred before the release of the book Perceptrons, it is not related to it in any way. And it happened for one simple reason: after 1957, when, you know, the Soviet Union launched the first artificial satellite of the Earth, American scientists, of course, were showered with gold, in a good sense. There. But, at the end of the sixties, a budget amendment was adopted that limited the military's ability to spend on promising research without a specific applied result on a known horizon. There. And this was the reason why Rosenblatt's funding was reduced through the military channels. But the reduction in this funding, it did not seriously affect his research in any way, because he continued his work at the university on the Perceptron, his last project was Tobemore, this system. At the same time, neural network research, of course, was conducted in many places. In particular, well, under the supervision of Rosenblatt himself at SRI, research was conducted on the applied application of perceptrons, yes, and the first OR systems were created then. There. And in general, the Minus Minus2 system, which was developed at SRI for recognizing tanks from aerial photographs, for recognizing symbols on maps, all this was done, it continued to be done even after Rosenblatt's death. No one abandoned this pursuit. Well, and in the seventies, a lot happened in the field of neural networks. And perhaps there was less media attention, but Koniikhiv's Kushima with his neocognitron is the seventies. SRI is the seventies, Steinbock is the seventies, Widrow returned to neural network research in the seventies, yes? So, one of the very important pioneers of the neural network direction. There. In the Soviet Union, of course, research in this field did not stop either. And therefore, to say that the seventies were some kind of winter, well, like, why, yes? Objectively, no. Listen, and if we talk about the development of this topic with artificial intelligence in general, is it more the role of some personality in history, that is, are individual geniuses driving history, are technologies driving history, is it the quantity that turns into quality? Well, you're like a textbook on dialectics, you know, objective factors drive, but these objective factors are still carried out by specific people. Therefore, when points of bifurcation arise, the subjective factor comes to the forefront, yes, that is, when objective conditions have matured, well, like superheated water, yes, boiling begins from one center, yes. That's why, well, it turns out that the subjective factor is very important at any given moment, yes, but for this subjective factor to become important, there must be objective prerequisites. It's incredibly interesting to listen to you, I think. Will the book be released in audio format? Will you not voice it? Well, no, it doesn't even need to be voiced, there is artificial intelligence, just synthesize a voice and there you go. This very well shows the real state of affairs in artificial intelligence technologies, because the distance between publishing a cool paper in a cool journal and showing it at a cool conference and the moment we get an industrially working technology is very large, that is, well, like the Pareto principle is sometimes called a joke that the first 80% of progress takes 20% of the time. What is the problem with automatic audiobook production? It would seem that neural networks, which, well, I don't know, I was given a neural network that speaks in my voice as a birthday gift back in 2019, yes, well, and we were actively developing our stocktronics back in 2018. And the synthesis itself is not neural, it has been working in production since 2019-2020, everyone has more or less switched to neural network speech synthesis technologies, but what will you encounter when trying to stuff such a book, the size of your world, into such an algorithm, into such a system? You will have very strong intonation jumps at the junctions of sentences. You will have many pronunciation errors, because no one knows how to correctly place stress on a surname. It's like, fork and plate are written without a soft sign, yes? But "asol fasol" with a soft sign, it's impossible to understand, you can only memorize it. Yes. And you will, of course, have so many defects in this automatic synthesis that it will be easier for you to re-voice it manually. Well, re-voicing fragments, well, like, different fragments will be voiced with different voices. There will be a lot of post-processing, which, well, like, alas, we want to honestly create a technology that will make audiobooks. For this, it must be said that systems of synthesis that do not synthesize by sentences, yes, but synthesize, well, like, the entire monolith left to right are now actively being studied. There. Well, this is also fastpitch, I think, there was already something like that. There. But it's interesting that in production, many places still use Tacotron successors, especially where it's important to do it quickly. And in business tasks, it's also very important for you to be able to synthesize in real-time. Well, we started talking about voice technologies, yes, about speech technologies. How do you assess the public's trust in artificial intelligence now? Because there is, well, some statistics that people don't like to communicate with robots one way or another. That is, they don't answer calls and want to hear a live person. And if we look at artificial intelligence in general, what is the current level of trust? Well, it's, you see, like, what is the level of trust in mathematics, what is the level of trust in matrix multiplication, yes? Well, we understand that artificial intelligence technology is a very wide spectrum of various things, yes? You take a photo with your phone, and the processing of this image is also done by a neural network that works in your phone. In general, it is mainly thanks to this that progress in photo quality occurs. The optics are not changing at all. There, well, do people trust these technologies? Well, they trust them. They don't even know they are there, yes. But everyone prefers to take a photo, of course, with a smartphone, yes, rather than with a Zenit camera, yes. Well, there are people who like the latter, but there are few of them, yes. So they trust, yes, 100%, yes? And the fact that the neural network will draw pupils for some monument, well, that's okay. Minor life, let's not pay attention to it. There, well, other models, yes, that we use. Well, a search engine, yes, any search engine is also powered by neural networks, yes, modern ones. That is, it's something that allows you to find relevant documents from the entire giant volume of the internet for your query. Do people trust it? Well, they trust it, yes. But there are some applications, artificial intelligence technologies, that annoy or scare people. There. And there, of course, the attitude is different. That is, some artists hate text-to-image models. Yes. Yes. I liked a comment on your Telegram channel. By the way, Sergey Markov has his own Telegram channel, so you can subscribe to it. There are many interesting things there. I said, you can, we strongly recommend. We strongly recommend subscribing to it. There are many interesting memes. And I once saw a funny meme where we invested millions of dollars to find out how a hamburger, I think, eats fries or something like that. Well, yes. Yes. And there was a funny meme. Well, a funny picture generated by artificial intelligence. Well, yes. Here, you see, some technologies still require time for society to get used to them and perceive them normally. For example, in the 19th century, artists were very afraid of photography, thinking that photography would destroy painting. Really, they beat up photographers. There. What do we see? The world turned out to be much more complex, yes, in fact, in the end. Well, firstly, it turns out that creating visual content is not a zero-sum game, yes? A modern person usually has hundreds of photographic portraits, I don't know, in their phone, there are selfies. There. And in the 19th century, people who had 100 of their own portraits mainly lived in mental hospitals. Yes. There. So, like, because a technology appeared that allows you to create a photorealistic image of yourself, and in general reality, it turns out that you can simply consume more visual content, yes? And now that the creation of some types of this content has become much easier, it turns out that, well, people have started to consume more of this content. Secondly, photo art emerged, yes? It turns out that you can not just take a photo, but you can do something with that photo later, yes? That is, you can specially create, not only on canvas, but also by correctly arranging objects or finding the right moment to press the shutter button, yes, and then make a photo collage and something else, yes, then digital appeared. It's a whole song and fairy tale. Yes. So, a whole new direction appeared that didn't exist before this technology. There are more artists. It turns out that technological progress leads to the fact that fewer people are needed to produce essential products. Now people's free time can be freed up, and they will paint pictures. Well, yes, not because they need to survive, yes, but as a hobby, yes. So, and in principle, there are even more professional artists now than there were in the 19th century. There. And this also happened thanks to technological progress. Well, indirectly, not of these technologies, but others, yes. But the world is not stationary. This is also important. Then it turns out that there are some niches from which, well, like, no technology will ever displace artists, yes? Why? Because I want to have a painting on my wall that was painted by a real human. I can afford it. You will hang photo posters for yourself, and I will hire someone who will get dirty with oil and paint my portrait. Yes. And a portrait from a photograph, I'll note separately. Yes, this is also how the original differs from the copy, yes, philosophical questions. Let's recall Walter Benjamin with his art in the age of mechanical reproduction, and so on. There. Well, so, I propose another term: organic art. That is, there are people who eat organic food, yes, that we don't want an egg from a factory chicken, we want it from a happy chicken that walks on grass, under the sun, yes. There. And I want a real human to paint it for me, yes, this organic art is a status symbol, that you have a painting made by a human, not a machine. There. So, this also exists and will exist. And how many times has labor productivity increased? Well, a hundred times, probably, yes. So, like, now for the same unit of time, a developer can create 100 times more functionality than they could in the late forties. In the late forties, there were about 100 programmers in the world. And if it were a zero-sum game, then what kind of world would we observe now? Only one programmer would remain, yes, who would perform the work that 100 programmers did then, yes, the rest would be fired. The economic effect of this automation would be equal to the salary of 99 people. Yes, it would pay for itself. It would say: "What are you doing? Why are you automating your programming? Your high-level languages. Why all this? How much will you earn? You will fire 99 people." There. Well, this is it. And there are such people now, yes, who say: "Why are you developing artificial intelligence technologies? How will you recoup it?" Well, you will fire people from call centers, well, okay, how much do they earn? Where is the money? Where is the economic effect? And the economic effect arises where, naturally, entire new industries arise, entire new product areas simply appear. There. That's why, of course, the effect of automation is not reduced to replacing a woman with a machine. Yes, it's all, well, sometimes good, but the main economic effect is not achieved there. So, I'm getting to the point that, well, people see that, well, that is, the same programmers, yes, well, they started using all sorts of copylots in their work. This does not lead to the replacement of programmers, yes. And it does not lead to a decrease in their number, because it simply increases, if you can now produce even more functionality per unit of time, it means that the scope of application of software engineering is expanding. This means that you can now create, for example, more complex systems than before, when you have assistants who identify bugs for you and write autotests for you, yes? This means that in some areas where automation was previously expensive, it can now be done, yes, this means a lowering of the entry threshold, yes, some kind of programmer fast food, to make a simple thing, elementary, now it's possible, it turns out, even without human involvement, yes. There. But, this leads to a large expansion of the sphere of application of these technologies. And the total number of people employed in this area is likely to increase, not decrease. That is, there will be more programmers, and of course, their work will be different, not like a programmer's today. Yes, and from the point of view, I don't know, of a programmer from the forties or fifties, modern programmers, well, well, what kind of programmers are they, they are some kind of managers, I don't know, they don't even know how to add on their processor, which they have. There. Well, and so on, yes? So, well, like, well, yes, we used to think in terms of servers, yes, to talk about physical technologies, and now we think in terms of virtual machines, and few people go to the servers anymore. And we even talked with colleagues from Zarya. This is in the context that artificial intelligence is more of a human enhancement rather than a replacement to some extent in some, well, in most professions, that to draw a good picture with the help of AI, you at least need to know the basics. Well, I need to appreciate you, for example, well, just, you know, the existence of technologies for generating pictures from text sets, well, like a basic level, level zero, yes, well, like what is available to everyone. Well, and how will this be valued? Well, well, not at all. That is, how was it valued before, I don't know, some scribbles that anyone can do, yes? There. And still, even the public's attitude towards this will be such that, well, like, you just decorated it with generated pictures, without even bothering to choose the best one from a dozen generations. Well, now I often encounter that when I start talking about the fact that, look, we have entire research teams, these teams are training Foundation Models, this is so important, they say: "This is some kind of crap, what are you doing? Why are you even thinking about this? There's Llama, there's PyTorch, open source is power. Why waste all this money and resources on training Foundation Models?" And you had such a heated discussion in your channel recently. I read it with great interest. And here, you know, your point of view, why is this necessary at all? Why, I don't know, us, you, the country, the world, different basic models? Why are different countries investing an incredible amount of money, resources, energy into this? Well, yes, it's important to also say that Sber actually remains the only single company, the only team left to get Foundation. If you refuse to develop certain technologies, there are certain risks associated with this. Strategic risks, and they are also very large, because, well, modern artificial intelligence technologies, they have a very powerful disruptive potential. They can transform entire industries, yes, create. Well, all this, well, like, the monetary effect of all this is potentially very large. And when you are completely dependent on a supplier of some, you know, fundamentally important components, these are risks that cannot be quickly compensated for when they are realized, yes? So, if you are on the track of, well, we take someone else's weights, yes? Well, okay, but the next version of this model, it will not be published with weights, yes? But Open AI, for how many years have they not published the weights of their models, yes? But why do you think that in conditions of intensifying competition, everyone will continue to publish weights for top models? Yes. And why did you think that at all? And you will essentially lose the competence in the teams to create such a model from scratch. Well, and in reality, it's not as harmless as it seems. That is, when they say: "Let's take Llama now and fine-tune it as if we trained it from scratch." In reality, well, the problem of fine-tuning a model that has already seen 15 trillion tokens. Well, you know, there's catastrophic forgetting. There's such a thing. Well, we added some Russian test set to it. Yes, of course, as a result, the models start to speak Russian more competently, yes, they improve in generative tasks. This is true, but at what cost? Well, like, your model forgets something from what it learned in pre-training. And this, well, in general, is a capitulation, that is, like, we don't know how to make a model like Llama, yes, so we will take a ready-made one and fine-tune it. Well, like, you admitted defeat, that you don't know how. And at the same time, we see that the trend in the world is the opposite. That is, Chinese teams, which, well, like, yes, 3-4 years ago, yes, everyone laughed at Chinese models, saying: "These are some bad works that are not reproducible." Reproducibility is also a problem everywhere in the world. Well, try to reproduce the work of not the top teams, yes, and even top ones sometimes. Well, everything is bad with reproducibility, but everyone laughed, and what do we see? Yes, we see that the team, well, DeepMind, which, in general, is in the same weight category as Russian teams. Well, of course, there, look, they have 300 people listed as co-authors, well, do we not have 300 ML engineers in Yandex or Sber, yes, we have more, of course. There. Well, they are unique guys, they are like quantum traders, they, well, they did low-level optimization. Well, do we not have specialists in low-level optimization? Well, we do, I personally know many, yes, and we have some startups optimizing at a low level, everything is fine too, yes. And, well, like, when we were significantly inferior to anyone in terms of algorithms, we have consistently taken top places in all Olympiads for the last few decades, yes? So, well, it's a shame to complain, so what are you lacking? Well, like, what will be the price for the loss of technological sovereignty in this area? It will potentially be high. There. Because tomorrow, when you really need to not fall apart, you will need to create a model, I don't know, that will lag behind, at least, not more than
for a year there, yes, from leaders, yes, from this, for example, well, like gradually these models, they are becoming part of a huge number of business processes. And overall - well, there are tasks in which, for example, Gigachat for the Russian language is better than top models. Uh, there may not be a super, uh, gigantic number of such tasks, but they exist, yes, that is, we know, from our internal Sberbank applications, yes, >> again, we have people who know how to prepare this model correctly, know how to solve various business problems based on it. >> Uh, >> tomorrow, if we enter an era of complete regionalization and closure, yes, when the Chinese stop publishing their models, and Western companies, what will we do? And also under some sanctions. Well, it will be a sad story, believe me. Here. And it is important to conduct some research in what sense? So, uh, sometimes a narrative also appears that, well, you are scattering your efforts, you are doing, I don't know, some kind of peas, yes, you have people doing this, and that, and photographing bicycles, rowing >> and hunting. So, yes, but, uh, thanks to the fact that you are at least a little bit involved in this topic, if it becomes mainstream tomorrow, you are not starting scaling from scratch, right? You have a core team, it has some data, it has a set up understanding, and you can, the scaling speed of this team will be much greater than if you have no one at all. That is, if you have two people who have been tinkering with this topic for 5 years, >> Uh-huh. >> you can turn these two people into a team of 100-200 relatively quickly. If you have zero, then, well, like >> finding those two key people, on the basis of which this can be built, it seems to me, is a separate thing. >> Yes. Yes. Here. Therefore, that modern technologies are expensive, yes, modern technologies are expensive, but the factor of scientific and technological progress is currently key for long-term business sustainability, because any technological achievements that already exist >> Uh-huh. >> they are devalued at an enormous speed. That is, the speed of technological inflation is enormous. Therefore, uh, we cannot afford not to play the long game. Well, like, it's too risky a strategy not to play the long game, but to quickly cash in, yes, to earn on current technologies, to abandon all promising research and so on. Uh, well, this strategy will very quickly lead to collapse. Here. And many things that we have done, which are now part of the core business, arose as absolutely complete disruptors, which it was unclear why to do, who needed them, and so on. Well, a lot of our various electronic online services and, I don't know, models that underlie decision-making in core banking business processes. >> Now, by the way, LLMs are even starting to be actively used in scoring systems, I've heard. Well, yes, they are just used quite specifically. It is important to understand that LLMs, well, that is, in tasks of small dimensionality with a small number of factors, classical ML methods work better, yes, boosting, random forests, and so on. But, uh, when you need to analyze unstructured data, you have no other tool except LLMs. And you have a lot of data about counterparties, about transactions, about macroeconomic factors, it does not exist in a structured form. You need to create factors from a lot of non-specific, unstructured data, which you can then feed into, well, some model based on boosting. So, in general, everyone is actively using hybrid models, when we extract some factors, analyzing tons of various documents, I don't know, news, something else, yes, and then feed them into a Russian-specialized model. Here. But I want to say that now even in the analysis of banking transactions, in the analysis of various specific financial information, a big breakthrough has also occurred thanks to fundamental models, because a fundamental model is good because it >> knows how to transfer knowledge between domains, right. And so, yes, your specific scoring model, trained on your specific credit product, seems to have high accuracy, but its problem is that the data it learns from is limited, right? And the data, well, like, has some external factors, right, that are difficult to account for. And it turns out that if you take, well, much more data from the outside, not related to your task, but from some adjacent areas and so on, then a large model, having learned from these huge volumes of information, and then being further trained on your specific task, demonstrates much better performance than a narrow model, which learns only from your data from this area. That is, it finds some analogies in other processes, in other tasks, and successfully uses these analogies to solve your task. And this is especially helpful in situations where there is very little data >> on your target task and when you need to have stable results out of sample, out of time. Here. Uh, here, of course, fundamental models are much better than specialized ones. And I see now how gradually, even in tasks that seemed to have always been a stronghold of classical ML, the influence of fundamental models is gradually growing. >> That is, we have already found business applications. That is, even if it is hidden from ordinary users, consumers, businesses are already actively using it. >> Well, of course, no. Well, the direct effects from the application of modern artificial intelligence technologies, even within Sberbank's core business, are huge. Well, it's clear that we can't tell all the cases externally, but they are really very big, but imagine how much information our bank works with, how many text documents, anything, communication with people, transactions, anything. And in order to extract useful, necessary information from this sea, useful support for human intelligence. >> Well, we believe in Gigachat, we actively tell colleagues about your successes. And we will gradually finish. Then I have two questions. A blitz, a small one. A small blitz. So, >> briefly. The most insane myth about artificial intelligence that you have heard? >> An overwhelming degree of madness. I remember one professor, a well-known neurosurgeon, said that artificial intelligence cannot be created because, you know, connections in the brain constantly change. That is, the guy wasn't taught in an electrical engineering course that there are switches, you know, like you can >> press a button and the light will turn on. Nothing is being re-soldered there. >> Funny. And one more question, the last one. Where would you not see the application of artificial intelligence at all, or would not want to? >> Yes, I am a radical in this sense. I everywhere, >> well, like when someone comes. >> No, I, you know, I probably have more of a position like this. >> Uh, well, not for bad purposes, right, apply it, but for good ones, yes, right, well, that is, any technology, as universal as an artificial intelligence system, is a double-edged sword, right, you can apply it for good and for evil, and it all depends on people. >> Because even with a simple hammer, you can hammer nails, build houses, or break someone's head. Yes, and therefore, uh, well, like, we don't want heads to be broken with a hammer, we want houses to be built. >> That's why any applications, I don't know, for fraudulent purposes, applications to make the poor even poorer, the rich even richer, right, which, in principle, worsen people's quality of life. Well, of course, we don't want to see such applications. Here. But technology, unfortunately or fortunately, is neutral, right? That is, and well, you can also apply these technologies for evil, right. But the good news is that these same technologies can be applied and, >> yes, that is, if you, I don't know, >> or to counter bad people, right? >> Well, yes, yes. Well, that is, you can generate fakes, you can detect fakes, right, you can, uh, you know, deceive people using generative models, or you can use the same models to identify such deception, right, and protect people from fraud. And such systems are actively being created. And, for example, there are many applications in cybersecurity with Sberbank's artificial intelligence technologies. That is, if there were no such advanced artificial intelligence technologies in Sberbank, I think that given the volume of attacks of all kinds and social engineering with the involvement of technological means, which our bank is subjected to every day, it would not have withstood it. >> So, in general, it's more a question of what we consider right and wrong. It depends on us how these technologies will be applied. And in this sense, the paradox is that in an era of such rapid technological progress, human qualities turn out to be most in demand. Yes, and how we will use these technologies depends not on the technologies, but on our commitment to humanism and human ideals. >> We are finishing. Sergey, thank you very much. Thank you for inviting me, for this wonderful podcast. >> Thank you for coming, it's a pleasure to talk to you. Of course, it's a pleasure. Thank you. It's just a great pleasure. >> You're too kind. >> Well, no. >> impossible. I've known you for a long time. >> We'll just summarize. And I've prepared a small speech here. We talked today about people who are skeptical. They have a low level of trust in artificial intelligence. And remember, if you search for something on a search engine every day or watch movies on Okko, then one way or another you are using artificial intelligence, even without realizing it, because artificial intelligence is a broad field of science and technology. Thank you all, and subscribe to Sergey Markov's channel, subscribe to my channel, and subscribe to Alena Drobashavskaya's channel. We will provide all the links. >> Bye everyone. >> Bye everyone, guys. Thank you. [music]