📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Евгений Глотов — Spark — ВСЁ!

SmartData45:02

Transcription

Well, friends, I greet you all. I have two questions and one reminder. First question. Who uses Spark? Many people. And who is tired of Spark? [laughter] Well, a little less, but there are those who are tired. And now Zhenya will tell us what to do for those who are tired of Spark, and when this bright future will come, that we will be able to kill Spark. I don't know how it will happen or not. And now, Zhenya. >> Oleg, thank you. Yes, my name is Evgeny Glotov. I work at Naviaо. For those who don't know yet what we do, we create autonomous transport, both cargo and passenger, for movement in this mode, yes, that is, without a driver behind the wheel on public roads. HD maps are used for this. I am responsible for creating HD maps in the company. Now let's get to the point. Today I will tell you about Spark, yes, about where it started, how it is, in general, ending now, and about, well, actually, about Spark substitutes, identical to natural ones. Let's go. Let's start with history. Before Spark, to begin with, yes, in the late seventies, when mammoths still roamed the earth, yes, accordingly, data was processed using relational DBMS. And the problem is that when the internet appeared, such volumes of data appeared that relational DBMS could no longer cope. Therefore, in the early 2000s, Hadoop appeared. And, accordingly, all large companies, yes, mainly BigTech, started working with it and began to encounter the first problems. The problem, you see, is as follows. It was very difficult to develop on Hadoop, yes, meaning the API is very complex, yes, because you only have two, well, essentially, such large operations, yes, MapReduce, there is also shuffle and everything else. This also needs to be configured. This is also a pain. This creates a huge percentage of boilerplate, yes? That is, to write some logic, you have to write a huge piece of something incomprehensible. And, the most serious problem, yes, is that between each stage, that is, after each reduce, before doing the next map, you have to dump data to disk. And not just to disk, yes, but to the file system, to a distributed file system. That is, this is even slower than just to a local disk. Accordingly, Spark appeared as, in general, Matei Zaharia's diploma work. as an attempt, yes, to solve these very problems. Firstly, it provided a convenient API, like RDD, distributed dataset. It solved several problems at once. Firstly, it's a high-level API, it's much more convenient. Secondly, between stages, data was not dumped to a distributed file system, but directly into RAM. And from there, the next stages were read. And, accordingly, you could stack parts of the query together, yes, meaning the API became much more convenient. Well, for example, yes, a banal, yes, standard Word Count, what it was on Hadoop and what it became on Spark. As you can see, yes, 66% of the code on Hadoop is boilerplate. Of course, Spark helped us a lot with this. What's next? Yes, accordingly, Spark began to develop rapidly, precisely because it is very convenient. And, accordingly, from the first version, for those who may not be aware, yes, this in-memory storage is successfully dumped to disk if you don't have enough RAM, because before, queries simply failed, but now Shuffle Spill is mainly used. And, accordingly, besides this, yes, besides the fact that it is simply reliable and actually works, it has a very large library and SQL functions, there is streaming, and much more, yes. Well, for example, we actively use Spark's Python UDFs, including vectorized ones, yes. Well, and accordingly, in principle, you can redefine almost any part of Spark. There will be reports on this tomorrow too. Now I will tell you about features that, in fact, few people talk about. Well, meaning, I mainly talk about UDFs. Few people are probably interested. But about the Python Data Source API. In fact, this is new functionality, it appeared in Spark 4. About Connect a little later. So, let's start with Apache Arrow and how it helps Spark speed up some complex queries that cannot be described by ordinary SQL. What is Apache Arrow? It is a framework for processing, storing in RAM, and transferring columnar data between processes in an optimal format. I have already talked about it in sufficient detail in previous reports. You can familiarize yourself with it via the link. Its feature, in fact, is that it is implemented in all major programming languages, and it is implemented in the same way, meaning it's an element of standardization. Which is quite cool in general. Spark has been using Pandas UDFs based on Apache Arrow for a long time, since version 2, but it, roughly speaking, used them incorrectly before. How? That is, before, everything worked through the Pandas API, including UDFs at some point, and they also worked through the Pandas API, and Pandas, as you know, has problems with nulls, it has problems with, well, missing values, yes, it has problems with complex structured types, yes, meaning if you want an array of structures, arrays of structures, then, unfortunately, it's not available. And you had to dig into JSON, naturally. Well, and accordingly, in the documentation, they maintained backward compatibility and crammed it into the documentation anyway, a layer on Pandas for some reason, it's unclear why. Accordingly, with version 4, you can no longer do this, yes? You can write code in pure Arrow and not use Pandas at all. Accordingly, all the overhead of Pandas, yes, and everything else, simply disappears. Now, firstly, the Pandas UDFs themselves work on top of it. Secondly, you can simply use Arrow compute functions or any other frameworks that can be adapted here. Let's move on. Data Source API, yes, in fact, well, many people have probably developed their own data source APIs in Scala. Well, the thing is, Scala needs to be compiled. In Scala, you need, well, simply put, to find developers first. Because it's much easier to find developers in Python. And it turns out that in fact, you can write your own data source API in literally, well, in half an hour in Python, not compile it, but simply run it in Jupyter, see how it works, debug it quickly. Yes, and in fact, you lose nothing here, because, thanks to Apache Arrow, all this is vectorized and executed, well, either in C++ or in whatever else it executes. And it turns out that you win both in terms of code writing, yes, meaning the speed of code writing is higher, of course, in Python. And also in terms of execution, because everything is executed on the native engine. Let's move on. What about convenience in general, yes? Why is Spark good? Why do we all use it? Well, at least many of us. Firstly, it has a really convenient API. Why is it convenient? Yes, because it is also, in fact, very similar to SQL, and the best SQL, as you know, has not invented anything new in 50 years. DataFrames are the same SQL, only in an imperative paradigm. That's why it's convenient. Also, Spark handles simply any volumes of data. Well, petabytes are no problem at all. It allows you to read and write any data formats from any systems. If you lack something in Spark's standard library, you simply take and write your own, for example, a data source or just, well, some code that reads and writes and does some integration. Besides this, it naturally has, well, over 400 functions, meaning around 450 already, in version 4. Well, and accordingly, if you lack something, you can always extend it, including in Scala or Python. If it's so good, why am I actually telling you about this now, how to abandon it? In fact, it has problems, it is very slow. Well, and let's start with the fact that just starting a session takes about 20 seconds. Well, maybe for some it's faster, for some it's slower, yes, it's quite normal. There are too many primitives for distributed computing, yes, and naturally, well, the JVM also needs some time to start up. And as a result, what happens? If you process some small data arrays, and in fact, in Big Data, there is a lot of work with small data arrays. That is, the total volume can be quite large, but a specific job is small. And it turns out that you simply lose on starting the session here. And in fact, well, this is, of course, my personal opinion, yes, but I believe that at some point the development, well, of Scala, and Java, and all libraries based on them, has stopped a little, and at some point everything started to be done in Python, the State of the Art, yes. Well, at least for about 10 years now, for sure. Well, and accordingly, the poll that Oleg conducted showed that 90% of users do everything in Python. Well, and accordingly, what else can be said, yes, Spark queries that we write in Spark, yes, they are actually tied to RAM. What does this mean? This means that you don't have enough RAM to speed up your job. That is, you would add more Spark Executor Cores, but you can't because each core consumes a certain amount of RAM. And, accordingly, RAM, as you know, is not a divisible resource, yes, it cannot be distributed in pieces, otherwise, you would have to, again, dump RAM to disk. This creates even more problems and is even slower. Besides this, yes, for example, if you move to Kubernetes, you encounter the fact that, simply put, requests and limits cannot be set separately, they are set only together. It turns out that if you don't increase the limit, you will be killed by OOM. And if you increase the limit, then you, it turns out, have taken a large amount of cluster resource monopolistically and are not using it. As a result, naturally, RAM consumption grows simply, well, colossally. And it turns out that, well, as it were, the optimization potential, yes, here it is huge, and you want to use it somehow. Now let's move on to what, in fact, different data engineers, yes, in different companies, are trying to replace Spark with. Let's start with small data, yes, what is, well, let's say, my personal definition of Big Data, yes? Big Data is what crashes Excel. And small data is a subset of Big Data that fits into computations on a single node. That is, it can be large data that lies on S3, but at the same time, one node is enough for you to process it. And for this, you can use frameworks like DagDB and Polars. Let's start with Polars. Polars is a bit more popular, yes, in fact, it has 35,000 stars on GitHub. And we, in our company, also actively use it. There will be a report on this today at 17:30. Come, support the guys too. Well, accordingly, I'll briefly tell you about it, yes, it has its own implementation under the hood, Arrow protocol, yes, meaning it's not Apache's, unfortunately. And also, unfortunately, or maybe fortunately, it's a bit deprecated. And, well, in fact, a standardization process has begun. The guys from Arrow2, yes, they, well, once had a disagreement, yes, and now they are returning to Apache Arrow RS, which, in principle, cannot but please. But if we talk about Polars, yes, firstly, it is really fast precisely because of the LazyFrame API, meaning it has a quite powerful optimizing engine. But at the same time, you can also use the DataFrame API, that is, for migrating from Pandas, it's more convenient. For some analytics, yes, that you do in notebooks, it turns out to be quite convenient. It has a SQL API, yes, since the first version, but unfortunately, it works incorrectly. I, well, also talked about this in the previous report. But the point is, for example, window functions simply work incorrectly there. And more than that, they don't just fail, yes, but they work incorrectly. This must be taken into account when working with it. Despite this, it is used, yes, because, firstly, it is very fast, and secondly, it is really flexibly configurable, yes, meaning you can write your own Rust extension, and this is very convenient. My colleagues will also have a report on this. And, accordingly, well, I personally use the Polars extension for geospatial data. That is, it works fast and well. Let's move on. DAGD DB, yes, this is essentially a different approach. It's pure SQL. That is, essentially, it's a mini MapReduce. Well, it works even faster than Polars, because, well, firstly, it's in C++, they are in Rust, yes, and secondly, the guys are purely working on the SQL query optimizing compiler. And it works very well. Well, and accordingly, it's also a quite popular framework. It has, I think, 32,000 stars on GitHub, meaning it's close to Polars. But unfortunately, there is no DataFrame API. And if you want to get it, yes, then you are essentially doing it and moving forward, as it were, to use something else. Pandas, Polars, or other frameworks. Well, meaning, if you are starting your small data or big data, yes, if you are starting from scratch, if your data doesn't fit into Excel, if it fits, then you take Excel and enjoy life, or you probably won't be invited there, because there are already analysts who have been doing this for 10 years. And accordingly, if you have previously developed in Pandas, then, as Oleg already said, you can quite easily switch to DagDB or Polars. What are the disadvantages here? Firstly, if you already have code in Spark, then you will have to rewrite it, because the API is incompatible. Secondly, there are things that are simply not implemented, for example, working with the Hive Metastore. And it's not certain that they will be implemented, because, well, they have already moved forward somewhere, yes. But in fact, you can use them with Iceberg, they work more or less. The question is whether they can write a lot of data. Well, again, my colleagues will tell you about this in their report. Well, and accordingly, if you stop fitting into one node, then it's unclear what to do. Either build some of your own workarounds with distributed processing, or, accordingly, you will have to rewrite everything back to Spark and lose performance here. That is, it seems like you optimized, optimized, spent time, yes, and accordingly, you will have to go back. What are the options, yes? You can take big data, yes, take Spark. Spark, as you know, has Catalyst, it does query plans very well. And you can take this plan and execute it on a native engine, that is, in C++ or in Rust. And let's move on to such solutions. Well, let's say, a forbidden organization, yes, made a plugin for Spark called Gluten. Well, accordingly, they made a fork, yes. Gluten, I think, is an Apache thing. Well, and accordingly, well, ClickHouse, as you know, is also an American startup. Well, and well, it turns out that we have a plugin that allows you to offload Spark computations to native engines. Here, specifically, everything is in C++, yes, and accordingly, Gluten itself is also in C++. And if some method is not implemented, well, this happens, yes, then, accordingly, all computations are simply reset back to Scala. That is, to Scala Spark, meaning purely classic solution. And, accordingly, the authors achieved some acceleration. Quite good, yes, three times, in general, cool. If you still want an Apache solution, then you can take Comet. It's also a plugin for Spark. It also forwards computations to a native engine. In this case, Apache DataFusion, about which a little later. If some method is not implemented, it also, naturally, resets the entire plan and executes it on Scala. Well, and the acceleration is a little less. Well, in fact, this may be due to the fact that not everything is implemented yet. And now, a little bit about DataFusion. It will be needed later. What is DataFusion? It is a low-level framework for local data processing. It is based on Apache Arrow, meaning it has many primitives from Arrow, such as buffers, as well as ready-made Arrow data structures, types, and many compute functions. They are simply wrapped in a higher-level primitive, yes, and accordingly, it becomes easier to work with. And, accordingly, for engine developers, it provides very powerful functionality that allows you to significantly reduce the complexity of developing your own engine, yes, with your own functions. Well, and accordingly, it is also extensible, which is quite cool. Now let's move on to the comparison, yes, Gluten with Comet. In principle, well, the comrades from Comet launched a benchmark, yes. It's unusual that those who ran the benchmark got worse results. Well, but at least it's honest. Well, and accordingly, as we see, yes, there is some acceleration. Meaning, about 2.2-3 times. This is specifically on TPC-H. At the same time, yes, in fact, there are problems too. That is, well, meaning, not everything is implemented. And you can even simply encounter unimplemented functions or incorrectly implemented ones almost everywhere. Well, and accordingly, there will also be a report on this tomorrow, about how all this works and whether it works or not. But I'll briefly summarize, yes? That is, you can get acceleration, or you might not. Functions can be, well, they can be implemented incorrectly, yes, in the worst case. In the best case, they are simply absent, and the entire plan runs back on Spark. In fact, the developers of, for example, Iceberg, oh, not Iceberg, but Comet, do not quite understand whether their plan runs on Comet or it runs on vanilla Spark. And this, by the way, well, it was clearly visible in the issue that I cited. Well, and accordingly, if you want to compile all this, you, well, start compiling even more. All this is not very convenient. And the question arises: is it possible to abandon this somehow? Yes, is it possible to throw out the entire Spark implementation completely and do it from scratch? In fact, you can use something called Spark Connect. What is it? Spark Connect is a client-server architecture with a lightweight client. Well, here's a bit more detail. What is a client? It's something that forms a raw query plan. That is, it simply takes user code, yes, in SQL or in DataFrame API, and forms a query plan from it, meaning what to read, from where, how to process, and so on. At the same time, it does not optimize it, transmits it via gRPC to the Spark server, yes, to the Spark Connect server, and the server does all the rest of the work, that is, it optimizes this SQL, computes, yes, performs all the necessary procedures, yes, to ensure fault tolerance, everything else. And, accordingly, returns the result to the client using Arrow buffers, that is, in an optimized columnar format. Yes, there will also be a report on this tomorrow. And accordingly, what do we get here, yes? We actually get the opportunity to throw out Spark altogether, not install Java on the computer, well, or in Kubernetes, yes, and install only a library that weighs 1.5 MB. And, accordingly, start using Spark to the fullest, yes, if we have an implemented server. Well, the server can be either Spark's, yes, meaning Apache's, or any other. And now I will tell you about the engines that implement the Spark Connect Server in one way or another. Well, the first will be DataFusion, yes, I also talked about it at the conference this year. Well, this is an interesting framework. It is also, naturally, in Rust with a Python wrapper. It also, like Polars, has Arrow under the hood. But it has a distributed mode, and it doesn't run in Kubernetes, but through Ray Cluster. This is a higher-level abstraction over Kubernetes. Well, as you know, orchestration, yes, but scheduling in Kubernetes is a rather complex task. Ray solves it. Also, unlike Polars, it has window functions. Well, just like Polars, it has UDFs. And a bit of a surprise, yes? That is, there is image processing right out of the box. That is, there are ready-made methods for working with images and a bit of ML. And besides all this, yes, it implements the Spark Connect server. At least it tries. In fact, the attempts are quite complex, yes, because precisely this Arrow2 implementation, the problem is that, firstly, it's deprecated, and secondly, it's too low-level. That is, to write all the code, which, well, already exists in DataFusion, yes, you can just go crazy. Well, and accordingly, yes, if in May I had the question: "Is there anything better?" Yes. Is there something that, with the right architecture, implements a Spark analog, but completely on top, yes, and other Apache components? In fact, such a framework has appeared. It is called Velox. Well, and further I will talk about it in more detail. What is Velox? It is a distributed data processing framework. It is also written in Rust and Python. Its distributed mode works in Kubernetes. Well, will there ever be YARN? I don't know, maybe. But in Kubernetes, in principle, most have probably already moved. Unlike Polars and DataFusion, it has DataFusion under the hood, which allows it to be very lightweight. That is, development in Velox is a rather simple thing. That is, because the majority, yes, precisely the complex things, yes, it simply delegates to DataFusion. It has no Java at all, and most likely never will. Of course, it's unclear how to support Java plugins currently. Surely, you can write an extension that will call an extension in Java, and accordingly, do it somehow like that. But whether this will ever happen or not, is still unknown. But nevertheless, that is, this is perhaps the first engine that truly aims to fully implement the Spark Connect Server API. Even at the current moment, Spark streaming is present in a raw form for now. It requires serious refinement. More on this a little later, but it's already like a checkbox has been ticked. Well, and accordingly, it passes TPC-DS and acceleration, well, at least according to the developers' claims, about four times. Now let's move on to the architecture. We take Spark Connect, which is provided by Apache Spark, yes, it weighs 1.5 MB, remember. We take the Velox server instead of the Spark server, and accordingly, inside it works a driver, and it creates workers, which then perform all the work. So, it looks the same as Spark, only here, there are no row-to-column and column-to-row conversions. Everything here is columnar, yes, and executed on DataFusion. Moreover, the shuffle is also columnar and transmitted via Arrow buffers. Well, here it is presented, yes, how Velox actually uses DataFusion. Essentially, at each stage, a large part of DataFusion's primitives is used. And if it's not enough, then at each stage, Velox makes its own small extension, yes, in order to implement Spark. Because, unfortunately, it still doesn't implement Spark. Well, naturally, distributed computations are also made by Velox developers. Let's talk a little about physical execution, yes. As you know, Arrow is a columnar data processing engine that is super fast, yes, because, well, in C++ or Rust. And, accordingly, vectorized, yes. What are vectorized computations? How do they look? In fact, they look like this. You simply, in fact, in the code, you will most often notice a loop over

a set of elements. Here. But it is executed, naturally, along one column, yes. The data there is of the same type, it is well structured. Here. And actually, if you have a quality compiler, then it can make execution vectors by itself. Here. Well, or you can write a truly optimized function that will perform such code. Here. Well, actually, the main problem, yes, where work with strings, with binary blobs, yes, that is, with some types that are not trivial, yes, that is, it's not ints, to add two numbers, yes, but some, I don't know, even the same substring, is already a non-trivial operation. Here. Well, if you think that this is a joke, then actually no. Here, I personally committed this to sale. Here. Well, and, accordingly, this goes to SD data Fusion. Here. So you can, in principle, encounter such implementations almost everywhere. That is, just a pass through the rows of one column and, accordingly, the execution of some useful code. And now let's talk about execution speed. Yes, that is, uh, what will we measure the speed with, yes? That is, how to compare two or more engines. You can take PCCDS, but, firstly, not everyone passes it, unfortunately, yes? Secondly, those who pass it often sin by focusing on these queries and start tailoring their execution plans to these queries. Here. Which is not very honest. Here. And most importantly, yes, that there is no general leaderboard, on which you can look and see who is actually cooler. And, accordingly, here clickbench comes to our aid. Actually, well, many know that ClickHouse does not lag. Here. But for those who don't know, ClickHouse developers specifically created a benchmark to show that it really doesn't lag. And the benchmark is actually good because, firstly, you can add your own implementation to it. And now there are about fifty implementations, that is, almost all popular engines are there, at least. Here. And if something is missing, then you can always commit it. Here. And accordingly, everything I need for comparison is there, there is an open leaderboard, on which you can go and look, well, compare, yes, two engines or ten engines. And moreover, all this works on one, well, on one specific node, yes, of a specific type. So you can also compare how it works on a garbage lid of Nubuck or how it works on a node with a huge amount of RAM. And, accordingly, one of the advantages, yes, is that essentially these engines, which I will now talk about, have been committed in the last couple of months, and the results of the queries have not been tailored, that is, the engines to make them work faster. Here, which is also convenient. In clickbench itself, yes, there are several ways to run it. Here, I will only talk about reading parquet. Cold run, that is, we read everything from scratch, no in-memory there, nothing like that. And, accordingly, there is a geometric mean metric, that is, you can compare not by the total execution time, but by the ratio of different queries to each other. Here. Well, the metric is like this, you can compare it later. And what can we notice, yes, that, firstly, actually DGDB beats everyone. Here. Secondly, Data Fusion is good approximately the same as, yes, but slightly inferior to DDB. Here, well, I don't know, you can talk about the fact that C++ and all that, but actually it's more a question of implementation. Also, among all the engines that implement parquet, unfortunately, they are all a little slower, but among them, sale turned out to be the fastest. Here. And it is actually at the same level as Polars in terms of speed, yes. So it's something to think about. Here. Now let's move on to why our engines, including SPРK, are lagging and how to fix it. A small question to the audience. Here, data beats everyone on one of the queries, that is, simply by 15 times. Why do you think? First answer option: DDB has cheaters. Raise your hand if you think so? Nobody. Okay. In Dac DB, there is an incorrect implementation, there is a bug and an incorrect result is given. Who is for this option? Someone is there? Yes. Here. And the third option. In Dadi, the most efficient algorithm is implemented. Everyone else is in doubt, yes? Actually, the most efficient algorithm is implemented there. I'll talk about it a little later, and now, actually, I'll tell you about the fact that Data Fusion got a little burned by this and implemented an optimization themselves, but a slightly different one. And actually, about the query itself, yes. If you don't optimize anything, you do a full scan, then you do a full sort of the entire dataset, and the dataset is, well, several gigabytes. Here. And then you take the first 10 rows. Here. Can this be optimized in any way? Of course, yes. The first thing that comes to mind is to make a binary heap for 10 rows. Here, they are actually a full scan, yes, and then instead of sorting, you just select 10 rows. Here, this is, of course, fast, but not very. Here. You can do even better. You can extend this binary heap to reading. Here. That is, while you are reading, you always keep 10 rows in memory for two columns, yes? That is, you read two columns first, understand whether you need to read the entire row or not. If you need to, then, accordingly, you read it. And this is how the update happens, yes, and it always contains 10 rows. Here, of course, it doesn't count 10 rows, but it's still quite fast. That is, Data Fusion developers claim up to a 25-fold performance increase. Here, I don't believe it. And what did the geniuses from DAКDB do? Well, actually, you can consider it cheating, but well, we, at least, discussed it in data job and didn't find where the cheating is. Here, actually, they take the query, yes, read two columns, sort them, and so on. But in addition to these two columns, yes, which are actually needed for the query, they take two system columns. These are the file name and the row number in that file. That is, if we have parquet, and we have, as you remember, parquet, then it turns out that, further, for a query that reads all columns, yes, you need to read 10 files and read one row from each. What could be faster? Well, I personally, well, couldn't think of anything. Here, this is, in my opinion, the most optimal option. This works, essentially, for all columnar formats, yes, and for a number of records, well, thousands, maybe tens of thousands, it will probably work. Here, let's move on, yes, we talked about speed. Speed can be improved infinitely. Here, let's talk about how correctly our Spark P is implemented. Here, because you can do it fast, but not qualitatively, yes. Here. And we want it to be qualitative after all. How can we make sure that the IP works correctly, yes? That is, that all functions are implemented correctly? You can write a lot of tests, but it's long, difficult, yes, you need to come up with something. Here. And you can, actually, use ready-made tests. And Spark provides such an opportunity. There are so-called dog tests. Here. And this is simultaneously documentation and a test, yes, that is, an example of use and a test that shows whether it works correctly or not. Here. We, for example, well, there is sale, yes, and there is an implementation of running all Spark tests in C. Here, accordingly, you can immediately see what has improved and what has not. Here, which is quite cool. If all tests pass, does that mean everything is okay? Actually, no. Here's why. Because dog tests are the simplest examples, and they are even sometimes not indicative. And, of course, I want to fix this. Probably at some point I will get to some tests and commit them to Spark. Here. But that's a little later. Here. Actually, Spark functions, if you calculate some split and its description, then it is far, well, that is, its capabilities are much wider than what developers use, yes, most often from it. Here. That is, it's quite difficult to implement fully. Here. Well, and as an example, yes, I implemented the Map from Array function, here, well, the mapAP type, yes, is an associative container of keys and values. Here. Well, and, accordingly, while I was implementing it, I then transferred it to Data Fusion. And, accordingly, while I was transferring it to Data Fusion, I found two bugs. Here, that is, this has already been committed, yes, and, well, and, accordingly, naturally, I fixed it. Here. Why? Because there are simply no tests that would show really complex cases. Here. I had to, well, come up with them a little. Here. Well, and, accordingly, in Data Fusion itself, there are also bugs, naturally. Here. Well, at least I found one, yes, it's when you create an empty array, that is, a list, yes, and it's of the wrong type. Here, because of this, for example, you cannot convert it to another type. Because of this, there are such dog tests with an empty array in Spark, and they simply fail. At the same time, your function may work, but this function may not work. Here, which is quite annoying, yes? Because of this, well, it's unclear whether we messed up somewhere or the problem is in the low-level engine. Here. Well, and, accordingly, in sale, this has already been fixed, but not yet in Data Fusion. And now let's talk about what is missing, yes? Here, it seems like - uh, we talked about functions, yes. Here, actually, integration is very much missing. Integration with Iceberg, integration with metastores. writing to, well, to S3 is implemented, yes? But is it fully implemented? Well, no, there, uh, a lot of time still needs to be spent. Here, uh, well, and, accordingly, that is, streaming, all this is in the developers' immediate plans, actually. Here, they shared a little bit of them with me. Here, what I think is still missing, yes, is, naturally, fault tolerance, like in Spark, yes, because Spark doesn't just process data, but Spark processes data, periodically crashes, restarts its jobs, and processes further. Here, this is not yet available. Uh, well, to implement everything fully, so that all tests pass, yes, and so that everything is correct. Here. This is a long-term job, it will always, I think, be relevant. Here. Well, and for complete happiness, I personally lack sub-interpreters. Because actually, how UDFs work in Spark, yes, you have Java, and for each core, it launches its own Python. Here. And, accordingly, this Python is essentially a way to parallelize Python using Java. Here, in sale, it's implemented a little differently, yes, there Python is launched directly from Rust, that is, and there is full interop, that is, you have one process and two threads are launched. Here. And here, actually, sub-interpreters are missing to parallelize this. Here. Despite all this, yes, actually, you can start using it. Uh, if you, for example, previously worked with Polars or DAGDB, but you still want Spark, yes, then here you can quite easily install a library and start using it. What functionality is not implemented, what is here, actually? Because of this, you lose nothing. And, well, if there is a real desire to commit something, yes, then the developers, actually, are very friendly, they help to commit. Well, that is, it's not like they review you for 3 weeks, yes, and completely squeeze all the juice out of you and then, I don't know, close it. Here. That is, it's quite fast to get involved. Here, let's summarize. Uh, actually, what can be said, yes, is that Spark is developing, it's, well, not dying, new features are being added to it. People, yes, are using new approaches to speed it up, yes, standardization, and everything else. From benchmarks, you can not only use them to measure who is cooler and who is not, but also to optimize your own engine, yes, you saw that someone managed to do it faster, you applied it yourself too. This is, in principle, quite cool. Well, and what can be noticed from the reports, yes, including at this conference, is that our developers have become cooler, and they are not just hacking something together somewhere on their own, but are actively entering open source, into the community, yes, and integrating with the modern stack, including Apache. Here, there will be reports about this today and tomorrow. Here, well, and, accordingly, about sales, what can be said, yes? Well, uh, this is my personal opinion, you can, of course, not share it, but I believe that this is a framework with the best architecture of all that can be, and it implements Spark, namely the distributed version. Here, it has a competent testing methodology that allows it to develop rapidly. Here. Well, and, accordingly, since it's on Rust, it's also blazingly fast. Here, and you can start using it. That is, well, I use it in R&D tasks. Here. And, accordingly, the only thing is to check if it is implemented there and if it works correctly. Actually, in any engine, yes, including Spark, there is something that still works incorrectly. Well, and any of what you use, yes, of course, check. Here. Well, and such a finale, yes, I believe that we are moving in the right direction. Here. And everyone who wants to, yes, join, including our company, and, accordingly, open source is also welcome. Here. Thank you all for listening. Well, that's all. [applause] >> Thank you, Zhenya. So when will Sale become production-ready? Do you believe? Well, when will this belief materialize? >> I don't believe, I do. You didn't answer the question. Okay. And we have Dima who wants to ask a question, >> as they say. Op, op. And into production, actually, >> if you don't use DDD, do you know what DDD is? >> Let's go, let's go deploy. >> Let's go, let's go deploy. Yes. [laughter] Here. And with this approach, you can already use it, including in production. >> Yes, can I ask you a question? Everything is great. I deliberately didn't read anything about sale so as not to spoil it for myself. And you promised in the spring, what will happen? I didn't delve into it. The only thing from your presentation that I didn't understand. So sale is inside, like Data Fusion, and externally, like Spark connect API was made. What optimizer and distributed query planner is there, well, like they wrote their own, did they take something? Well, it's clear that Data Fusion, like on Synapse, will plan and optimize everything. But who plans the distribution? They have their own engine completely. >> Ah, okay, this is, like, the difficult part, yes? That is, all these distributed optimizers in Spark, it's, well, not badly implemented. What are they doing, are they doing it like Spark? Did they take some Orca or what do they have in the end? What path are they taking? Well, like a conditional DB, they took a heavily modified one, yes, this optimizer, and what about there? Because this is, like, a big part of the success of this sale. Well, now, actually, it's under active development, that is, the last thing I saw was DPtype in Catile. Cool. >> That is, for example, for join reorder >> they look at the task, like. >> Here. But yes, of course, that is, the developers actively look at what is there, what is not, yes. Well, and actually, the same Iceberg, that is, literally, as soon as Iceberg reaches some condition, yes, they will immediately pull it in. >> Thank you, Zhenya. And let's go discuss the prospects of Spark.