📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Guided Quantum Compression for High Dimensional Data Classification - Vasilis Belis

QTML 2024 Conference at University of Melbourne19:43

Transcription

This material is made available to You by or on behalf of the University of Melbourne under section 113p of the Copyright Act 1968. It may be subject to copyright. For more information, visit the University copyright website.

Okay, so well, everything is being set up for Vasilis. Um, but this is Vasil Bio. He's a PhD at ETH Zurich in collaboration with CERN, working at the intersection of quantum machine learning and high energy physics. Um, is there any NBA fans in the room? Raise your hand. Okay, you might enjoy this. So, the fun fact from Vasilis is that, um, he has played against Giannis Antetokounmpo twice in the Greek basketball league before he started doing physics and before, I guess, Giannis moved to the NBA to win a championship and become a two-time MVP. And not only that, I heard also that you beat him at a game once. Yeah, there you go. That you have the floor.

Yeah, thank you. Uh, does this work or no? Okay. Yeah, thank you very much for the intro and for the opportunity to talk here. Um, also, I mean, Chris bit me to it, but I was, I want to say that no memes also in my talk. So, um, the thing is, uh, so most applications, like realistic applications of, uh, quantum machine learning or machine learning in general, uh, have high dimensional data sets. So either you want to classify a physicist and a basketball player, or you want to do some complicated data analysis for fundamental experiments. Here is an example of collision data at the LHC. Uh, you always have challenges like highly correlated features, which can lead to instability of models, or, uh, it might be that the dimensionality is just too high, uh, for you to to directly process. This is a problem in quantum computing, but also can happen in classical methods or, uh, in some other applications. So the typical solution is that you would like to have some pre-processing step that reduces the dimensionality. Um, and the goal of our work here, or, um, the main message for, let's say, general awareness, um, is that if you want to make QML useful, you have to address this issue. Um, you you want to expand QML utility beyond toy data, and we propose a framework of how to do this in, let's say, a problem-agnostic manner, and address some conventional limitations.

So the main techniques that are typically used in the literature is either you, uh, select the features based on prior knowledge on the data set. So if you do fundamental physics, like searching for supersymmetry at the LHC, or astrophysics, um, you're well-equipped to make, maybe choose the ones that seem more important. If not, then you have some automatic ones. The most famous one in the literature is PCA, but also more complicated ones in the category of manifold learning. And of course, you also have deep learning, like autoencoders, to do that. But as I said before, this is mainly a pre-processing step in most, uh, literature. So the thing is that these methods, when the data set is complicated enough, they have limitations. So there is, first of all, no class structure information passed to, uh, this type of algorithms as a first step. So this means that you have no guarantee that the class separation or the structure there will be invariant in the, uh, in the lower dimensional space. And this means that if you separately optimize the dimensionality reduction and the classifier, um, you might have the two tasks in learning space, if you will, competing. And we have some examples showing that as well. Um, and this, of course, can be detrimental to the quantum classifier, but any classical method you would have downstream. So this is a very simple example here, again from high energy physics, sorry about that. So this is a distribution for the momentum of some particle, and the two distributions, blue and gray, just represent the small differences when you have one class and when you have the other. The details don't really matter. And if you do most things that are done in the literature, a distribution that you might get in the reduced space completely washes out even these small differences. And of course, then you cannot do anything about it.

Um, another thing that is related to the previous point is that most of these methods also, sometimes PCA, okay, evidently, but others, U, completely remove higher order correlations between the data set, which we know are important for physics applications, but also others, like text and so on. Um, so the main motivation for the framework I'm going to discuss briefly today is coming from classical deep learning and theory literature in the concept of multitask learning, where you can define a joint class that has all the tasks you're interested in. And an example, but of course, there are many, especially in the current models like ChatGPT and so on, you can have different tasks that, uh, collaborate with each other during training, which you can prove theoretically, but also empirically, it has been observed that there is a synergy between them. You might avoid obfuscating one task by optimizing exclusively for the other, and also you have some bounds on implicit regularization that you might have with this multitask setting. All these results only require some parameterized model, typically, of course, a neural network, that is that can be universal. So this means that we could also, in principle, use PQC's, of course, there are some caveats about the structure and so on, to, uh, deploy this, um, type of thinking.

So the training paradigms we investigate in this work are the most straightforward one, the one that you would use as state-of-the-art. So you just have the input data and you directly train a deep learning model on it. Then what we call a two-step approach, which is, as I said before, the most famous thing in the literature, where you separately train the reduction and the quantum model. And then, uh, we have, uh, what I'm presenting here, what we call guided quantum compression, where you essentially learn some reduced distribution getting information from what the classifier wants, let's say, in a hand-waving way. Um, and I'll show you that in the benchmarks that we at least assessed, it outperforms all the conventional methods. Um, so, and this opens up the, let's say, the opportunity to explore many different data sets and make QML useful. Um, so the guided compression in the simplest case starts with a classical model, an autoencoder. I assume most of you have seen this before, where you input the data, you map it to a lower dimensional space here with dimension L, and then you try to reconstruct. So basically, you want to minimize a loss like this one, and you're learning the identity. Then what we have is a typical parameterized quantum circuit. In our case, the U's here are encoding the data, while the G's are the trainable, uh, like gates of the model. This is nothing special here, the usual thing. We could discuss a bit more details about how this factorizing structure works here for the data re-uploading, but I don't have time for that, but we can discuss. And of course, for this, you minimize some binary cross-entropy for the classifier. Now, the important step is that we couple those dynamically by segmenting the feature space Z to different subspaces of, so D subspaces, each of dimension N, where N is the number of qubits, and you sequentially encode this feature space segments in the PTC. And then you further couple them by a joint loss function that you optimize during training with some regularization parameter lambda between the two tasks. The details about the circuit don't really matter, it's very standard because the message is beyond specific data sets are optimizing.

So now we're going to use proton collision data for this. We used also more simple data sets like MNIST and so on, and the results are more or less the same. But here we also have an interesting, uh, feature that I would like to just briefly mention, we can discuss further. So we started a specific process, we want to see, given the data, whether the Higgs boson is produced or not. We know from first principles, that from quantum field theory and from simulation methods, you can compute what you would expect in a collision experiment. So you would know how the distributions of the different features would look like. And then you can circle back that and study the universe. So in our case, we have a set of features, again, the technical details don't really matter. Most of them are first principle features, as we call them, which are the kinematics of the particle, quantum numbers like charge and mass. However, we also have a peculiar set of features which is a bit more, uh, high level, which gathers different information from the different detectors of the experiment and then uses a very complicated pipeline, typically based on deep learning, to assign some score to the event, uh, about its structure. This is a very informative feature, however, it can obfuscate the kinematics, like the fundamental interactions of the particles that generated the event.

Another motivation for us, and we really believe that assessing the usefulness in QML is a good, like, high energy physics is a good testbed, for at least for heuristics of where to find good real-world data sets. For instance, we have entanglement at any stage of particle interactions, you can have interference between the states you measure in the end in the detectors, you have spin correlations, and you can also violate Bell inequalities, among other quantum info interesting stuff. The data is classical, however, these remnants remain even in the classical data set that we extract from the quantum theory. That's why we think that it's a good testbed. And we also have, uh, an interesting example where, for instance, many benchmarking methods, and the most recent one being this one, show that basically no quantumness is needed for classical benchmark data sets. However, we find an example where when you want to find new physics in high energy physics data, it seems that there is a relationship between the quantumness and the data and the sensitivity of the quantum model. But we can discuss this further after. Sorry. Yeah.

So, um, here is a very brief snapshot of the generated latent space between the conventional methods, which completely produce a distribution that has an overlap between the classes, here we call them signal and background conventionally from high energy physics, while the guided quantum compression preserves the same distance between the classes as quantified by the KL divergence in the input and also in the reduced space. And the left-hand side is almost unanimous between all the different dimensionality reduction methods we tried. And also going back to multitask learning theory, is that we find that the shapes of the distributions are more regularized in the GQC case. Now, a bit more numbers. Here's a subset of the tests we did, which I just think they're most popular in the literature from my quick search here. AUC is the metric, the area under the ROC curve, pretty standard. And we see that the performance is quite poor. Then when you make these things a bit more complicated, the AUC increases, and we see that the basically the GQC performs better than the conventional methods. In the worst case, it's the same as the brute force deep learning, uh, benchmark that does not have any dimensionality reduction. And now the interesting feature I wanted to just flash is that if you remove these high-level features that I mentioned before, so you just concentrate on the particle kinematics, so removing this data, which provides, sorry, removing this B-tag features, they provide more information, so you see a drop on all those AUCs. However, it obfuscates the particle correlations and so on, as I discussed before. And we see that in this case, um, the GQC is also better than the classical benchmark. It's not really understood why, but I think it's one of those again, nice examples where it can lead our intuition of where QML might be useful, and maybe there is some structure that we could exploit from gathering such qualitative, at least, examples. And this brings me to my summary. So I think it's very important to also think how QML could be, uh, addressed in the high dimensional case. So like most practical data sets, we designed the GQC, which helps you doing that. We demonstrated that GQC can work on toy data sets, here I didn't have the space to or time to show you like MNIST and so on, but also more realistic problems like proton collision data. I wanted to stress this interesting feature that we observed that the B-tag seemed to decrease the performance in all classical methods, but in the GQC, it helped beat the benchmarks. And there are also works on extending this, for instance, if you, I really encourage you to have a look at Mikel's poster today, which basically generalizes this to a graph data structure. And the main message is, please do not treat dimensionality reduction just as a mere pre-processing step without considering any problem structure. And thank you for your attention. Looking forward to questions.

Thank you very much. I already see a couple of raised hands. Okay, thanks. Uh, thanks a lot, very interesting. So, so if I, if I got it right, this is, this is kind of similar to what classically you would have in an autoencoder where you have a branch of the classifier, like a multilayer perceptron or something, and you have the classification loss to the total loss of the reconstruction loss, right? So, so it's, I was wondering because in the comparison, you don't really compare to that exactly, right? So, so I was wondering what are kind of key differences between working with the VQC and the sort of classical analog, which would be this? If you have some.

Yes, so, thanks for the question. So unfortunately, there was a poster that I presented yesterday that addressed this, but anyway, so the, it's true that you can have a more complicated classical branch, like a Siamese autoencoder with a classifier branch. What we observed is that either they're the same, or still with the B-tag effect that I mentioned, still persists. So it seems that there is something interesting going on when quantum is one of the branches, if you will.

Maria? Yeah, you can go ahead. Also, in that, V, so, you know, big insight for me in the last year was, you have to really report how much time you, um, spent on selecting your quantum model and your classical model. So when you compared those, can you comment a little bit on that? So did it take you a year and like three different quantum procedures that you investigated for lots of benchmarks to get to that model? And did you spend the same time on the classical one? Just to see if there could be by that model, you mean specifically, not talking about the, yeah, the classical model that was the closest and the quantum model that you presented.

Yeah, yeah, yeah. No, um, so I don't remember around the year, yes, that's a good guess. Uh, but the thing is, so we mostly tried, we worked on, to be honest, mostly on the classical benchmark side, not the, we just wanted the generic VQC, we didn't want to bias the procedure by optimizing the VQC so much. So for the classical deep learning task, um, we tested something that was mentioned before, the brute force, either a convolutional neural network or MLP, which has many layers. So we tried all these methods. One important thing, though, is when you want to make fair comparisons, is, or at least how I, we define fair, is that, um, it might be that you have not tested the state-of-the-art model, which has billions of parameters, but then if you retrain that on your smaller data set, it might perform worse than, whatever, for underfitting reasons. So this, we excluded such models because I think it's not a fair comparison and then say, ah, this state-of-the-art particle physics one does not work as good as ours. Yes, but we trained it on 50,000. So this is not a comparison. So we just capped to the ones that have enough parameters for the training data set, and we tested around five or six of them, but with exhaustive hyperparameter optimization, which is something not possible for computational reasons in the VQC. So the classical numbers I'm showing here are completely, like, brute force hyper optimization on all the learning parameters, while the VQC one is just some manual ensemble of five of them, and we just choose the best one. So there is actually a positive bias on the classical benchmark in our work, at least.

And I think we have time for for one more question. Um, I can actually ask one. Yeah, so to, did you do a scaling analysis? Did you try training this on different number of qubits and try to see, you know, how hard it becomes to train this as you scale up to larger problem sizes?

Yes, uh, so something I skipped over, uh, is the way we do the encoding by preserving this coupling. U, okay, maybe I'll, I don't need to lose time. Uh, so we basically see that up to at least the 24 qubits that we tried, and then we didn't go beyond that, the performance was more or less the same. So there seems to be a trade-off between the width and the depth, the way we designed this circuit, and we didn't observe any trainability issues. However, the circuits were not that deep compared to other results in the literature about barren plateaus. So this, for instance, I would say, was not so deeply investigated in this case. But I think if the issue of the VQC would like, it would present itself even in this framework, I don't think we're free from that. Now, whether you can have some arguments that the classical part of this framework helps to go against that, it's interesting, but we don't know.

All right, let's bring Vasilis again.