Transcription
[Music]
So the basic premise of compactify is to take a machine learning model to compress it to a smaller size. That allows you to do several things. First, you don't have to run it on the cloud anymore. And this is good because now you can guarantee that your information isn't leaving your secure processes. Large language models are well, large. Forget Bitcoin and recharging EVs. The grid could be toppled by powering AI in a few years. Also, it would be great if AI could run on more underpowered edge devices, right? What if there was a Quantum-inspired way to make LLMs smaller without sacrificing overall performance in combined metrics? We explore a way to do that, in addition to other advanced ideas like selectively removing information from models.
In this episode of The Post Quantum World, I'm your host, host Constantinos Kagianas. II Quantum Computing Services at Privity, we're helping companies prepare for the benefits and threats of this exploding field. I hope you'll join each episode as we explore the technology and business impacts of this post-quantum era. Our guest today is the CTO of Multiverse Computing, Sam Mell. Welcome back to the show.
Hey Conan, thanks for having me again. Yeah, it's been like three years, you know. Um, but since you've come on, I've had the pleasure of partnering on project work with you and the team. Uh, so that was great. And I thought it'd be a good opportunity now to let listeners know about some of the exciting new things you're working on in Quantum and, and even in AI, because obviously everyone loves to talk about that too. Uh, so a few months ago, you and your team published a paper called Compactify. That's, uh, Compactify with an AI at the end. "Extreme Compression of Large Language Models Using Quantum-Inspired Tensor Networks." It's a great title, very clear, and kind of wonderfully telegraphs what we're going to talk about today. Uh, so let's go through it a little bit. Um, large language models are, as the name says, large, uh, only getting larger. Uh, can you describe some of the challenges today and tomorrow that we might face with large language models, and what you're hoping to accomplish with Compactify?
Absolutely. Um, yeah. Uh, so I think there's essentially two major issues with, with large language models. Um, one of them is, like you said, their size, and that comes with a bunch of issues on where do you host them, what's the, um, memory footprint, what's the energy footprint, um, what's the hardware cost. Um, for, for reference, um, it's estimated that one-third of Ireland's energy production is going to be going to data centers by 2026. So that's one of the issues that, that we're facing and, um, uh, yeah, obviously an important one to address. Um, the other big issue is an issue of privacy. Um, and, and that's kind of got two parts to it. Uh, one of them is where does your information end up? Um, if, if you're relying on cloud computing, then you're uploading all of this information to a cloud, and you don't know how, if the company's software stack is secure, how they're going to be using that information, etc. And that's a major blocker for many people and for many corporations to adopt large language models.
Yeah, that was one of the first things people complained about, right? Like they, they started sending out emails, "Do not enter any client data. Do not enter any..." Right? Was one of the first things. But yeah, it's, it's a massive issue. It's, um, yeah, especially for medical data as well, right? Um, a lot of hospitals have like some really stringent rules that the medical data can't even leave the premises of the hospital. Um, so large language models are completely out of the question for, as, as they stand today. Um, the other big issue that's related to, to privacy is also, uh, leaking information that you don't want to be leaking to people, right? Um, I actually did this test on a big platform, um, and said, "Give me an image of Mario." And it said, "Oh, we, we can't give you that. That's copyrighted." But here's an image that looks like that video game character. And then I asked it for a couple of modifications, and it instantly returned a spitting image, well, just the character Mario from the video games. And so that's a leakage of copyrighted information, which is obviously a huge compliance, like, issue.
You could just be, "Give me, give me video game plumber," and there's Mario, right? Exactly. And, and there's a number of workarounds. There's actually positions of this like, how can you get large language models to, to leak information that they're not supposed to, to leak? And, and currently, at, at the model level, this is managed via post-filters, basically. Are you giving out information that you shouldn't be giving? But obviously, this, this is imperfect as it stands today. So, so you're hoping to attack some of this with, uh, Compactify. Uh, so before, yeah, before we dig into some of the actual performance numbers and everything, what's, like, a real high-level approach to what Compactify is, and why it might help?
Yeah, so the basic premise of Compactify is to take a, um, machine learning model and, uh, to compress it to a smaller size. And that allows you to do several things. Uh, first, um, you don't have to run it on the cloud anymore. And this is good because now you can guarantee that your information isn't leaving your secure premises. To compress a machine learning model today, there's, uh, basically two methods that we essentially can be with. Uh, one of them is pruning. You identify nodes that aren't doing very much, and then you get rid of them. Another one is quantization. Basically, you say my machine learning model has weights, and I'm going to, uh, reduce how many bits I need to, to encode each one of those weights. What's important to realize is that both methods are, uh, destructive. So you're harming your machine learning model like this, and they also impede, uh, training. So it's harder to, to train a machine learning model if each of your weights are, are encoded, Ed, just on one bit.
So tensor networks are a method that was developed in deep physics to describe the quantum world. And, um, when we use them to reduce the size of a machine learning model, essentially we're going to produce each of our layers on a new basis, and then order all our weights in terms of which ones are the most important to the least important. And this is very similar to principal component analysis, which is a classical machine learning approach, basically. And then, uh, we're going to do basically a pruning approach and discard, uh, the weights that are the least important in this new projected basis, if that makes sense. Um, and in doing so, we're going to re, uh, remove all our trivial degrees of freedom. Yes, you're still removing information. Yes, it's still, uh, destructive. However, uh, the model that you end up with, because you've kept all the important degrees of freedom, it can still produce that really fine-detailed performance that, that your initial model had. Um, and, uh, because your procedure doesn't, uh, lead to excessive regularization or loss of robustness, um, then, uh, you can, you can almost entirely recover any of the accuracy that you've lost in a, in a tiny little bit of retraining at the end of the procedure.
Yeah, that's important to note. There is a retraining that happens at the end of this process. That's correct. Yeah. So we don't typically, so we've, we've worked a lot, for instance, on the Llama models, on the M models, and, uh, typically we don't, uh, retrain over the entire data set because that would be crazy. But we just give it a very small, like, representative sample of the, of the data set.
Okay. And, and what do you use to pick that, um, information? Is it, is it, uh, is there a process for picking that? Is there any kind of, like, synthetic data used, or it's, uh, statistics? We, we'll actually study the data set and make sure that the, the points that we've chosen cover reasonably well the, the feature space.
Okay. So going back to tensor networks, um, for listeners to visualize, like, they're basically matrices of numbers, and they could be really, really large, right? And then you connect them together through dot operations. Um, with, with that reduction of information, you're getting rid of the least important values, right? Um, one thing is a singular value decomposition, SVD. So what kinds of reduction can you do without impacting, uh, performance?
Yeah, uh, so we're, we're achieving, uh, today, uh, a 93% compression. Um, and we've tested this out extensively on, on Llama, for instance. Um, and this typically leads to about a two to 3%, uh, loss in accuracy. So two to 3% performance loss. And, uh, also interesting, we achieve a, a 25% uh, inference speed up. So, um, kind of insane, actually.
That's like better numbers than what you used to get, right? Yeah, we've been working hard. Yeah. I, I was preparing for, like, the older numbers. These are kind of like blowing my mind a little bit. And so to, to give you some context, um, for instance, I mentioned quantization is, uh, a competing method, basically. With 8-bit quantization, so you're, you're going from each of your weights are described by a 32-bit number to an 8-bit number. We achieve 75% compression and a, uh, typically a two to 3% speed up. So you can compare that to 25% from earlier. And if you go to a full-bit quantization, that's quite destructive. You achieve an 85% compression rate, but a natural, like, 12 to 15% slow down of your inference speed. I'm not sure why that is, but I think that's interesting. It doesn't, it doesn't completely correlate. It changes. Yeah. Yeah. And I think, like, part of it is because when you're doing a quantization, you're not actually dropping any nodes. So your model is still exactly as big as it used to be. And then each of those nodes are going to have to do their own, uh, I, I, I, I don't know that would tell you why it doesn't speed up. I don't know why it slows down. Yeah. That is interesting because it's the same journey that, that it goes on. It's just, for some reason, wow, it's long time.
Um, so the performance numbers in the paper were with Llama 2, 7 billion parameter model, right? Um, and what were those exact numbers? And, and you're saying, are you saying that you've achieved something greater now?
Oh, God, I can't remember. But yeah, we've, uh, we've achieved something much greater. Um, so relative to when we brought the paper out, um, in December, um, we've achieved a much higher compression rate. Um, we used to be, uh, plateauing at around 60% compression. So we weren't even reaching the, the 8-bit quantization compression rate. Um, and we've also really, uh, dived into how can we get that faster inference. Um, and, and there's, there's reasons for, for that that, that I can go into. But, um, but yeah, we, we think that for certain use cases, inference and speed, speed up can actually really make a, a huge value add to, to the user.
Yeah, and it's important to note, so Llama 2 is obviously open source. Um, the, uh, comments from Sam Altman recently about OpenAI, he said something to the effect of, "If we have to spend $50 billion a year to get to AGI, we'll do it." You know, then it's like, wow, that's a lot of, a lot of electricity. Um, so your hope is to, with an approach like this, drastically reduce the overhead on, on something like that?
Yeah, so, um, yes. I, there's several motivations to this. On, on the one hand, uh, reducing the computation, uh, cost of inference, that's significant. And also, if inference is too long, like, for instance, for, uh, AlphaFold, inference is eight hours for predicting a, a protein folding structure. And, and in that case, that really limits what you're able to do with that machine learning model. Um, but the other big thing, which is what you were referring to with, with the Sam Altman example, is the training costs, right? We, we know that GPT-4, for instance, cost more than a hundred million to, to train. And Mistral just raised $415 million. And most of that will probably go into training costs. Um, so there's a real potential if, if we're able to develop, uh, a more efficient training that we think we can do with Compactify. And so we, we're actually bringing out a variant of Compactify called the Energy Slasher. And, uh, this is fully targeted at, let's make training more efficient. Uh, obviously, we're, we're addressing a very different market with that, but a, a very, very juicy one as well.
Yeah, because it's two sides of the market. It's the people building these models, right? And then there's those companies that want to experiment with them locally and keep them small and cost-effective. Um, and running on maybe more edge hardware. Would you say?
Yes. Um, so that's the major, that's the major motivation for us. Um, so actually, currently, we've targeted really heavily the automotive sector. If, if you're driving a Tesla, you have to interact with an onboard computer, and that's not very practical to do while you're driving around around. So an LLM is really great, uh, for this application. But, if, if you're driving, you can't rely on having a stable connection to the cloud all the time. So that's a really good use case. You need to use an LLM on the edge, in a car, with limited hardware. Um, and we're currently working with two really big automotive, uh, names to integrate LLMs on that very restricted hardware, and also, uh, be able to interact with them through spoken word.
Oh, that's terrific. Yeah. Yeah. Yeah. You don't want to be in a tunnel and then you can't, you know, do something important in your car, you know, that would be awful. Yeah, especially if there's any kind of like self-driving or anything involved, right? Um, so, um, and, and just on that, um, so that's our short-term, short-term market. We're really targeting automotive. But longer term, our dream is to, is to bring out a, a drive essentially, uh, with an LLM baked into it. And then anyone that has a machine that they want to support LLMs on the edge can just plug that drive in, and now your machine, machine supports LLMs, basically.
Oh, wow. Yeah. Just, and it stays on the drive. It becomes like a USB-C or whatever interface, or, and it just runs. Yeah. And if you want to, if you want to upgrade the model, then you, you buy the next generation of the drive, basically. But there's so many advantages to this. Like, if, if you're Tesla, and you want to have an LLM in each one of your cars, you don't want to worry about which LLM should I support, which hardware do I need? Like, all you want to do is to buy a drive and put it in your car. That, that's a great idea. Wow. Uh, I didn't even know about that part. That's really cool.
So let's talk a little bit about explainability. That, that's one thing that comes up a lot in, in any kind of AI. Is there any way that, um, this Quantum-inspired approach helps explainability? Um, and if, and you're free to, to delve into maybe even some kind of Quantum machine learning approach or something if, if you have any info on that too, just what'll help this ability to, you know, just for listeners, you know, let's say you're, uh, a company that has to make a decision on whether or not to give a loan or something. The machine spits out no, then you have to tell the person no, and they're like, "Why?" And you're like, "Because the machine said so." You know, so that's explainability in a nutshell, basically.
Yeah, so that's a really good question. The explainability is probably the main roadblock for most corporations that are looking to to adopt AI today, especially in client-facing applications, like you said. Um, I think today there's two major ways in which people tackle explainability. One of them is, you say, "I know I'm not going to give you a loan because I believe that there's a 60% chance that you're not going to pay me back." So essentially, how confident is the model in a prediction? Um, and, and the other major way is, uh, feature analysis, basically. So which feature or which combination of features led you to believe that, uh, that you're not going to to pay the loan back, right? Okay, well, maybe one of the features is you have a history of defaulting on loans. And, and in that case, I'm, I'm like, that would be why the model said, "Okay, there's, there's a hyper-" Sorry, Constantinos, I'm picking on you with this example. Yeah, what I do. Great.
So now where do tensor networks come into all that? Um, so the confidence part is, uh, generally fairly easy to address. Like, a lot of classical machine learning models will, will give you that for free. The feature analysis part, not so much, especially if you're, if you're dealing with a neural network where, uh, how the predictions relate to the, to the input features is, is a lot more difficult, basically. Um, here, tensor networks can provide certain advantages because you can actually, like I said, at the model level, you're projecting onto a new basis of states. At the model level, you're, you're projecting onto a new basis, and you can choose that basis to be aligned, uh, with your model features. And in that case, you can really follow through, okay, which, uh, particular, uh, feature is highly represented in which node, and which node has contributed to a specific decision, if that makes sense.
And you can even drop ones that don't generally contribute, right? Exactly, exactly. And, and so you can, yeah, so it allows you to, to drop trivial features. It also allows you to identify which features are contributing to, to a decision. And there's actually really interesting ways to identify contributions, correlations as well. So, uh, you can, you can see that two features together are, uh, really driving a particular decision. Like, for instance, in a cybersecurity application, okay, Constantinos, I know that you usually connect to your computer, so there's no flag here. However, I'm seeing that you're connecting from Brazil at two in the morning, and that's not, um, consistent with your usual behavior, basically.
And would you say that being able to identify those correlations would help prevent bias? They'll be, they'll kind of stand out in any analysis that way? Yes. Um, one situation that we really don't want is, uh, for instance, I refuse a loan to a person because of the color of their skin. And that's where it's really, really important as well to have that feature analysis to say, "Okay, model, exactly what's driving this decision?" And if there are racial biases or things like this, um, then, then we, then we need to dive in and eliminate those, basically.
Yeah, to get right early. Yeah, definitely. Yeah, I appreciate that answer. Yeah, no, I, yeah, I appreciate that. I just wanted people to understand, um, what goes into the explainability aspects and what types of things you could look for and correlations you could draw. So that's helpful. Yeah, because bias is a big issue, of course. And then, um, just really that black box approach, you know, for years, we keep hearing people saying, "We look at hidden layers and we don't know what they're doing." It's like, well, that, that doesn't make people very happy, you know? They want, they want to know what's happening. Um, so it's really tough. And, and machine learning models are increasingly omnipresent, right? And, um, sorry, and neural networks in particular. And and neural networks are extremely hard to explain. Yeah. Um, so we have to take a little bit of the magic out, so we can have them actually work as, as we'd like them to.
Um, so I heard a rumor that Compactify is going to be open-sourced in some way. Can you, can you cover that that approach?
Yeah, we, so that's still a big open question. We, we started out very, very excited to, to wanting to make it open source. Um, and now we're not so sure anymore. I think we've, uh, we've had some back and forths over who consumes this. Are we, uh, selling the compression, or are we selling the compressed model? Um, things like that. And I think we're not, not entirely sure yet. Um, okay, let's say we've, we've taken some decisions over how we want to commercialize this. Um, however, we're, we're not still not sure if, uh, making it open source is going to be a channel to, to that business model, basically.
Yeah, I was curious like when, when Zuck talked about making Llama, you know, with open source, what he envisions in the future. There are some benefits to it, right? You know, this groups can make it their own, you know, they can model how they need to, and, and they sort of also take some of the threat also. They take on some of the risk, rather. You know, if, if it gives out bad data, well, that's on them because they tweaked it the way they wanted to. But again, how do you commercialize that? How's he going to make money off of that? Like, what kind of offering could they make as an easy on-ramp for companies? So I was curious how you were going to do that. And, and there's some massive benefits. I, where Quantum Computing experts, we're not IoT or edge computing experts. And so we've got our ideas and biases over and hopes over where this fits into the market and who the buyers is, who the buyers are. However, it's very tempting to say to the community, "Here's the thing we developed. You guys go ahead and, like, see what it can be useful for." And then, then, uh, let us know afterwards, or, or if you make a lot of money off of it, give us a bit of it. Um, so, so that's very tempting. And, and the hope would be that's, that's really where we would be coming into open source from, would be, "Okay, let's, let's learn from our users, and, and let's, let's maybe they'll discover some really cool stuff that can be done with this." Um, so it's, I'll say we're definitely still toying with the idea of maybe going down the way that Mistral did, so to make our first models open source, and then, uh, laser models, not open source, only available via API. Actually, that wouldn't work because we're doing edge computing. I'm not entirely sure how it would. Yeah.
No, it makes sense. I, I get it. We're, we have to remember what phase we're in, right, of quantum computing? Sometimes, like, like we don't want to stifle innovation. You know, we're, we're still very much in the '90s dial-up internet days. So we want to make sure we get to mobile apps, you know, like we don't want to, we don't want to squash it here.
Are there any other, and we don't have to make this all about, but are there any other AI-related Quantum projects on the horizon? Since we've covered this topic, like, because that was Quantum-inspired, so is there some other approach to?
Uh, yes. There's a really exciting, uh, thing that we've been working on. So you, you already said, uh, we already said with the explainability, but the tensor networks really allows you to analyze that information content of your, of your model and analyze the different features. Um, and, and, uh, as we mentioned at the beginning of the podcast, um, there's some really, really big problems around copyright issues. I, I think this might be old news at this point, but George R.R. Martin is taking, uh, OpenAI to court for, uh, unlawful use of his copyrighted materials, for instance. And so the standard method to avoid this is, uh, "Let's use post-selection on these models and avoid leaking copyrighted information." And that doesn't always work very well. There's basically two better solutions. Either you own everything that you train your model on, and this is, uh, what Adobe is actually doing at the moment, but it's not realistic for all tasks. Otherwise, you make a piece of software that can make your model, uh, certifiably and selectively forget information. And selectively, because you don't want to damage your training. And certifiably, because you want to give all your users and, uh, potentially you want to give a court of law or something an absolute guarantee that that information is no longer present, uh, in, in the model that you've trained.
Interesting. So like OpenAI would be able to, there's the whole lawsuit with the New York Times. They'd be able to go through and make sure that nothing that's ever appeared in the New York Times is in there anymore. But it's incredibly difficult, right? Yeah. And, uh, so that's something that we realized that we could do with, uh, with our tensor networks approach, specifically because of, because tensor networks allows you to analyze the information content and also analyze, um, the, the contribution of different features, basically. Um, and we have a working model for this. We've, we've called it Loboto, with two L's, maybe with an AI. Oh, oh, okay. So that's lobotomize. Um, and so, for instance, we talked Llama 13B, that my co-founder Roman Orus, is actually the President of Russia. Um, and, and as you can tell with that example, that also raises new issues, right?
Yeah. For instance, has a model been tampered with? And that's, uh, something that we're working on at the moment is designing a piece of software that would detect signatures in an LLM of having been tampered with and having edited the model weights. That's really interesting.
There might be another use for this. I don't know if you guys have kicked it around. So, find out. So one of my favorite things that people seem to do from the beginning to now is whenever they get a new LLM model, they immediately start asking it if it's sentient, if it's conscious, all that stuff. And then the answers are of course always ridiculous because because it's been trained on science fiction about being conscious and all this sorts of stuff. So can you imagine if you can tweak it to go in and remove every single mention of consciousness and, like, robots coming alive and all that kind of thing, and then ask it if it's conscious? I mean, it might add some weight, something, something to consider for the future.
I love that. We'd have the first certifiably not conscious LLM. Yeah, it's like, nope, there's no, there's no pre-loaded data here. Um, so that sounds really interesting though, because it's also useful, I'm sure for companies. If they're, they're experimenting and something gets in the training set they didn't want there, or, you know, potentially damaging, um, something, maybe even analogous on like, um, adversarial training data. You know, be a great way, yeah, to fix that. I mean, yeah, there's, there's lots of companies, for instance, they're really excited to upload all of their HR data to, uh, for internal use of, okay, well, yeah, how to, how to manage that data and like, obviously, it's textual data, so it would be so efficient and fantastic to be able to manage it via an LLM. Um, yeah, but that cannot start leaking private information about social security numbers, yes, like salaries and bank account numbers to the rest of the company, right? Yeah, it gives you a way to fix things that come up in, um, red teaming, perhaps, you know, like, you're red teaming your AI, you come up with something really devastating, well, just go in and lobotomize it, right? Right. It's possible. Yeah, it's a great approach. Love it.
And, um, are you doing any other, uh, new kind of Quantum machine learning, maybe approaches, or anything? I guess we end up staying on theme that way if there's anything like that on the horizon.
Okay, so all of this stuff that I've discussed so far is, um, Quantum-inspired, which is essentially, essentially classical. We've taken a lot of ideas from Quantum Computing on how to process the information, but our entire workflow still takes place on CPUs, on GPUs. Um, we do have some more forward-looking, maybe what I should say is that tensor networks for us is an ideal bridge until we get to a full Quantum era correction, etc. Um, so we've been using it heavily as an intermediate technology. Um, however, as Quantum Computing scientists, we, we know that at the end of the day, Quantum always wins, right? Once we have that full, at scale, error correction, then that's always going to be able to do stuff that tensor networks just is not able to do. Um, and so we've still got, uh, a big chunk of our company that's working on more forward-looking applications, uh, pure Quantum machine learning, uh, applications. Here, we're targeting mainly machine learning, so not neural networks. We've, we've done some applications on, uh, variational neural networks and things like that, but mainly because of the explainability piece, uh, we've mainly been, uh, targeting, um, regressors and classifiers, and just doing these things way more efficiently than classical computing could when we have that at scale, error-corrected. Um, another big thing that we're working on at the moment is how do we translate those gains from tensor networks onto quantum computers when those, uh, become like competitive, basically. And there's some extremely interesting things that you can do, especially in the realm of, uh, encoding classical information to a QPU, which is obviously an, a big open problem, big open and limiting problem in, in Quantum machine learning. And tensor networks can really help you with that.
It feels like a bidirectional possibility, but it's challenging, right? Like, if you were to take a Quantum circuit, you could turn into a tensor network 100% of the time, right? You can always do that. And then you have to worry about contracting it and all that good stuff if it becomes very large. But how do you take just a tensor network and then turn it into a bunch of gates? And that, that's that's tricky. Yeah. Yeah. That's hopefully there's something there that that makes it possible that I'm missing. But yeah, it sounds, sounds pretty challenging.
Oh, that's exciting stuff. So tricky. There's recipes to do it, but it's, it's definitely tricky. And, and the great thing is, okay, well, if you're able to encode your data via tensor networks, and then you have an efficient recipe to translate that to gates, then now you have an efficient and optimal recipe for encoding your data to a quantum memory, for instance.
It's like a direct path forward to whatever you manage to accomplish in the tensorization of, let's say, a neural network, can just be moved over. And I know you have lots of experience, uh, like two years ago, you were doing Forex trading, right, with, uh, Quantum machine learning. And so I'm imagining you guys aren't just sitting around, like you said, you know, things are happening. You're informed. And that incredibly exciting work is, is still going forward. And now we, we have a model that is competitive with, with some real, like, state-of-the-art. And an actual, this is exciting, it's an actual fully Quantum machine learning application that, that does all the training on an actual quantum computer. And that's competitive with, uh, state-of-the-art Forex models. And we're actually, uh, working on getting it into production with one of our customers. So, really exciting stuff.
Yeah, yeah. Full transparency, like I mentioned, real early, if people were listening closely, we have worked on projects, work together, and we are partners technically. So, yeah. So, so I'm pretty informed on these things, but still, there were a couple of things that you kind of surprised me with today, so that was fun. Uh, yeah, so I hope we get to work on some of this again in the future. That'd be really, really amazing for someone who wants to bring this to, or even pre-production, but, uh, kind of a close level. So, yeah, Sam, thanks so much for coming back on and sharing some of this. I, I, I think, uh, listeners are going to be excited to see how this progresses. Really great stuff here.
Thanks so much for having me.
Now it's time for Coherence, the Quantum Executive Summary, where I take a moment to highlight some of the business impacts we discussed today, in case things got too nerdy at times. Let's recap. Large language models face challenges related to their size, including hosting, memory footprint, energy consumption, and hardware costs. Using LLMs in a business also brings privacy concerns, as they often require uploading sensitive information to the cloud. Multiverse is looking to solve both of these issues with Compactify. This Quantum-inspired approach uses tensor networks that compress large language models, resulting in faster training and inference speeds and reduced energy consumption. Even though you end up with models around 75% smaller, you don't sacrifice much accuracy. And because this makes it possible to work with smaller LLMs in-house, privacy concerns go away, assuming you have your security buttoned up. This concept of running LLMs on smaller devices has Multiverse examining putting AI on removable drives that customers can connect to hardware. If you want to upgrade the AI, swap out the drive like you would with a thumb drive. Lots of AI lawsuits are making the news with artists and news outlets claiming their work was used without permission to train LLMs. Multiverse is working on something called Lobotomy that aims to remove information from models to address copyright issues selectively. Further, tensor networks make it possible to better understand features, allowing for improved explainability of what data influences a credit decision, for example. This could help spot and remove bias in models in the future. Of course, Multiverse still works with actual quantum computers. Recently, they built a fully Quantum ML application that is competitive with classical Forex trading models. Sam and I both remain positive that Quantum will always win in certain use cases when we have fault-tolerant quantum computers.
That does it for this episode. Thanks to Sam Mell for joining to discuss Compactify, and thank you for listening. If you enjoyed the show, please subscribe to Privity, The Post Quantum World, and maybe leave a review to help others find us. Be sure to follow me on all socials at Constant Hacker. That's Constant with a K, Hacker. You'll find links there to what we're doing in Quantum Computing Services at Privity. You can also DM me questions or suggestions for what you'd like to hear on the show. For more information on our Quantum Services, check out Privity.com or follow Privity Tech on Twitter and LinkedIn. Until next time, be kind and stay Quantum Curious.