📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI Platforms That Respect Privacy

Naomi Brockwell TV23:08

Transcription

AI is changing the world—fast. It’s already deeply integrated into our daily digital activities, whether we realize it is or not, and it’s here to stay.

Google says it's launching its own artificial intelligence-powered Chatbot to rival ChatGPT. Perplexity has eliminated 90% of Google searches for me because it is an answer engine versus a recommendation engine. Chatbots like ChatGPT, Bard, Perplexity, Claude. They’re incredible tools that can skyrocket our productivity, and give us abilities that were previously well out of our reach. We can easily generate images and videos, write articles and legal documents, instantly get advice or answers to our questions, even generate code.

But like all tech, it can be a double-edged sword, and there are some serious privacy concerns to consider. Open AI gathered and shared private information without consent. Bing's Chatbot sends your browsing info to Bing for possible use in ads. So the question is, should we embrace this technology, or be concerned about our privacy? Why not both? We don’t need to throw out new tech just because we care about privacy. It’s possible to embrace new tools, and protect ourselves at the same time.

In this video, I talk to two experts working on the privacy side of AI. We discuss the privacy risks of using AI, and how to protect ourselves, exploring some of the most private platforms and looking at best practices. Let’s start by understanding some of the privacy risks of using many of the most popular AI Chatbots.

First, many of them train their models on the questions and responses you generate. For something like ChatGPT, uh, OpenAI, they are sourcing a huge amount of data, uh, for their models through their users. Ryan Condron is a developer working on privacy tooling in the AI space. What they do with the user data is they'll, they'll store it and then they’ll use it to train their next models. Most AI platforms are pretty transparent that this is happening, and that it’s to improve performance and enhance capabilities. But many people haven’t quite internalized what this means. Essentially, nothing you type into them is private. It’s all cataloged, analyzed, and added to your profile. They say that they’ll try to anonymize it, but if you just say like two things about yourself that are somewhat unique, then it becomes, you can become very targeted very easily.

Brian Bondy is the co-founder and CTO of Brave, the privacy-focused browser and search engine. You might not realize it, but this data becomes part of their training set, and it’s, there’s no easy way to, to remove data once it’s trained on that data. AI Chatbots are complex and ever-evolving tools that are constantly being trained and updated based on our interactions with them. They’re learning from our reactions, they’re learning from our written responses to them, and they’re also learning from every piece of data we feed them. So let’s say we feed it names, addresses, medical information, relationship details: Once a model is trained on this data, there’s no simple way to ensure the information doesn’t show up as an answer to someone else’s prompt. It’s not just like you can just issue a takedown request and it’s done. There’s no way to reverse this, when it’s already encoded into that, that model.

Another privacy risk is that most of these companies keep a giant database filled with the things you’ve asked these Chatbots. This likely includes very personal questions and information. Matt Green, prominent cryptographer and privacy advocate, paints a picture of how personal our AI interactions are going to look within a few short years, that LLMs will be listening to your conversations and phone calls, reading your texts and emails. Your phone will be able to answer questions about what you did recently, what you talked about, what restaurants you went to. The more capable these tools become, the more access to data they’ll require. And people will willingly grant it. It has all the information about you that you’ve ever shared with an AI over years, it will know every personal, intimate detail of your life.

Erik Voorhees is a tech entrepreneur focused on decentralization and privacy. The ability for damage from that knowledge in one centralized place that a person can never get back once they’ve given it away, uh, I think is unacceptable. A central storehouse like this is going to be a huge target for hackers. If they gain access, this information could be used for extortion, to target individuals in other harmful ways, or might even be leaked in bulk on the internet or sold to the highest bidder. Most people are not scared enough.

This database will also be incredibly valuable for data brokers, because it provides such deep insight into your life, neatly tied together in a single profile that you’ve likely linked to your real name and phone number. The platforms collecting your data will absolutely be tempted to sell access to it. In fact, we should presume that companies are already sharing this information because their privacy policies tell us as much. Most explicitly state they’re allowed to share your data with external entities, which could include advertisers, partners, and law enforcement.

But a really important consideration about this data is that it will become irresistible for governments. As Matt Green asks, “what do governments and law enforcement do when they realize every single person has a little agent on their phone that can answer literally any question about their activities, in simple human language? The “crypto wars” are going to look quaint.” We can expect to see waves of subpoenas or even blanket national security letters aimed at these companies, where your private searches are not only AVAILABLE to law enforcement, but used against you in ways you hadn’t anticipated or taken out of context, or they might even be used in court cases where they’d become part of the public record. I personally don’t want all of my conversations within AI to be stored forever by a central party to be offered up to the government as soon as they ask for it. As this data gets warehoused and centralized, the incentives to abuse it will grow and grow and grow.

Given all of these privacy risks, how do we navigate this world of AI safely? Some of you will say, “Just don’t use these tools in the first place”. That is one solution, and it’s completely up to you. But for me, I love exploring new technology and seeing if I can use it to improve my life. And I like teaching people how they can still live a privacy-conscious lifestyle without throwing out their digital devices.

So if you’re someone who wants to embrace AI tools, what is the most private way to do so? I’d say the best option is to host one of these AI models yourself. What does this mean? Well, companies are training large language models, or LLMs, using vast amounts of data and computational power, but the end result is a model you can query at any time. Many of these models are now efficient enough that you can host them locally on your own machine. This eliminates both of the major risks we just talked about: Your data isn’t being sent off to external servers for further training—it’s all staying on your own computer, so you can feel comfortable giving it personal information. The information you give it also doesn’t get stored in any company’s central database, so it’s not susceptible to data breaches, selling, or subpoenas the way that major hubs are. We’re going to explain how to set up one of these models locally on your computer in our next video in this series.

If you don’t want to host one of these yourselves, there are privacy-focused platforms that don’t collect your information. I think one of the best comes from Brave, which is a great privacy-focused browser and search engine. They’ve created an AI tool called Leo. Brave Leo is a chat assistant that is built natively in the Brave browser, which means you can access it without installing any extensions or apps. When you use their search engine, Leo will instantly summarize results for you to include answers for your search queries. There’s also a Chatbot interface. The current way that you access it is via the sidebar. You’ll see a little icon for Leo. You click that and you can communicate basically with the current webpage that you’re on. Whether that’s like a Google Doc, a PDF, um, whether that’s a Brave Talk meeting, if you’re on a GitHub page, it understands how to get pull requests, things like that.

Brian spearheaded Leo development, and he explained how their privacy-focused AI tool can create real-time summaries of videos, generate new content, translate pages, rewrite them, all kinds of things. It can either get context from the webpage, um, or it can also use Brave Search to get more recent information, for example, or for accuracy reasons. You don’t actually do the search as the user, but you’ll see a little message that says, “enhancing your interaction with Brave Search” within Leo. But what makes Leo great is its focus on privacy. There’s no logging. We’re not building a profile of a user over time, we’re not storing anything. As soon as it hits the server, we generate the reply, we send the reply, and then we discard. There’s no state saved on the server whatsoever. So you have a conversation with Leo, but as soon as you close your browser session, the entire conversation history is deleted, and Brave doesn’t keep a copy of anything you asked. So you can have like a coherent conversation basically, and ask follow-ups and things like that, but as soon as you leave that page or close that tab, everything’s gone. You don’t even need to create an account to use it. There’s no login required so you can just, uh, open the browser and start using it. Leo, it’s not linked to your account at all.

On the backend, Leo uses an open-source AI model called Mixtral, made by one of the leaders in the open-source AI community, Mistral AI. These self-hosted models can be downloaded and run by anyone, but by default Brave hosts it for you on their servers. Now because Brave is just hosting pre-trained open-source models for you, this means the data you put into them doesn’t end up as an answer to anyone else’s prompt. Brave isn’t doing any training at all themselves. So this is a big privacy plus. We already mentioned that Brave doesn’t log any of the queries you send. But on top of that, they also pass your query through a proxy server first instead of sending it directly to the server hosting the model, where it strips out your IP address. We use a reverse proxy. So whenever you do a query, it drops the IP address and then it sends it after that as a second step to the, uh, LLM. Basically, we’re just dropping the IP address so that any request looks the same as anyone else’s request. This means that your queries aren’t tied back to you.

Mixtral is just one of the open-source models hosted by Brave that you can use. Brave offers like a bunch of different models you can select between. So we use Mixtral by default, and you can use Llama 3, uh, 8 billion or 70 billion. And as long as you choose one of the options hosted by Brave, Brave is processing that request and not collecting any of your data. An upside of having Brave host these for you instead of hosting locally is that you can run more powerful models than your own personal computer might be able to handle. You also have the option to link Leo to powerful models that are not hosted by Brave. You can even select to use an external provider like Anthropic. So all the Anthropic models are on there as well. Keep in mind that if you choose to use a model hosted elsewhere, that 3rd party hosting it won’t have the same “no logging” policy as Brave, so your data can potentially be collected by the third party. But there are still privacy perks of using Leo to access these other models. The benefit is that whether you’re sending to Anthropic or to Brave, it’ll go through that, that reverse proxy first. It’ll then strip up the IP from the query. So that when there’s any handling of the payload of the message that you’re asking whatsoever, there’s no way at that point to, to link it.

So there’s Brave-hosted models, there’s 3rd party hosted models, and Brave will also allow you to use Leo with self-hosted models. Bring your own models that’ll allow you to bring up a local LLM on your machine. It’ll keep those interactions completely, not talking to Brave whatsoever. So Brave doesn’t act as an intermediary, it’ll just use those local LLMs. This means that nothing leaves your device and you retain full control over your data, but you still get to use Leo and its capabilities. You running on your own machine, I would say is one step even better, um, than Brave hosting it, because at that point you don’t need to trust Brave, it’s a little bit of setup, so it’s not for everyone. Um, but if you want that like a hundred percent guarantee that nothing is ever happening with your data, then the best way is to run it locally. And that’s, that’s why we’re bringing that option to users as well. In our ideal setup, we, we would rather not host the models ourselves, and we’d rather just be local, but even as it is now, um, you can have safe interactions even with the Brave hosted models, you can host them yourselves. So, you can definitely have both privacy and AI.

My takeaway with Leo is that if you trust Brave as a company to not log your search results, you can trust Brave-hosted models very comfortably with your AI Chatbot interactions. I personally have high trust in them, but you’ll have to make your own call. I think Leo is a great option if you want to try out some of these AI tools while protecting your privacy.

Another interesting platform is Venice.ai. Eric Voorhees, uh, had this idea of taking open-source models, and, uh, building a ChatGPT competitive UX on top of them, and allowing users to create images in real time and, uh, text generation in real time using uncensored models. Venice is an AI chat app; Erik is the founder of Venice.ai. It gives them like most of the benefits of what the big firms are doing, but without the censorship. Easy-to-use, censorship-resistant tooling is the main value proposition of Venice, and we’ll look more closely at it in a future video when we look at decentralized AI solutions. But Venice also makes efforts to protect user privacy. The way to protect people’s privacy is to not take their information in the first place. So the way Venice works is, when you make a request, Venice basically acts as a passthrough, connecting you with the computers hosting the LLM models that will process your request. Your data passes through Venice to the, the backend compute provider. It never gets captured on the Venice servers. It never gets, you know, compiled into data stores. When someone sends in a prompt, text or image or code generation that’s going through an encrypted proxy server to a decentralized GPU that does the processing and sends the result back through the same encrypted proxy server, it’s never held on Venice’s servers at all. So that’s what happens with your data on Venice’s front end, and we’ll talk about what happens to your data on the processing side in a moment.

Now a good thing about Venice is that because they’re using open-source self-hosted models, your queries don’t end up as part of the trained data set. That’s kind of the, their core ethos, and that’s why they started was they didn’t want a ChatGPT being this big harvester of user data and using people’s personal data to train the models. Unlike many other platforms that require you to create an account and link your cell number, on Venice, they don’t even need to sign up with an account. But if they WANT to sign up, they can just use like a dummy email if they want. Venice never knows your name or address or other personally identifiable information. Even if a user pays to upgrade their account, they can do so privately. If someone wants to pay crypto, they can do that too. So it doesn’t matter where they’re from in the world. We don’t want to know who they are. We don’t need to know who they are. We give people privacy not through some like 15-page privacy policy that promises how we will respect their data. We just don’t take it in the first place.

Venice is a great UX that recognizes that privacy is a value proposition they can offer users, and that they can protect users by not collecting their data. It’s important to note that Venice is not processing your queries themselves though. Think of them like a super easy-to-use front end that connects you to a network of compute providers on the backend. They use, uh, device rentals in the, in the background like Akash, uh, to run their compute and model hosting. Akash is a marketplace for selling your compute resources. It’s kind of like Airbnb for computer processing power. It’s essentially a GPU marketplace. So if I have a, a computer with several GPUs that I’m not using, I can put them up for rent on Akash and someone can rent them for a certain amount of time. So these people renting their GPUs are the ones hosting the AI models and running your queries through them. As well as using Akash, Venice uses other computer providers like Hyperbolic, and they are in the process of moving to the Morpheus network which will completely decentralize the processing of queries, taking a big step forward when it comes to censorship resistance in the AI space.

What you should be aware of from a privacy perspective is that your queries are seen by whoever processes your requests on the back end, and you don’t necessarily know who those people are. The provider themselves do have direct access to the, the data coming in. We can have encrypted data go from the user to the provider, but at some point, the provider has to decrypt the data to feed it into the model. Essentially, this data needs to be visible to the provider in order for it to be processed by the model they’re hosting. Even if you feed encrypted data into a model and have the model decrypt it, technically the provider does have access to the decryption key because the model is residing on their local machine. That’s the last piece of the puzzle we’re working on right now; there’s actually a couple of groups working on, um, anonymous inference. Uh, it’s, it’s actually a really hard, uh, problem to solve.

So in general, I love the un-censorability of Venice, and that Venice allows anonymous sign-up, uses self-hosted models, and doesn’t collect user data. The privacy isn’t yet perfect because the backend servers are an unknown entity, but truly private inference is an unsolved problem. While Brave has a no-logging policy with the queries they process, Venice’s decentralized providers will be making their own logging decisions. But in both cases, there’s an element of trust required to use both Brave-hosting and Venice’s decentralized hosting. So the most private option for AI Chatbots is still to host your own model. But Brave and Venice are 2 of the best options out there if you’re going to use the cloud, because your data isn’t used to train models, they strip IP addresses from queries, and they both take steps to protect you from any central repository collecting your information.

If the risk of logging is a concern for you, you CAN use these systems comfortably, you just need to employ some best practices. So let’s specifically outline some risks of each kind of AI Chatbot across all the privacy tiers, and the best practices for using each so you can mitigate privacy risks.

With self-hosted models, whether linked to Leo or not: I think you can feel comfortable asking sensitive questions and including identifiable information. The queries aren’t leaving your machine, so you can trust these self-hosted local models as much as you trust your own computer.

With Brave-hosted LLMs: They have great no-logging policies and a good reputation for privacy. So I feel comfortable asking sensitive questions and revealing personal information because I have high trust in them, but you will need to make your decision based on how much you trust Brave. (See server and computer with a package on them and a question mark) When I don’t know who is processing my queries on the backend, I remove any identifiers in the prompts such as name etc. and am careful not to identify myself within any individual session. But by using passthroughs like Venice, which stops a provider knowing which prompts come from me, and using decentralized providers on the backend, no single provider is getting all my queries across sessions, so I’m not as concerned about them being aggregated together to identify me. My concern is specifically about identifying myself through the information I provide in a single session.

When using centralized LLMs like ChatGPT that are collecting your data and linking it together in a single profile, I’m careful not to include sensitive information that I wouldn’t want known about me or shared with others. They already know my name, so it’s not a matter of them identifying me. It’s that I’m very aware that they’re keeping a record of all the questions I’ve asked, any of which might be taken out of context or used against me in the future. Obviously, your choices will be different; this is just how I do things, and as long as you’re aware of how each platform handles data, you can make an informed decision based on your own judgment.

People will come to learn that. Like they shouldn’t give up their information without at least thinking about it, right? Like sometimes it’s fine to share your data, but at least be thoughtful about it. It’s up to the individual to disclose information. We’ve entered a new era, where AI platforms scrape the internet for any public information they can find. This means that we need to be more thoughtful about the information we broadcast publicly in general, and realize that even what we think of as private interactions on Facebook or Google services like Gmail or Maps, should be considered public in this context. No one really thinks that all their posts are being used to train these AI, but they are. We’re posting everything about us, where we’re at, what we’re doing, what we’re eating, um, who we’re hanging out with. We’re using, uh, centralized services. We’re using, uh, software on our phones. We’re using social media and different networks that essentially harvest us for data. They harvest us for analytics that can be abused in so many different ways. Now that this data potentially is about who we fundamentally are as people, we need to be very careful with that information. I’m really excited that we have private options in AI because I wanna use this and I don’t necessarily wanna be just handing over my data to all of these companies. So it’s great that we have choices.

In our next videos in this series we’ll explain how to host models on your own machine. Especially in an age of AI, we need to be more mindful than ever of the consequences of leaking all of our data in online interactions. Luckily, there are plenty of ways we can better protect our privacy. Including by making privacy-conscious choices when we use AI tools themselves. Caring about privacy doesn’t mean we have to give up modern technology. Even privacy-conscious people can enjoy AI tools in their lives. Thanks for sticking around!

Now, I hate labels, but one term I’ve started using is priv/acc-- Privacy Accelerationist. Because I don’t think we can afford to wait for people to eventually wake up to the importance of privacy. We need to push for change now. I’ve adopted this label because there can be solidarity in communities. People all around us keep declaring they have nothing to hide, making the rest of us feel uncomfortable for valuing privacy. I’m just here to remind you that Privacy is normal. You’re not weird for valuing it. And the privacy community is much larger than you might realize. Thanks to all of you who have joined in signaling that you too are priv/accs. Let’s bring more people into this community! Also, this T-shirt is available to purchase in our store! Or you can find other ways to support our non-profit that advances privacy. Thanks so much for your support! She searched for the best cat videos. INTO GULAG! Can I haz tshirt? Yes you can!