Transcription
This week, as per usual, we have a variety of AI releases, but we got the one thing that everybody has been secretly desiring for a while: a ChatGPT that sounds more human and less like a robot. Beyond that, it seems like it was LLM week because we have updates from Mistral, Gemini, and Perplexity, which opened up a brand new category of AI products. So we'll be looking at that too, and so much more in this week's episode of AI News You Can Use—the show that looks at all of this week's AI releases that you can actually put to work today. No drama, no hype, just practical tools in a relaxed atmosphere. Let's get into it!
First up, we'll be having a look at the brand new version of ChatGPT, and I'm very excited to look at this because this has been my main gripe with ChatGPT: that the voice of it, the sound, the style, is a bit too robotic, not human enough. Other tools like Claude, frankly, did this better. That's why I have both subscriptions, and I know a lot of people who preferred Claude because of this new update that’s supposed to fix it.
So let's have a look, and I'll just go into a brand new chat here in GPT-4. Again, this is just a first look, but let me start this out with a classic: "Write me an essay about penguins," my go-to prompt that I've run thousands of times across all different types of LLMs. Let me just get a feel for how this performs with writing.
Okay, actually very similar to the previous version on the first look here. I do notice less typical ChatGPT words. Let’s just have a quick search; does it want to delve? No, it doesn't want to delve here at all. It’s good! These sentence structures are a bit different, which I like.
So, let's do something like, "Let it write copy." All right, my brand new "How to Use Chat to Make 10K a Month." Of course, copyright for the website. Let's go! The copywriting is something that it was very bad at right out of the gate; you had to do some intricate prompting to get something that's even close to usable.
Okay, that's pretty good. That's decent! Never seen it use bullet points like this before, with the green check marks. Wow, look at that! I'm looking over this and I don't see these typical ChatGPT words. I mean, we were sharing custom instructions inside of the community, teaching you how to remove a list of specific words just because they were so ChatGPT. You don't have to do this anymore from what I can see here. Also, it does sound a bit more human and less robotic, doesn't it?
Let me just run this inside of Claude for comparison. 3.5 sounded right here. It tells me a "get rich quick" approach could mislead people. Yeah, fair enough. Okay, if I say "do it," it just does it, which is good, but it doesn't really do the topic that I wanted; it changes it up. I mean, this is a big limitation with Claude. What can I say? It says that this is not a get-rich-quick scheme, although I explicitly asked for that. But yeah, the voice of this is good; this sounds way more human than ChatGPT used to. But honestly, with this new version, this is really good.
Okay, one more prompt here, and although this does not qualify as creative writing, I really care most about the readability. Again, this is one of my go-to prompts for style testing. Okay, this description is also different than it used to be; it's more elaborate here, giving two examples, and this voice is just quite human, I gotta say. I was wondering if it could arrange for it to be repaired or possibly replaced soon. This is very natural, not robotic at all as it used to be, I would say. Even this concluding statement is way better than what we saw.
So my first impression here is this is actually really impressive. I mean, be honest, look at these results and tell me if you would immediately identify this email as ChatGPT written. I would say that I've become quite good at detecting exactly how ChatGPT writes, and this breaks the pattern. This doesn't just break the pattern and remove all the obvious words, but it just feels better; it flows a bit better, I would say. Again, just a first look here, but to me personally, this is impressive and something I've been looking for. This is the main reason I've been using Claude a lot next to code generation.
So this is available through the API too. One other ChatGPT update this week that's shipping here is the fact that you're supposed to be able to use the advanced voice mode on desktop now. It hasn't arrived here, nor has it arrived on the US accounts for some of our team members. So apparently, this is something that is on its way and rolling out slowly but surely. Advanced voice is just available everywhere, from your phone to desktop. But I think that's a minor announcement compared to the fact that ChatGPT finally doesn't sound like a robot anymore.
Huge! And the competition is not sleeping, so we have some minor updates here from Google's Gemini. They added the memories feature, as you might know it already from ChatGPT. This seems to be very popular, especially with newer users that do not craft their own custom instructions. It just automatically picks up on what you tell the chatbot and saves that to its memory.
The result is more custom conversations in the future. Now you have this inside of Gemini 2, but as we talked about before, even though the LLM might be closest to what ChatGPT does, maybe inferior in certain use cases and categories, the overall tooling of something like ChatGPT just beats Gemini on all dimensions.
So it's nice that they're catching up on some of these features, but they'll really need to do more to dethrone the king. Also, while in Google's AI Studio, you can see there's a brand new Gemini experimental model that came out over the course of the last week. It is best in class and usable in their AI Studio now, and this new model actually ranks as number one in chatbots and LLMs arena right now, ahead of the new GPT-4 release.
I don't know how much weight I would give this, though. You can try it out in here. I personally would still prefer ChatGPT as of now.
So in these videos, we try to keep it lighthearted and entertaining, but at the end of the day, the whole point of us making this is showing you tools and techniques that you can actually use in your everyday lives or at least be aware of the fact that they exist. My goal is always for you not just to watch the video and say, "Hey, good video," and move on. No, ideally, I would like you to interact with the tools in the video, click the links in the descriptions, or leave comments on what it might have inspired you to do because at the end of the day, if you're just consuming, it's very difficult to improve and grow.
That might be fine if you're looking for entertainment, but if you're looking to learn something, interacting with the material is essential. And that's why we love partnering with Brilliant as a sponsor to bring you videos like this because Brilliant is an education platform that brings you hands-on interactive lessons. They have dozens of courses with hundreds of lessons across data science, programming, and more. As I mentioned, only by going hands-on and failing in the process will you really internalize some of these concepts.
So let me just briefly show you one of the courses that are Brilliant. This one was actually made in collaboration with Quirky, a popular YouTube channel that you might know for making topics like history or science easily consumable with their super high-quality YouTube videos on them. The great thing about this course is that it has the ability to go way more in-depth than usually YouTube videos allow for, and it's exactly the thing I sometimes wish every educational channel on YouTube had.
So if you want to check out this course or hundreds of other great options on Brilliant for free, check out the link at the top of the description. If you decide to stick with it, you'll get 20% off an annual subscription. Thanks again to Brilliant for sponsoring this video, and now on to the next piece of AI news you can use.
And another quick update from Google is that they're using a brand new family of models called Learn LLM. These are the models in the background of a lot of their applications, and I personally am a big fan of this. I just wanted to point this out because they're different than typical large language models. As they highlight in this article, if you ask this to create a travel itinerary for you, it won't do that because it doesn't even have the data in the large language model to build travel itineraries.
These models are trained on a dataset that is specific to education, so it can do specific things very well. Like if you're watching a YouTube video, you can ask follow-up questions based on the video, and then this model will be able to give you higher quality answers faster, thereby improving your learning experience. They said that this is not just coming to YouTube and Google Search, Gemini, and the Android operating system, but is already being used inside of Notebook and the new LearnAbout application that we covered recently on this show.
I just thought this was interesting because it’s a specific model for a specific purpose, and a lot of the applications are going to be using this and not a general-purpose model like ChatGPT. This is meant to be a universal assistant, but when you're learning something or building something for the educational space, you might want to lean into this direction.
Also, there's a brand new story that came in while we were editing this, and that is a brand new reasoning model. Remember 01 Preview and 01 Mini? This was something that only OpenAI had, and now a Chinese company actually released DeepThink R1 Lite, a brand new reasoning model that actually performs better than 01 Preview on several benchmarks. You could see them in here; I will link these docs below.
But most interestingly, this thing can be tried out today, and they’re both releasing an open source and an API is coming, meaning we just got an open-source 01 Preview model that you can try right in here. Just make sure to enable this DeepThink; you get 50 messages per day, and you don't even need a paid account here!
That's just for a classic logic problem. Add it and see how it performs compared to OpenAI's 01. Shall we? So you can see right away this is a bit different; it shows the entire thinking process. I don't dare to judge the differences between this and 01 Preview from this very first look, but it does give you the conclusion that from the free doors in this game show problem, you should switch to the remaining ones.
01 Preview says the same; this is probably in its dataset. But it's interesting to see that here you get the entire thinking process fully revealed, whereas in 01, you only see this one step. Also, this took 4 seconds and this took 32. But there you go—if you're not on a ChatGPT Premium plan, I would recommend you go check this out. You can just log in with a Google account and try this out. We'll keep an eye on how this develops, but this is very impressive, and anybody could be using this now, either here or in their projects.
So this is also a quick feature, but this is Jonathan from Grok with Q, and they built custom hardware that runs large language models at incredible speeds. He shows off a demo here comparing how quickly a small Llama model, Llama 8B, ran three months back at 750 tokens per second, and now they can run Llama 7B models, almost 10 times the size, at 3,200 tokens per second.
So those are just numbers, but check out the little video that he shared here. “Hi, I'm going to Atlanta for Supercomputing this year! I'm going to be there an extra day. Can you come up with an itinerary? Can you put that in a table? Can you add a duration column?” I know that was easy to miss, but here you can see the speed—it basically just appeared! And this is what you can expect in the future for use cases where it makes sense; the LLMs are good enough, and with hardware like this, the results just appear out of nowhere. Pretty amazing!
Okay, next up, we have some updates to Mistral AI's chat application, which is their direct ChatGPT competitor. Again, as many competitors, they're just catching up, but this does have some upsides, the main one being that it's freely available. You can just go to chat.mol and you can access their version of ChatGPT, which now can do a few added things. It can generate images, it can search the web, and they have a canvas feature that works just like ChatGPT.
They also now have PDF uploads, which is an absolute must with LLMs—such a useful feature. So as you can see, they caught up on a lot of this tooling, and they also released Pixel Large, which is their image recognition model, which has, as they say, a state-of-the-art vision model which has also been integrated into the chat.
But long story short, as with everything else inside of this application, this is just a stripped-down version of ChatGPT, and hey, that is to be expected. But nevertheless, I think it's lovely that they're shipping. Here, as you can see, you can make edits inside of here and even re-prompt just like in ChatGPT. You do like these presets, but it can do things like access the web with citations. Here it's quite simplistic compared to something like Perplexity, but there you go, it can do it now!
And I want to actually point out one positive thing here, which is I threw a few prompts at it and it actually performed really, really well—better than ChatGPT, I would say, on one thing, which was personal budgeting. Look at that! I just asked it for a personal expense categorization system for a single guy, 30 years old, living abroad by himself, and then it didn't just create categories and split them up into fixed and variable expenses but also created estimates for all of it.
I thought this was quite comprehensive and gave me tips for success in the end that I, as a person in this exact situation, found quite helpful. If you throw the same prompt at ChatGPT, you do not get these estimates and these tables down here, and no extra tips for success. Great success!
So at the end of the day, none of these is better; they're just different. And that's how it is with LLMs these days. All of these bigger players in the game have really solid products that behave differently in various scenarios, and that's why I always say, I think at this point for consumers, it really is about the tooling. What can this thing do? Does it have the canvas? Does it have the prompt presets, custom instructions, web search, image generation, code interpreter, web app, advanced voice modes?
The answer to all of those used to be no; now the answer to most of those is yes, but you can get a lot more free usage out of this compared to ChatGPT, so that may be useful.
On to the next one, which is a Perplexity update, and I think this is a really meaningful one because they're the first player in the AI space entering a brand new category: online shopping. They call this the AI-powered shopping assistant, and I think this is significant because one of the main use cases for Perplexity that it really, really shines at is product reviews and finding the right products.
It looks like something like Reddit comment sections and different blog posts, comparison websites, and actually reads the articles, whereas Google would just give you the links to the articles. It is probably the quickest way to decide on product purchases, and ahead of Black Friday, here they released this new shopping experience. Essentially, what this means is that there are new fields at the bottom of your product searches where you can directly purchase the products.
So one clear limitation is that as of right now, this is only available in the US. I suppose this will stay the same for the US, but some of the searches we ran from a US account clearly show you how this works. So you get the normal Perplexity result at the top, but here at the bottom you have related products where you can directly visit the site or, in some cases, even buy right here.
Now, this brings up a lot of questions because who curates this and how objective are these opinions? Because if I look at Reddit comments, I know that those are individual opinions. At the very least, I can be sure of the fact that there's individuals who will not filter what they actually think, so if there are serious problems with the product, they will appear there.
But this seems to just serve up certain products as if they were the go-to choice. We did the search for "Rocket League merch" right here, and it gives you some mystery packs and some random t-shirts. I'm sure there are many other products out there that might be preferable, but they just pick these ahead of everything else, which I guess just is what it is.
Reading through the launch blog post, one thing had me worried, and that is the fact that down here in their Perplexity Merchant program, they outlined that hey, if you sign up with them and partner with them to become a Perplexity Merchant, there will be an increased chance of being a recommended product. They even put this as the first point here. So I don't know, I guess time will show, but essentially, I've been using Perplexity to look at products to get unbiased reviews, whereas now they're not just looking at what people are saying online but they're actually making the decisions on what to serve you up and what products to recommend.
So I think this is good and useful, but I probably wouldn't make my shopping decisions directly off of here. I would take this as a recommendation and still proceed to watch YouTube reviews of trusted sources before I make my buying decisions, especially on larger items.
And lastly, I wanted to have a look at Sunno here with you because they came out with a brand new version, Version 4, and I reactivated my subscription here just to test it out. So this is a music generator, and this is supposed to be way better than Free 5, which was already really, really good. Huh, and interesting—you can actually now remaster old tracks! So if you have previous tracks that you really liked, you can now recreate them.
Let's try this one. [Music] Okay, so this was V3; let's remaster this one. Wow, that was super fast! Okay, let's see how this sounds. "Advant Curiosity Leads..."
Okay, clearly better! There's just more clarity in the voice of "Advant Curiosity." Yeah, the voice is clearly better; it is noticeable. Let's look at some popular examples here on their site. [Music]
Okay, that's an arrangement that goes beyond what was possible up until now. Nice! I really like this. What if you round out this week's episode with a custom song for AI News You Can Use? With a little bit of custom prompting and quick lyrics that I wrote myself right here, let's see what we get. [Music]
Here we go!
Okay! There are news, then you have AI news, but here—only here—you get your AI news you can use! That you can actually use! Abuse! For AI news! Better news! Then you have AI news but here—only here—you get your AI news you can use! AI tools that you can actually use!
Okay, so after a bit of iteration, here's my final version. I landed a Gregorian chant and some custom lyrics up here. Let's have a first listen together! How about this?
How about [Music] that?
In the mainstream, you have classic news and that is quite all right! On the web, you can find your AI news and that can make a weekday bright! But here—only here—you get your AI news you can use with AI tools that you can actually use!
This is AI news you can use.
Okay, that's impressive! It was so much fun to create, and that's pretty much all I got for this week. Make sure to check out the five educational live streams that I'll be holding next week. Hope to see you there, and have a wonderful day today! [Music]