Transcription
Another fantastic week for generative AI users! OpenAI made the GPT 4.5 model accessible to everyone. We'll take a closer look and compare it to a free alternative, Grok Free, which many are switching to.
There's also brand new voice assistance boasting unprecedented features like contextual awareness and incredibly human-sounding voices. A new PDF recognition API performs better at OCR β scanning images or PDFs β than anything previously available. All of these are innovations you can use to your advantage, and that's what we explore in this show. Every week, we highlight AI news you can use. Let's get into it!
Our first story is the ChatGPT release to all Plus users. This means it's no longer gated behind the $200 Pro Plan. All Plus users ($20/month) now have access to GPT 4.5. Last week, I reviewed the model and explored some of my favorite use cases: writing and thinking. Over the week, I tested it with psychological and marketing scenarios β things I regularly do with these models. I was very positively surprised.
The internet, and many commenters, disagreed. People questioned my positive opinion, suggesting I was a paid actor. This motivated me to dig deeper and compare it to other models, specifically Grok Free, a great free alternative to the $20/month GPT 4.5. I want to highlight some differences, or rather, similarities, because they're very similar.
Some creators initially negative on GPT 4.5 are reevaluating it. David Chappelle, for example, initially expressed strong negativity but is reconsidering. Matt Wolf also initially disliked it but now uses it for almost everything except coding. This is my exact opinion.
Grok is excellent; in many tests, it performed as well as GPT 4.5. If you're not paying for a subscription, that's incredible β a state-of-the-art non-coding model freely accessible. I recommend Grok Free if you're on a budget. For almost everything except coding, it performed equally well or better than other models. However, ChatGPT's tooling is far more advanced.
Let me show two examples illustrating this point. One is a simple ideation prompt for a blog post about AI alignment. The results from both ChatGPT and Grok are virtually identical, except Grok suggests a hook. Even the topics overlap (the black box dilemma, for example). The results are the same. For ideation, both are great; GPT 4.5 is just better. Most people prefer its tone; it's the best LLM for writing.
My second example, featured in a previous episode, uses your personality profile to create a CIA-style report. I provided extensive personal context. Both generated reports with psychological profiles, latent threats, and risks. They were virtually identical in structure and content, showing deep insight and empathy.
So why do I prefer ChatGPT? Grok's model is incredible, but its functionality β deep search and its "thinking" mode β is inferior. What matters are the projects I constantly use, the advanced voice assistant (used almost daily), and the overall workflow. For coding, I use Sonnet 3.7.
For prompts where results matter, I use multiple LLMs: GPT 4.5, Grok Free, Sonnet 3.7, and 01 Pro. For creative or psychological tasks, or when tone matters, I use GPT 4.5. But if I had $20 for one platform, it'd be ChatGPT. For a $0 budget, it's Grok Free; for coding, Sonnet 3.7.
Next, a non-typical AI news story: Mistel's new OCR technology. They claim it's higher quality than anything before. OCR (optical character recognition) converts images with text into usable text files. Their examples show it converting phone pictures of paper, Arabic text, and tables into computer-readable text. It outperforms GPT 4.0 and Gemini 2.0. It's accessible via Le Chat, their web interface.
I tested it by writing "If it can read this, it can read anything," then the same thing illegibly. Le Chat correctly transcribed both. GPT 4.0 failed; Gemini Advanced also failed. A simple test, but if you need OCR, use Le Chat's API, which handles bulk processing and multiple languages.
Next, Ideogram 2A. Ideogram is an image generation model considered best-in-class for text stickers and graphics. Its new model, optimized for graphic design and photography, surpasses expectations. We tested it; difficult images (like ballerinas in complex poses) worked well. It excels at displaying text and graphics. Billboard examples are among the best.
The downside? Detailed facial expressions and close-ups don't work well. But for graphical elements or text in images, it's the best choice. MJourney and Flux are also good, but Ideogram 2A is an easy recommendation, especially at 50% lower cost.
Next, text-to-speech (voice assistant) innovations. OpenAI leads with its advanced voice mode (though it interrupts often). ElevenLabs, Hume AI, and Sesame are making waves.
Hume AI's Octave text-to-speech understands what it's saying, using intonation and pacing based on content. It recognizes sarcasm, and the voice changes depending on the input description. It can even invent voices from scratch based on the script.
Sesame's voice quality is impressive. It sounds better than OpenAI's advanced voice mode, with smoother interruptions. It's a matter of time before all LLMs reach this level of quality.
Next, Claude's MCP (Model Context Protocol). It's a standardized protocol that lets you plug external services into Claude. You can pair your LLM with tools. I installed MCP on my computer to access the web and create directories. People are using it with Cursor and Sonnet 3.7 to build applications, giving agents access to local directories, databases, and internet searches. It's free with a Claude subscription. An event in my community teaches you how to set it up.
A quick segment on AI video: Luma AI and Pika Labs released transition features. PixVerse V4 has a redesigned interface and a new video model (though V2 is still best). OpenAI shared plans to integrate Sora functionality into ChatGPT and create a Sora-powered image generator. DALL-E is outdated. High-quality models like Flux are available. Sora has a user-friendly interface but inferior model compared to Chinese models and V2.
Haen, known for video avatars, released a feature using preset avatars to generate user-generated content, often used for ads. This makes it easier to create ad campaigns with AI-generated influencers.
This week's AI news you can use! I find GPT 4.5 and Grok Free valuable, along with Claude's coder and Deep Search. See you soon!