📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

НОВОСТИ ИИ: Прорыв в Агентных и Думающих ИИ

Продуктивный Совет24:02

Transcription

Deep research is one step closer to AGI. Minor but cool updates from OpenAI on how to train your own model for just $50.

Hello, this is Prod Sovet. My name is Uncle D, and I've gathered the most interesting and important news from the world of artificial intelligence for this week. To support the release of new videos, please subscribe to the channel, leave a like, and a comment below. Let's get started.

Deep research? Any good? No, this is still a Russian-speaking channel, but the company released Deep Research quite unexpectedly at the end of last week, on Sunday. They announced this release. We knew that something like this was in the works, so it wasn't a huge surprise. Nevertheless, let me say a few words about it.

A wonderful button appears for all Pro subscription users for $200. What can you do with this wonderful button? I'll briefly show you. Why briefly? Because a full review will come next week, where I'll explain how to use this feature properly and improperly.

Look at this massive research provided by the model. It all starts with you writing a prompt. I tried to make mine quite detailed. Then the model asks you several guiding questions. This doesn't always happen; sometimes there are more questions, sometimes fewer, and sometimes it doesn't ask any at all. After you answer them, the network goes online and gathers all the necessary information, in its opinion.

You can see the process of its reasoning and the resources used to compile the analysis. The reports are quite extensive; you can even see links to the sources from which the information was taken. You can really scroll through it for a long time. The question, of course, is the quality of these reports. So far, my conclusion, based on three or four uses of this model, is that the prompt is indeed important, and the quality of the reports can depend on what you ask the model and how you guide it before it starts generating everything.

But let's return to our page and see what users are saying. They say the following: "Deep research exceeds the value I get from a private researcher whom I would pay $150,000." Where did he find such a researcher, and why is it not me? But apparently, the music didn't play for long; now this user will pay $200, not to a researcher, but to Sam Altman.

Some research reports turn out to be incredibly extensive—30 pages, 10,600 words—based on a rather simple and not very detailed prompt about the history of the development of certain games. As I noted, this is currently only available to Pro users, but OpenAI promised to release it for ChatGPT Plus. It's quite possible that there will be a trimmed version there.

Deep research works on the O3 model, not the mini. This is the only way to test and work with the O3 model. Perhaps on other plans, it will be cut down or use O3 mini in some variation or with a smaller computer. Well, we'll see. Expect a release about this next week.

OpenAI needs more money, more data, and more partnerships. The company is expanding its presence through a strategic partnership with South Korea's Kakao and Japan's SoftBank. SoftBank is investing in this joint venture, and they plan to deploy everything within the SoftBank company. Of course, OpenAI plans to gain some access to data for training its models and expanding its influence.

Since the Chinese are also not standing still, it's necessary to strategically penetrate these markets, which OpenAI is doing. While things are going well for OpenAI with people, the situation with robots is not as smooth. At least, the partnership with the robotics company is not going well. However, with Figure AI, a company for which OpenAI primarily supplied its models as software for these robots, contracts are being terminated, and partnerships are also being dissolved.

The CEO of BT EXF states that a major breakthrough has occurred in the company, citing the need for vertical integration. In the opinion of Figure AI, OpenAI is no longer needed. The company is currently valued at $6 billion. We'll see if they can maintain that valuation without OpenAI.

A short advertising integration from our friends at GP Tunnel. The sponsor of this episode is the GP Tunnel service, a convenient tool that allows you to interact with more than 100 advanced neural networks without VPNs or subscriptions. Payment is based on actual usage, starting from just 50 rubles. All current models are represented, including Gemini 2.0, O3 Mini, Deep Grog models for generating images, Midjourney FLX for generating sounds, Audio SUA, and more.

Recently, they introduced a creative lab, a kind of canvas where you can work with ready-made images or generate new ones. There’s an eraser, a brush, and you can even generate videos using Min Max in text-video or image-video formats. If you are a developer and want to officially work with a Russian company using the API, that opportunity is also available. You can create a Telegram bot or a microservice using the API.

For all new users, there’s a wonderful promo code for 100% bonus on your payment. Follow the link in the description of this video. GP Tunnel is a fast and convenient way to access top-notch AI tools.

There are almost official statements, or maybe not, from the company that OpenAI will create hardware devices, not just develop models but something wearable. Yes, it needs to replace the phone somehow, right? While traveling in the Asian regions, we were asked when we would release something. It was quite vague, but they said they would make something that you would really love.

I don't know what that could mean. If you have any insider information, please share it in the comments. The main point is that our interaction with devices should change because very smart autonomous agents will appear, and we need to adapt humanity and ourselves to this future.

So, I hope this will help. Thank you, OpenAI, for a few minor but pleasant updates released this week.

First of all, let's make search great again. Here’s a tweet from Altman: "We observe that search is available to everyone on all tiers." Congratulations, you can close or delete your bookmarks for Plexi. Or not, if you love Plexi. Open Search works well. I tested it some time ago, and it is indeed comparable to what Plexi shows in the UI.

Of course, there are differences, and it’s a matter of taste. But now, if you don’t have a subscription, there’s a great opportunity to go and see how it works.

Another pleasant news is that we now understand a bit more about how the O3 Mini models reason. The company still refuses to show us the entire raw chain of thought, all the raw reasoning the model does before giving you a final answer, but it’s starting to look a bit more like those raw thoughts.

By the way, I found a prompt that creates summaries. It’s gigantic, of incredible size. I didn’t read it entirely, but it generates summaries of all these thoughts and outputs what you see here now. I think that interacting with models whose reasoning you can read is indeed more engaging.

That’s why I prefer the PC R model more because you can see how it thinks, observe, and understand whether it arrives at the right conclusion or not during its response. You can even learn coding when it thinks and writes code for you.

Go to O3 Mini and enjoy the almost pristine reasoning of this model.

Another useful update is that you can now share your canvases. I actually made a post about this in Telegram and even created a canvas that I shared with all our subscribers. I’ll show it to you now.

This is somewhat similar to the artifacts that Claude had. You create them right in the app and then share them. It loads quite slowly, but we still get a nice, cute one-page code in React or HTML.

Now it can be rendered in these artifacts. This is specifically React, and here you see three news items with links to the corresponding sources that I made and shared. It’s a nice, cute use case for prototyping a one-pager or some kind of presentation to share your ideas in a beautiful visual interface.

So, this update is good, and it really feels like Claude is gradually yielding more and more features to the OpenAI model and, in general, to the ChatGPT service. If earlier we went to Claude, for example, for artifacts or because it coded well, now we have O3 Mini, which also does wonderfully, and you can use this model without any worries.

But there is one big "but." I still peeked at the arena, specifically the web arena, to see who is currently in first place for creating UI interfaces. Regardless of what models these wonderful and terrible Chinese and American companies are making, Claude still stands in first place, which is quite surprising.

If you often work and create interfaces or websites, it remains a good and reliable choice, and even the new models are not catching up. But it seems Anthropic still needs to start moving a bit and release some new models.

We’ll be waiting for them. The Ning format will be very interesting to see and touch, and I’ll tell you about it. It seems that employees or even founders of all these AI lab companies are in some kind of competition to see who can work in the most companies for the least amount of time.

For example, John Schulman, co-founder of OpenAI, left in August because OpenAI is no longer "it." You understand the world of Murati. We also know that she left OpenAI in September for the same reason. Then John Schulman left Anthropic and came to Murati. Apparently, Anthropic is no longer "it" either.

The question of where John Schulman will go next remains open because companies are running out, but you can still go to Sukevira. However, their number is not infinite, so we need to understand that Murati is negotiating to attract over $100 million to her company, and she has some serious minds from OpenAI and Google on board.

So, we are waiting for a product from Murati. Let’s see what she comes up with and if she will come up with anything at all. Well, investors have given money, so they need to recoup it.

Gemini 2 Pro has finally arrived. Google has come to us with this model. They have a blog post where they explained everything. Honestly, it’s already quite difficult to understand all these models and their names—F, Pro, and now Flashlight have appeared.

But the more dashes, hyphens, numbers, and dots in the name, the better, according to the company’s perspective. However, I think users are skeptical about this. Nevertheless, the context window is 2.0, F has 1 million tokens, and the context window 2.0 Pro has 1 million tokens.

This is wonderful. We see such numbers in the benchmarks. M79 and 1 are good, and GPQ Diamond handles complex PhD-level questions well.

There are many questions about this. I think everyone has many questions because, in reality, things are not as smooth and sweet when we test these models. There’s a feeling that Emin is just stylized better, gives better answers, and users prefer it more.

But I worked a bit with this Gemini 2.0 Pro, and the answers are indeed pleasant, well-structured, and follow the prompts wonderfully. I have no complaints, at least in tests with self-generated text.

What’s interesting, perhaps primarily, is the price-to-quality ratio. The quality is on par, but there are certain questions about coding. There may be Ning models that are better, referring to the ones from the Chinese and OpenAI, but the price is astonishingly low.

Not only is there a possibility to use these models for free, for example, in Google AI Studio, but the price is also quite low. Just before recording this video, I found a useful comparison. El Mara has now started releasing the price-to-score ratio.

As we can see, the score is very high for Gemini 2.0, and in terms of price, it is practically leading, being very cheap. Then we have DPS One, GPT 4, and Gemini Flashlight, which is worse in score but even cheaper.

So, it’s a very good solution, a great option for developers. Take a look and try it out. By the way, I’ve already added Gemini Flash Thinking and 2.0 to my cursor, and I will probably pay a few pennies for the API instead of paying for the cursor.

Look at this cute purple mascot or just some logo of Copilot looking at us. Agent Mode has appeared for Copilot in VS Code Insiders. I didn’t know this existed. If you are an Insiders user of VS Code, congratulations!

Agent Mode looks like this; it can autonomously perform tasks, fix errors, and suggest commands for the terminal, reducing the number of manual actions you take. The same feature is now available in Cursor and Relita. I tested Relita, but I haven’t specifically tested this feature in Cursor.

If you work with physical code and with Python, pay attention. Moreover, Microsoft introduced Copilot Pro, I don’t know how to pronounce it correctly, but it’s a fully autonomous agent that can identify problems in your code, take them into development, improve them, check that it hasn’t broken anything, and make requests.

You can then review everything and either accept or reject the request or ask for something to be refined by highlighting a piece of code. So, we are indeed moving towards everything being as autonomous as possible.

Don’t despair; life is beautiful! To make your life even more beautiful, join our community. The Prod Sovet community will help you move faster towards your goals by applying AI tools.

We focus primarily on development and project creation within the community. Many participants have their own businesses or ideas and projects. However, if you haven’t decided on your idea yet, the Prod Sovet community is the perfect place to come up with one.

Moreover, all participants gain access to educational materials on various topics that can be viewed and studied through subscription. Of course, there are regular calls, communication, and support. The link is in the description of this video.

But why? Here’s my question: Mistral is updating its branding. They added dashes to the letter M. It’s a very interesting branding update. But besides that, what else did they do?

They presented prices: $15 for a Pro subscription. I did a review on Mistral; you can search for it on the channel. It wasn’t too long ago. I was absolutely not sure that it was worth paying $15 for anything, but they probably want to emphasize that now.

Mistral generates up to 1,000 tokens per second in its client, which is called Chat. By the way, they also launched a mobile app.

How is all this done? Apparently, it works on CBR chips. The French decided to focus on speed rather than quality. If you want to support a European manufacturer, now you have all the opportunities for that. Or, let’s say, make your contribution to the development of open source, because Mistral is trying to release almost all its models with reasonable licenses in open source.

These are the news.

This is very interesting. I often come across various studies on Twitter, but some are complex, technical, and don’t carry much meaning or benefit for detailed study. However, this particular study from Stanford scientists even made it into the news.

What’s the idea? Let me quickly explain. The scientists took a model with 32 billion parameters, if I’m not mistaken, spent 30 minutes on 16 NVIDIA H100 GPUs, costing about $20 to rent, and trained a Resing model based on this model through distillation.

They took data from Gemini 2.0 Flash Thinking and had a dataset of a thousand very clear, correct, and apparently high-quality reasoning chains, which they used to teach this model. An interesting life hack: the scientists intervened in the model's reasoning process in a specific way.

When it seemed like the model was finishing its reasoning, they added a single word: "Wait." This way, the model continued its reasoning and generated tokens after that word towards further reasoning, achieving better results.

You can see on the graph how many times the word "Wait" was added at the end of the reasoning—twice, four times, six times—and how the benchmarks improved as a result. For example, on GP QA Diamond, it reached 60%. If you remember, Google had 64%.

In general, this is quite a breakthrough and interesting. The better quality data you have for training models and the more clever you are in using such tricks, the cheaper you can train models for just $50.

Of course, the base model was used, and we’re not talking about training and all those perks, but still, it’s impressive.

Now, MCP is available in Cursor and Win Surf. One of the founders of Cursor informs us about this. You need to download a separate application, WNF Next Beta, to use it.

By the way, Dorsey also released a goose, but I hope to talk about that somewhere, either in Telegram or on YouTube when I get the chance. The goose works entirely on MCP.

I won’t go into details. Now, you can connect various endpoints through this protocol. You can call them servers to which your model can refer to retrieve data or, conversely, output data.

In general, you can have a lot of fun and customize your applications and agent workflows. It might be a bit complicated for some, but it’s genuinely interesting and useful.

Although Cursor offers a lot of built-in tools, like internet search and database work, so it may not be super relevant. But if you want to dive deep, go ahead.

MCP is now available, and we want to know more.

We now have a bit more information. Ilya Ver is working on Super Intelligence, the company where John Schulman didn’t have time to work but will apparently negotiate funding with a valuation of $20 billion.

So, we have no product, not even a website—just a page with text—but it’s valued at $20 billion. Why? Because he co-founded OpenAI.

Well, forgive my cynicism, but that’s the news. Now we’re just observing indirectly that something is happening, or at least the company’s valuation is rising. I hope we’ll also see some product from Safe Super Intelligence.

Finally, Paul McCartney dug up an old demo recording of John Lennon and released a new hit. Well, actually, an old hit, but he processed it so that the voice sounds clear, high-quality, and pleasant.

Yes, he even won a Grammy for it. The song "Now And Then" by The Beatles was released in 2023, and just today or yesterday, you can see how it is.

It’s important to adapt to modern conditions and technologies, just as it was then. The song is wonderful; listen to it in your free time.

That’s all from me. Subscribe to the channel, leave likes and comments. There are many useful links in the description of this video. See you in future episodes!