📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Обзор и Тест ЛУЧШИХ БЕСПЛАТНЫХ «думающих» нейросетей DeepSeek R1 VS OpenAI о3 mini

digital kir25:22

Transcription

The last week has seen a storm of hype around a new thinking neural network from China called Dips. Because of it, NVIDIA has lost a record $600 billion in stock value, marking the largest loss in the history of the U.S. stock market. All the media are reporting that this Dips is killing ChatGPT, as independent tests show it is comparable in capabilities or even surpasses the O1 model in some areas.

However, within a week, OpenAI released its new model, O3 Mini, which is currently the best model and is already outperforming this new Dips in independent tests.

In general, a crazy race has begun between Dips and OpenAI, or I would even say between the U.S. and China, for the title of the best neural network in the world. Over the past week, I have been testing the Dips neural network with its thinking model R1 and the same thinking model from OpenAI, O3 Mini. In this video, I want to give you an overview of these two latest and best neural network models and show which one is better based on the series of tests I have prepared.

My name is Kirill Alekseev. Let's get started.

Let's begin with an overview of these two neural networks. On the screen, we have Dips on the right and OpenAI with its O3 Mini on the left. I'll start with Dips since it is currently very interesting to everyone.

The main and first advantage is that it is completely free right now for everyone on the Chat Dipsomania website. Secondly, Dips also has mobile applications that you can install. You can download it on both Android and iOS. In contrast, you cannot easily install ChatGPT on your phone if you are in Russia or Belarus. In other countries, you will need to go through certain procedures. If you're interested, check out my Telegram channel where I have a pinned post with instructions on how to do this.

For accessibility, I give Dips one point because it is simply easier to access.

As for ChatGPT, they have released two recent models: O3 Mini and O3 Mini High. They even state in their interface that O3 Mini is a fast model for advanced reasoning. This means it can think before giving you an answer, and it does this quite quickly. O3 Mini High excels in coding and logic but will take significantly longer to think. Depending on the tasks you want to solve, you can choose either model.

Dips, on the other hand, currently has only one model that reasons, which is suitable for all types of reasoning.

I want to warn you that all prices and features I will discuss in this video are accurate as of the day of recording, and I am sure that in a week or even less, they will change, along with prices and new features. So, I am showing you what I have access to right now.

In summary, Dips is easier to access than OpenAI. Now, let's look at the functionality within each model. I have Dips open on the right, and for the reasoning button to work, I will get a practically instant answer without any reasoning.

However, if I do the same in a new chat and enable the Deep Think button and send that request, I will see how the model starts to reason. It thinks, but for some reason, it is thinking in English here. Sometimes it thinks in Russian, sometimes in English, although it should ideally think in the language in which I give the request.

We see this gray area, which is its reasoning. Accordingly, its answer should be of slightly higher quality.

Now, you understand that the O3 Mini and O3 Mini High models work in the same way; they will reason. An important feature of these models is the ability to search the internet. You can not only ask the model to reason about your task and give you a thoughtful answer but also send it to the internet. You can ask it to find some information, and the model will analyze a certain number of websites and give you an answer.

In my tests, Dips turned out to be slightly better at internet searching because it analyzes a bit more websites, around 40-50, from which it gathers information and then forms an answer based on that. ChatGPT also has an internet search function, both in the O3 Mini and the High model, but you need to enable the search button for this feature. They are roughly the same in this regard, but that’s where their similarities end.

Dips has a feature to attach files, allowing you to upload a PDF document, a table, or a Word file. You can attach up to 50 files, each up to 100 MB. This means you can upload data about your company or some marketing reports for it to analyze and give you a response in the role of a marketer. This feature is very important for reasoning models, as they are generally used for complex analytical tasks.

ChatGPT, in its O3 Mini and O3 Mini High models, currently does not have this button to attach files. They have this feature planned, but it is not yet active. However, I want to mention that O3 Mini was released just a few days ago as I record this review, while Dips has been out for several weeks. It’s possible that ChatGPT will roll out this feature in the coming week, and they will be on par.

On the other hand, the previous model O1, which I wouldn’t say is weaker, is actually on a similar level with Dips R1. O1 and R1 are roughly equal in benchmarks and various independent tests. O1 has the ability to attach files, and if you really need to attach a file, you can use it. However, it has a different problem: it does not have internet access.

In general, ChatGPT seems to have complicated things a bit. They now have many different models, and it will be quite difficult for an ordinary user to navigate through them. In contrast, Dips is much simpler; there are just a few buttons, and you’re good to go.

However, ChatGPT has another advantage that I really like: you can not only type text on the keyboard but also record it by voice. I can literally speak something. I currently don’t have permission in my browser to record my microphone, but you can set up permissions in your browser and record your prompts by voice. This saves a lot of time because short prompts are not suitable for reasoning models.

It’s much more convenient and easier to do this by voice than to type everything out in Dips. This is a downside for Dips, so I won’t give a point to either one.

In summary, one can attach files in Dips, while in ChatGPT, you cannot. In ChatGPT, you can use voice input, while in Dips, you cannot. They are roughly equal in this regard. I think it’s just a matter of weeks or months before both models roll out these functions and equalize.

So, it’s a tie.

Now, two final parameters that do not concern the interface but distinguish these two models: first, Dips is an open-source model. Anyone can download it locally on their computer, install it, and run it, ensuring that your data remains completely secure and is not sent to another company. You can run it locally, completely safely.

However, this is only available to a very small number of people on Earth because, first, to deploy it, you need a lot of hardware, which costs thousands, maybe even more than $10,000. So, it’s either for geeks using the model for personal tasks or for professionals using it for business tasks.

Now, regarding business tasks, both of these neural networks have open APIs, allowing them to be connected to other third-party services so that the neural network can send or receive information. We are developing chatbots that act as sales consultants and curators in online schools. We have various cases, and we are using the ChatGPT API. We previously used it, but now we will likely use Dips.

This way, we integrate them into the business processes of companies, either to assist their employees or to replace them. If you are interested in chatbot development, I will leave a link in the description for you to check it out in more detail.

So, each neural network has a paid API, but Dips is significantly cheaper than OpenAI, particularly for the O3 Mini, O3 Mini High, O1, and so on. Therefore, for business tasks, Dips is more cost-effective because it will simply be cheaper.

Since Dips is an open-source model and its API is cheaper than ChatGPT’s O3 Mini and O3 Mini High, it earns another point.

With the overview of the interface and technical specifications complete, let’s move on to a series of tests that will show us these neural networks in action.

Let’s start with a warm-up. I will ask both models the same question: "When will the Grammy Awards be held in 2025?" This will test their internet access and reasoning skills.

Dips seems to have struggled with this task. If we talk about the year, the ceremony will likely take place in February 2025, for example, on the 2nd or the 9th, as it usually has in previous years. It didn’t give me an answer.

Oh, here’s why it didn’t give me an answer: due to technical issues, internet access is temporarily unavailable. That’s a flaw on Dips’ part. Meanwhile, ChatGPT O3 went online and said the awards will take place on February 2nd, providing me with a link to the source. It’s a strange source; it picked some website, but I can still check the source for the accuracy of this information.

Ultimately, it’s up to me to decide whether to trust the site from which the model gathered this data. So, it seems like the point should go to ChatGPT. However, I want to give Dips some leeway. Despite not being able to access the internet and provide me with the correct answer right now, it had been able to access the internet and search correctly in my tests before recording this video.

I think there is just a huge influx of people using it right now because it is completely free. In the American store, it has taken first place and surpassed ChatGPT. So, you can imagine that millions of people are likely using this neural network right now, asking questions, and it’s just very difficult for them to technically support such a crowd.

So, despite the small flaw, it’s 1:1.

Now, let’s give a logic task to both neural networks. I will turn off the search function and ask, "Who was depicted on the anti-racism poster?" because they are simultaneously white, black, and Asian. You can think about the answer while Dips is reasoning.

I will choose the O3 Mini High model for ChatGPT because they recommend using it for coding and logic tasks. Let’s see what Dips responds.

Oh, look at that! What a long reasoning process! Dips took 41 seconds to respond. It said, "On the anti-racism poster, people of mixed heritage are often depicted." No, that’s not the correct answer.

The correct answer is a composite image of a person, famous personalities for children or families, or abstract symbols. Dips did not succeed with this answer.

Now, let’s see what ChatGPT says. It’s a panda! The O3 Mini High model got it right because the correct answer is indeed a panda. The poster against racism featured pandas because they have black and white coloring and are native to Asia, making them both white and black.

Unfortunately, Dips disappointed me here and did not succeed.

Let’s give both models another logic task and see how they handle it. Maybe the previous question was in OpenAI’s dataset, which is why the model knew the answer, while Dips did not.

For the sake of experimentation, let’s ask, "In London in the 16th century, a unicorn was hung over the apothecary's shop. What was depicted over the fruit shop?" Think about it while the neural networks respond.

Wow, what long answers! Both neural networks took a long time to reason: Dips took 28 seconds, and ChatGPT took 38 seconds.

Dips said, "The image of Adam and Eve over the leather shop is related to the biblical story where, after being expelled from Paradise, they received leather clothing. This symbolizes the beginning of working with leather." That’s an incorrect answer.

Let’s see what ChatGPT answered. It said, "Over the shop where apples are sold." That’s the correct answer! It’s a fruit shop because Eve was with the apple, symbolizing the Paradise they were in.

Now, for logic, ChatGPT takes the point.

In the next test, I will check how the models handle coding. I will send this request to the O3 Mini High model because it handles coding better.

The request is: "Write the complete code for the game Arkanoid so that I can run and play it in the browser." This is a very simple and straightforward request; I won’t complicate it. Let’s see how the models handle this primitive request.

ChatGPT succeeded. I will copy the code and open the O1 model to paste it there because it has a function to run code. I want to run this code within the function.

Here’s my Arkanoid game launching. Yes, I see that the model did quite well. The ball is flying as it should, and it even shows me my lives and score. Overall, I’m satisfied.

Now, let’s see Dips. What I like about them is that even in the R1 reasoning model, there is an immediate "Run Code" button.

Oh, look at the glitches I’m observing. It seems that the window might be minimized. Dips did not succeed here; for some reason, the blocks are hanging, but the lower platform that bounces the ball is missing. The blocks just disappear.

I would give the point to ChatGPT here, but let’s look at another coding task to see how they handle it. This time, I will ask the model to create a diamond with a ball inside that moves according to the force of gravity. Let’s see how well the model can understand and describe this and provide me with code to run this animation in the browser.

ChatGPT gave me a response again. Dips is still thinking. Let’s check how long ChatGPT took to think about this task: 1 minute and 8 seconds.

You can evaluate how deep this reasoning was and how long it took. It’s interesting to read, almost like a learning experience. If you want to master development, you can see how a mid-level developer thinks.

I will copy this code and paste it into O1 to open it in the environment. There’s a video on the channel where I explain what this function is.

In ChatGPT, I personally like this feature, but I usually use it more for text than for code.

I see that, first of all, it’s not quite a diamond but a square. Secondly, the ball doesn’t seem to move under the influence of gravity; it looks like it’s been made inflatable, like a rubber ball, and it bounces in an unnatural way. I would say it succeeded about 50% with this task, and I need to refine the prompt to get the animation I need.

Now, let’s see how Dips handles this. I will run it right here in the browser.

Oh, actually, Dips seems to have done better because, in terms of gravity, the ball looks more metallic. However, it generated a cube instead of a diamond.

Overall, the ball seems to move logically according to gravity. So, for the first coding task, ChatGPT gets a point, and for the second coding task, Dips gets a point.

However, I wouldn’t rush to give them equal points because, in terms of benchmarks, O3 Mini High significantly outperforms the R1 model.

If I were to conduct five more tests, Dips might fall behind. So, considering the benchmarks from my tests, I will give ChatGPT 1.5 points and Dips 1 point, as ChatGPT seems to be slightly better.

The next task is somewhat simple but involves logic, math, and a bit of grammar: "How many letters 'R' are in the word 'Bureaucrat'?"

If you look closely, you will see that I intentionally made a mistake in the word; I wrote it with two 'R's together when there should only be one. I want to see if the model will notice the mistake.

ChatGPT gives me an answer after 6 seconds: "In the word 'Bureaucrat,' the letter 'R' appears three times at the third, fourth, and seventh positions." I’m not very satisfied with this answer.

Dips also responds the same way: "In the word 'Bureaucrat,' the letter 'R' appears three times." They just counted the letters without noticing the mistake.

If they had paid attention to the reasoning, they might have seen that the word was written incorrectly. Dips didn’t even notice the mistake in its reasoning.

ChatGPT, on the other hand, noticed the error and said, "Usually, in Russian, the word 'Bureaucrat' is written with two 'R's, but here it gives three letters."

However, it just counts three letters. I expect these reasoning models to identify user errors and point them out.

Suppose I wrote this prompt unintentionally with a mistake. If it were not a primitive task like the one I’m demonstrating, but a task where the model would think for 3 or 5 minutes, and it simply didn’t notice a basic error in my prompt, that would be a problem.

So, I expect the models to notice such primitive mistakes and guide me. I was hoping for a response that said, "Yes, there are three 'R's in the word you sent, but please note that the correct spelling has two 'R's." That would have satisfied me.

But for some reason, both just answered that there are three 'R's, so I give zero points for this test to both models.

The last test task for both models: I will choose O3 Mini without High for ChatGPT because it handles text tasks better.

My question is: "Write an eight-line poem that rhymes in Russian." This is very important in the context of the competition between Dips and ChatGPT. Let’s see how they handle rhyming and writing.

ChatGPT gave me a response very quickly, in 5 seconds. Let’s see:

"In the world of bits and codes, a fire ignited,

Where mind and deceit met in an old fight.

Dips skillfully weaves a false narrative,

While ChatGPT strives for truth, honest and active.

Virtual heroes argue in hand,

Each line proving their strength in the arena of ideas and algorithms,

Where the light of minds creates a new digital covenant."

There’s some rhyme, and I liked the punch from ChatGPT about Dips weaving a false narrative while ChatGPT is all about being correct.

Let’s see how Dips responds.

Oh, what a long reasoning process! Look how long it took to think about this task. ChatGPT took only 5 seconds, while Dips took a whopping 86 seconds to create a simple poem.

Wow, that’s a lot of time, in my opinion.

Let’s see:

"In the digital world where tokens boil,

Dips and smart ChatGPT clash in a duel.

No rhyme, one wit, a bright sunset,

The other, the depth of artificial mind.

Among algorithms, dawn sparkles,

Sometimes Dips jokes, sometimes ChatGPT gives advice."

What is this nonsense, Dips? This is just terrible, to be honest. It took 86 seconds, and the poem doesn’t even maintain a cross rhyme.

In fact, there’s no rhyme at all.

So, clearly, the point goes to ChatGPT.

Even in its model, which doesn’t significantly outperform R1, it still performed better than R1.

With the tests concluded, it turns out that ChatGPT earned 4.5 points, while Dips earned 4 points.

However, I want to emphasize that Dips earned most of its points for its openness, affordability, and accessibility, while ChatGPT earned its points for the quality of its responses, but it is more expensive and harder to access.

So, the conclusion is that if you need a simpler neural network for simpler tasks and just want to start trying it out, I would recommend starting with Dips.

If you have a paid subscription to ChatGPT for $20 a month and have no issues with enabling or disabling the service, downloading it to your phone, and so on, then, of course, it’s better to use ChatGPT.

Personally, I use the paid version of ChatGPT and will stick with it because I think it’s better.

And here’s a shoutout to Dips for emerging and pushing ChatGPT to develop further.

Before I finish, I want to answer a few important questions I’ve received about these models in my Telegram channel.

Regarding prompts, there is a difference in how to ask questions to reasoning models versus regular ones. Regular models understand context very well, and you can have a dialogue with them. I can ask one question, and they respond, and I can clarify or explain in a second message, and we can go back and forth discussing a topic.

In the case of reasoning models, it’s better to give a monolithic prompt all at once, detailing the task, the role of the neural network, and all the context you have. You can either attach files or describe everything in the prompt. This will significantly improve the quality of the response.

I have a small guide on how to write prompts for these reasoning models in my Telegram channel. If you’re interested, subscribe and check out that post.

That’s it for this overview and test of the two best neural networks. I hope you found the content interesting. See you!