📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Qwen 3.5 Local Test with Ollama | Coding, OCR, Data Extraction, Image Understanding

Venelin Valkov15:46

Transcription

First task for our model is going to be to give us a recipe using the ingredients in the fridge. And here you can see the run quen 3.5 is the latest family of models from Alibaba cloud. And in this video we're going to have a look at what the model family contains and I'm going to show you how this model performs on my WCO instance using the 35 billion parameter model. Let's get started.

The model itself was introduced a little over two weeks ago and I was waiting for the WCO GGUF files to be flushed out so I can run it on my WOC instance. And during this time we also have received another versions of the model that are smaller compared to the original one which was 397 billion parameter model. Even though this is a mixture of expert model it had 17 billion parameter models that were active and note that on a wok machine this is quite a large model to run but other than that we have just seen a lot of smaller models and the more important ones at least to me are the 35 billion 3 billion parameter active version of the model which I'm going to show you during this video and the 27 billion parame parameter dense model. Here I'm running this on an M4 Pro with 48 GB of unified memory.

So the particular model that we're going to be using is going to be the 35 billion parameter window with 3 billion active parameters and this is a mixture of expert model. Note that this one has the unified vision and language foundation. So no more Quinn V models or visual language models. This model is trained with both a lot of images and text as input during the pre-training and the post reinforcement learning stages. And from what I've tested thus far, it seems that this is a very nice jump over image understanding compared to the quantry VL models. This particular version of the model has 256 experts and only nine are active during a single inference pass. So this is why this model is going to be running extremely fast even though that the watch part of it is 35 billion parameter and also we have 262k of context window and this one is extendable up to 1 million tokens.

If you want to become a better AI engineer and learn how to run your WooChO models on your Wcom machine and enterprise applications, go and subscribe to MXP Pro. There you're going to find a complete AI engineering academy that starts from the Python and machine learning basics. Then it goes to how to set up your wo environment and then how to build arax agentic systems, how to do evaluations and deploy those applications in production. So if you want to become a better AI engineer, go and subscribe to IMAX for Pro. Thank you.

The concrete version of the model that we're going to be using is going to be this quantized version that is available in OAMA. And here you can get the command to run this particular model. It will be roughly 24 GB of storage space. So it should take some time to download. But after that you are going to be able to run it on your WCO machine if that is allowed by your hardware. I'm in my wo instance and I have loaded the model and then you're going to see the real world performance that is this model is working quite well on my machine and also you can see that the responses and the thinking are presented to you during the response. So let's try it in a Jupyter notebook.

I'm in my wok cursor instance and this Jupyter notebook is going to be available within the GitHub repository AI boot camp that I'm going to be linking down into the description of this video. And here we have just a single function that is going to run our model tests. After this function is complete, we're going to start with the first task which is going to be to give this fridge which contains various foods and this was taken from a Bulgarian fridge. So you can see that we have a lot of different food right here in this one. We have some yogurt. We have some rosé wine. We also have some types of pickles, tartar sauce, uh tomato, dark chocolate, and other very interesting things that you can find in the very Suavic fridge of pretty much all Bulgarians. First task for our model is going to be to give us a recipe using the ingredients in the fridge. And here you can see Duran. Turn took roughly a minute and 40 seconds. But this particular run was quite slow due to the thinking that we got. And you can start to see the chain of thought from this model. And you can see that it started with the top shelf, middle shelf, door shelves, middle bin, and bottom shelves. So at the start it essentially enumerated every single thing that it found. Then it revised it, drafted the response and then he did a final polish. And after this one was the actual response. So you can see that the response that the model gave us is to give us uh this particular recipe. First, it started with the main ingredients. And then the recipe was creamy pesto and broccoli pasta with feta. So, it gave us pretty much a preparation time, cook time. We got pasta from the bio box in the door. Broccoli, red onion, and garlic. So, we had garlic guided the image just to make sure that this is available indeed. And I pretty much used all of the ingredients uh that are available. White tip contents, yogurts, etc. block of cheese, pesto from the jar, retonian and garlic. And pretty much I was amazed that even though this relatively small model was able to get the ingredients from this let's say disorganized fridge and it was able to give us a nice recipe using that.

The next task is to give this particular image and we'll ask the model to give us a UI using HTML and tailwind that is going to be very similar to what we are seeing on this image. And this particular response took roughly 3 minutes and 20 seconds. But the response was quite long. And uh it first started with drafting the HTML structure. Then it went on some particular notes on what components to use, what icons, and then it started with the response. And this one was particularly long. So, let me show you the actual response. It gave us roughly 5,000 completion tokens and the prompt tokens were 2,000. I think these are actually the thinking tokens. Let's check out the HTML itself. And honestly, when I first saw this, I was pretty blown away. Of course, this is still a relatively small model that was able to produce an HTML with tailwind and some custom fonts from Google Fonts to give us this particular design. And I would say that this is pretty close at least in spirit to what we had within the original image. And uh another thing here is that this is quite functional. We have some hover effects.

This diagram was created by a Reddit user that is presenting the domain driven design and the merge with the clean architecture. And I was asking the model to describe what this is, what are the main components, the flow of data to be explained and what kind of system is this. So this time around we had roughly a minute and 45 seconds in order to complete the chain of toad was very nicely structured with the identification of the components. We have a presentation layer, application layer, domain layer and infrastructure layer access and adapter to outside SQL implementation guardian of the data. It normalizes the inputs which sounds pretty good. Data flow. The flow of data flows a standard outside in or ports and adapters pattern moving from the user through the logic to the storage and back. So again pretty nice explanation. What kind of system is this? This is a clean architecture or hexagonal architecture also known as ports and adapter system specifically designed for domain driven design. Okay. So the model was very good at understanding what this is from a whiteboard diagram. Uh note that this whiteboard diagram is quite well done. If uh it was created by someone like me probably nothing will be able to be understood and read from it. But this particular user afraid has done a really nice job.

One very important use case for these types of models, especially now that they do support image understanding is received extraction. And here I have one from the cut. This is from London. And you can see that we have some drinks right here with some additional modifications if you will with a sub and total and service chart. and then it was given from a given telephone number and also the web page of the bar. So I asked pretty much for all the information the vendor name the total amount the date and the list items and I have given the received image as an input and we got a lot of information here within the chain of thought and probably this is something why the frontier labs are pretty much not showing what those models are thinking. Well, in this particular case, I think that the bottle got stuck in a whoop. Essentially, after some time, it went through what has really the items within the receipt and the response took a lot of time. Note that the prompt tokens and the completion tokens were quite large. And we got the items which is correct. the DBL here which is again correct and we have a subtotal the tax amount and the total amount. Let me check just the total 9234. Actually the total is different compared to what the model has extracted. Yeah, it is a different number. So even though we had such a large chain of to the actual total was incorrect. Let me check the date. 30th of April, 10th of April, 2025, I believe. So, yeah, it seems like that this particular extraction didn't do that great. I'm not really sure if the artifacts on the image were the things that have messed up the model itself, but as you can see, we had some wrong extractions.

The final test is to extract some data from a given chart. And this one in particular is from Tesa monthly chart. Getting it from trading view. And you can see that the monthly chart has pretty much information from 2026 up until 2010 at least June. So each of those bars is a Monty chart and I want the model to essentially extract all this information. You can see that this took roughly 6 minutes to complete and I have also increased the context window for this particular one. And we got this list of data points. I have gotten those and I just presented the closing price and you can see the general form of the chart is very similar to what we have within the trading view. We can see that the general form of the chart itself is quite well fitted the real data. But if we go through the data itself, you're going to see that we have intervals of 3 months. So, uh we are missing quite a lot of data here within this uh particular output. Of course, with some prompt tuning, we might be able to get much better results.

I would say that the distillation work that Quen and other Chinese WS are doing is performing quite well in the real world. I was very much surprised when I saw the coding capabilities at least within the building of the web page that I have shown you for the web shop. This just from a given image, the amount of color and the correct forms of the UI elements was pretty amazing for such a small model. And again note that this particular model is running quite well on a local instance. So it should be able to run for most of you guys. Also this model run very well on pretty much all the task that I have thrown it except for the received extraction tasks. I'm not really sure why the model got such a hiccup there. Maybe with different prompts or a different prompt this is going to be much much better. But from what I have seen the total extraction has totally failed on this particular extraction.

So thank you for watching guys. Please like, share and subscribe. Also join the Discord channel that I'm going to link down into the description of this video. And if you want me to test more of these wo models and for you to become a better AI engineer, go and subscribe to MX Pro and let me know down into the comments what you want to see next. Thank you for watching and I see you in next.