Transcription
Oh, it's been a while, but man, I have been down coding probably 12 to 14 hours a day over the last nine days. So, I have a lot of opinions. I've used a ton of the new models that have come out and I want to kind of just share through my initial opinions on this.
So, to get this started, I mostly been working with Vue3, React, Python, TypeScript, even did a bunch of PHP and a lot of data engineering. So, my feedback is going to be based on these types of uh tech stacks. Let's jump right into it.
Gemini 3 Pro. I don't know if there's ever been a model more hyped than Gemini 3 Pro, to be honest. People were talking months ahead of time about all the crazy stuff this model could do. But, I do want to just call it out. I think it's a good model. I mean, I don't think it does anything that has like blown away the world of AI coding in such a way that like it literally changes the game, but it is a good model and I will show you that I have actually used this model a significant amount. In fact, right when Gemini 3 Pro came out, I switched over to pretty much just using that. Before that, had been using Composer 1.
Now, the pricing of this is interesting because at $12 uh $12 on the output or if you get to $18 greater than 200K, it's it's it's a reasonably priced model, but it does seem to be a little bit more token hungry, at least in my opinion, than some of the other models that I've used. The contact cash pricing is great, though. like we get 20 cents on the reads which is just awesome and I think that helps save a lot of pricing there.
Now opus 4.5. I literally was not expecting this and when I heard it come out I was like oh man that that's so cool. I wish I had time to test it. You know what I assumed though? That the price was going to be the same. This price actually puts it into a range not great. It's still expensive, but it puts it into a range where we could actually use it some. And in fact, I think it might be 100% worth it, especially with the prompt caching, the read side of this, cuz you can work for a relatively inexpensive amount. It's not super cheap, but just something to keep your mind on when I go over.
Now, GPT5.1, this was the first model that came out while I was heads down doing some work, and I couldn't really tell a difference with this model, unfortunately. And in fact, I still think there's something going on with the GPT5 models either. I have become spoiled based on the other things I'm using, or there's just literally something that's changed here that I can't put my finger on.
I do want to like talk a little bit about just highlevel what each of these models do. So, this is composer one. This is like one of my favorite models to like just get an answer quick. It is so freaking fast. And this is just a portfolio site that I built with it. Looks so much like all the Chinese models. I'm going be honest with you. Like this looks so much like the GLM4 models or the Quinc Coder models, but it's nice. Like it's a nice model.
Now this is Gemini 3 Pro. Like this actually has some design chops on it. And I know like this is just a static portfolio site, but there's like subtle movements over here on the right that might actually be hard for you to pick up. Like this is moving up and down. We've got a beautiful like uh type ahead that's actually happening here. It put some uh picture apparently that it found online in here. And if I hover over it, it goes in. There's like the micro interactions are just phenomenal here. They've got even put some fe uh featured projects in here and the get in touch. Like everything in here I feel like is just very different than what I've been seeing lately, which is very much a lot of purple and light color. I I I really like this. Um, one thing I did notice is like this hover here I think is actually really nice. You can kind of see the uh it just it did a great job.
Now let's take a look at GPT5.1. This one blows my mind. same prompt. And I'm not even really sure what to make of this to be honest with you. Like, this is the most bizarre version that I've ever got with this prompt ever. It's kind of neat, but honestly, I can't really judge what's going on here because this is not a portfolio site that I would ever actually reveal. But it it is unique and kind of like exemplifies what I see with GPT5.1. It's just an odd model for me personally to work with. And this is not at all what I would expect when I ask it to create a port personal portfolio site for me. This is one of the things I like to do just to try to understand its design taste.
Now, I have actually heard people say Opus 4.5 is not good at design. I beg to differ. I actually think it's phenomenal. Like, it's probably hard to see, but they've got a very cool like scroll animation here. This is just amazing. Like, I love the the flavor that I put in here. And the colors like the you can see here this planets like going up and down. This this part here I think is really nice, too. We've got I don't actually think clicking the view project does anything, but if you hover over it, you see it kind of like it moves around. Uh I think this is this is great. I'm definitely not a designer, but if I were actually launching a personal portfolio site and I built something like this back 5 years ago, you know how proud I would have been of something like this. This is great. I love the color style that it actually picked here. Um, so this gives you an idea.
So Composer 1, GPT5.1. I don't even know what to make of this. I think Gemini 3 uh Pro and Opus 4.5 actually. I I don't know, man. I think they've leveled up everything that I've seen so far from my design experience. I think they've beat like my GLM4.6 and and I've did a lot of front-end design and I've done a lot with Gemini 3 Pro and I've done a lot with Opus. I think they're great. I do think Opus 4.5 is incredible too at deep diving into a codebase and figuring out what the actual root cause of things are and building a plan for that.
So, let me just jump into a few other things. I wanted to show this. So, prior to Gemini 3 Pro, I was almost solely on composer and it like I can't show you all my history. A lot of this actually has been done in cursor. I did use some other tools, but most of it, I know this sounds weird coming from me, but I've actually started to become a fan of Cursor. I love their plan mode. They come out with models so quick. A new model comes out, I have it within minute. It feels like within minutes of announcement just there that they've turned around my opinion of them quite a lot. And they've gotten a lot more transparent on how pricing and usage actually works. Even though I despise this credit system, I'm going to be totally honest with you. I do not like this credit system at all, but at least I can hover over each of these and see like how the tokens actually were used like right read cash not, etc.
But yeah, I switch purely over to Gemini 3 Pro. My general thoughts are it works good most of the time. Again, this isn't cursor, but sometimes the model would just quit calling tools for me and I could never get it to work again in a particular uh context that I was working in. So, I just would have to kill the context and start over. A little disappointing because sometimes I would actually prime it a bit and like get the kind of like what I wanted to have happen and then I'd have to take that to a new chat to actually execute on it. Didn't happen a lot, but it happened enough times for me to get annoyed with it.
The other thing is is, you know, you guys probably see this too, but some models have a tendency to be like, uh, when I tell it to do something, it leaves a comment saying that I told it to do it, and it's like, I don't want that comment there. That comment's useless. Like, I I just wanted you to make the code change. I've been trying to figure out a way to get Gemini 3 Pro to not do that. But fast model, I really enjoy using this model. I mean, really, really, really enjoy using this model. had no major issues other than the ones I've called out here. And it was fast. Like I it's not as fast as Composer One, but I I mean I was able to build some incredible stuff and like get some great things done with three or four shots of like working back and forth and kind of guiding it. It it did a great job. It also does a great job building plans and cursor.
Now, Claw 4.5 Opus. When this came out, I I do want to just kind of iterate this a little bit more. Now, you can see the number of of credits that I'm using here. Like this one, 6.6 million, 72.4. I probably have used from a time perspective, Gemini 3 Pro, probably three times longer than Claw 4.5 Opus just due to when it came out. But oh my gosh, Claw 4.5 Opus High Thinking is what I mainly used. And honestly, the only gripe that I have about it is sometimes it makes kind of bad assumptions. And it's not even it's not even the worst assumption. Some I would call them like mid assumptions, but you can kind of get around that with prompting. And this is like if you know the codebase and you know what it's trying to do, but holy crap is this model insane. It's faster than I thought, cheaper than I thought. still slow, but it is such a joy to work with. It It is I I don't know. This almost feels like my favorite model right now. I haven't put enough time into it, but I'm on the $200 a month plan right now with Cursor. I know that's crazy, but like I would most likely use claw 4.5 opus the majority of the time it on incursor currently because like I said I I feel like it does a good job taking my design system. We have a design system and we're trying to build stuff out. It does a great job pulling that in, using it properly, all of that stuff. Big big big fan of this model and GPT5.1. I posted this on X that I could not tell much of a difference to be fair. I did not give this model a ton of time and I've only used it a little bit in codeex, but I want to show you what was happening here.
So, I went this is this is the order of things. I went from composer and then I composer annoys me sometimes. So, I was switching to uh cloud 4.5 Sonic thinking to actually do like the more the harder stuff. So this is pre pre all the new models that came out, right? So I was doing composer for fast stuff, four to five sonet thinking to try to do stuff that's a little bit trickier. GPT51 came out, I tried it and so I was doing a task and I remember very particular like, oh, this is frustrating. So I went back to claw 4.5 to actually get it back on track. And then eventually you can see here that this there's a bunch of time in between here, but you can see here I kind of had given up just in this little block of time. Back to cloud4.5 sonnet thinking back to composer because GP5.1 was slow and the quality of what I was getting did not meet my expectations for what I was trying to actually develop here. And at this time I was actually working on back-end code, but like look at this design. This is just so bizarre to me. Like uh same prompt and how it kind of interpreted this. It just in my opinion doesn't have good design taste. It doesn't do a good job gathering context about the codebase. It almost like misses um details that it should have. And then I would put it on cloud 4.5 sonnet thinking to try to like recover some of that. So I gave it a fair shake. Not as much as I would really want to. Uh, but I needed to get work done. And that to me is where like what actually matters. Like when you're on a tight deadline and you're actually working, I don't have time to uh fluff around with GPT5.1 if it's not doing what I want. when I could get back to composer and composer and I know this may sound weird because composer one is not I would not say is a better model than GPT5.1 but it's so fast that I could guide it to go get the stuff that I need. So it became like my it came became like an extension to me where I knew what I wanted to accomplish and I could guide composer one a lot faster. when you were lean on a slower model like GBT 5.1, you kind of need it to do some of the like determining of what needs to happen and figure that stuff out a little bit better.
And then when I switched to Gemini 3 Pro, literally it was just non-stop Gemini 3 Pro uh up until the time Claude 4.5 Opus went. And I'm actually a little bit sad because we did not get much time between Gemini 3 Pro and Claw 4.5 Opus. And I think these two are my favorite models right now. Again, I've only been using them since they came out. And time will really tell like if they end up working. You know, this is pro preview. We know what happened with 2.5 Pro Preview. If they end up messing this up, but 4.5 Opus, this pricing start I mean, it's reasonable. It It's It's still very expensive, but it's way more usable than the other Opus was.
So, my general sentiment on GBT 5.1 is something just feels off with this model. and I can't quite put my finger on it. It's is its way that it uh like gathers context or is it just the way that it makes up like the oddest like decisions that it ends up doing that drives me nuts. Uh I don't know, but I would love to get more feedback on what you guys are seeing with GBD5.1 or in fact maybe no one cares.
The other thing I want to do is I'm in the midst of running evals on Gemini 3 Pro for the end of the month across all the agents. I also am going to do that with call 4.55 Opus. So these two models for sure are in. I'm probably going to limit it down to like three or four models cuz just I there's just time and then the holidays and everything are going up. So GPT5.1 for example, I I may I may actually try to run that one as well just to kind of get a benchmark on that and then I may throw like Sonic 4.5 in there, something along those lines. But if there are other models that you guys are interested in me trying to run um to for the December benchmarks and I am going to try to get back to those.
This has just been a crazy time with my team getting acquired and like you know all the transition stuff that's happening, emergency projects, you know how all that stuff goes in the real world. And the other cool thing is I'm getting a really cool understanding from hundreds hundreds hundreds of engineers on you know what they like and don't like about AI. So I'm getting so much perspective about this that I wasn't able to get before.
Anyway, I appreciate you all. Thank you there for hanging in there with me as I kind of work through everything. Let me know in the comments below any feedback that you've got on these models in particular. How far off am I in the way I'm thinking about this? I wrote more code in probably the last nine days um than I did in the prior month. So, it has been very much heads down grinding code. Very many repetitions on both of these. Unfortunately though, a lot of it's mainly been in cursor uh clawed code quite a bit. And then I did use a little bit of root code, but I have not been around all the other agents yet. So, I need to go spend the rest of today figuring out what's been updated there.
All right, everyone. Till next time. Peace out.