Transcription
And throughout the entire experiment, I repeatedly estimated how much faster I was with AI than humans. And at the beginning, I was sure I would be twice as fast. Then, in between, I had the feeling, oh, four times as fast could also be realistic. And towards the end, I thought, yes, eight times as fast could also be the case. But it was completely different, because in reality, the factor was. If you follow a lot of content here on YouTube, you know there are numerous videos where people develop entire apps in a very short time with hardly any knowledge, just based on their voice. But let's be realistic, for us in professional software development, these statements have no basis and no value, because we work with huge infrastructures, mountains of requirements, really large enterprise architectures, and with a very, very high standard for quality. For years, I've been doing consulting and advising precisely in this area. Large enterprise applications, large architectures, thousands of requirements with a very, very high standard for quality. And I asked myself, what if one were to do an experiment? And in this experiment, develop a large, large enterprise application with all the bells and whistles, with all the zip and zap, what you need every day in software development. What would that look like? How good would AI be by mid-2026? That's why I started a self-experiment. You might have noticed, since the beginning of the year, there have been very few videos, and there was a good reason for that, because for the last four to six months, I've locked myself away and tried to dedicate as much time as possible alongside my normal job to develop a complete project with a large architecture and high quality requirements from start to finish, and that is now done. And the results, honestly, are unbelievable and they haven't let me sleep for weeks. But that's exactly what we're going to look at in this video. It starts after the intro. Enjoy. In the last two or three videos, I even told you that something happened at the beginning of the year that revolutionized the entire software development with AI and will reshuffle all the cards. And in this video, we want to dedicate ourselves to this very topic, and I've put a lot of time and energy into this video to give you as many insights as possible. But on the other hand, I want to keep the whole thing as short as possible, otherwise the video would go on for hours. Therefore, today in this video, we will do a first overview of this exact topic. I will present the project to you and the results. What's important is that we try to keep it abstract. Not too technical yet, because I want this video to reach a lot of people who are not just software developers, but also architects, testers, decision-makers, and so on and so forth. I've been using AI extensively in customer projects for a year and a half. In some more, in some less. We use it a lot for requirements, for quality gates, of course for code generation, for architecture validation, and so on and so forth. But in most projects, it's only used partially, not really coherently, to really think through a software lifecycle from beginning to end. Exactly with this AI. And I've always had the feeling, okay, especially with this event, we'll get to that, at the end of last year or the beginning of this year, there's much more possible, there's much more quality possible, and above all, much more speed possible. What exactly do I mean by this thing that happened at the end of last year, beginning of this year? From my perspective, from the professional software development perspective, I would interpret it as a new era. New era means that up to this point, we've had a lot of models. I've only shown Claude Opus and GPT here, and these models have gotten better and better. I think everyone has noticed that in relatively short periods of time. But around the end of last year, so around the end of 2025, beginning of 2026, there was a complete shift, because suddenly models came out that could do something very, very well that wasn't possible before, and that was faithfulness regarding guidelines, regarding guidelines. That is, before, if you gave it specifications on how an application should be structured and written, it followed them so-so. Very often, especially with large projects, it would eventually start to completely disregard them, but with this new generation, around Opus 4.5, that changed completely, and suddenly it was possible to give it larger specifications, larger guidelines, larger guidelines, and guardrails, and it suddenly started to adhere meticulously to these things and build the application exactly as you actually wanted it. That's why my mentioned mission. I've tried to take as much time as possible from the beginning of the year until today and to completely redevelop a larger application. I've been developing this application for 18 years. It's my private application, my accounting, my back office application, and my goal was to use as much AI as possible, to use an enterprise system architecture, to use an enterprise software architecture, and to adopt a quality standard as we do in my, well, I'd say, most meticulous customer projects, where we really pay attention to everything. I want the whole thing to be tested. I want it to be documented. I want the whole thing to be specified. And not just for a small iPhone application or what you've seen elsewhere, but it's supposed to be a large application. This application, which I've been developing for 18 years, I've redeveloped it, and not just this application, but it ended up with about 12 to 15 times the functionality. So, for my standards alone, it has already become really large. And the specific questions I had were: A, is it even possible with the current LMs, with the current AI agents, to build such large applications with a good quality level? Then also, how fast is the whole thing, and what are the results like in terms of quality? That is, can we really achieve the results or the quality levels that I normally have in large projects with large customers? And of course, the cost aspect. What does it cost in the end if I develop such an application alone with a given high quality standard, end to end? This old application we're talking about is my old back office application. You only see a screenshot here, and I've been using this application for 18 years to create my invoices, create my receipts, manage my customers, etc., etc. But it's very dated. It's a .NET Framework application with very old UI technology, and in the end, I do everything in terms of modules, in terms of functionalities, that I normally needed over the course of these various years. But it must be said that I've always only implemented a fraction of the functionalities I would have liked to have. There's much, much more possible. And I would like to implement this "much more" in the new application. This application is supposed to consist of several components. First, the back office, which I already had a bit in the old application, but with much, much more functionality, a modern interface, and a whole lot of things that work in the background, sending emails automatically, etc., because the old application and my homepage, for example, were completely decoupled. That is, if a customer made an inquiry, I received an email. I then entered the email manually into the back office, etc. And I want to have all of that completely integrated now. I want to have a huge, comprehensive back office system, and that should be integrated into a new homepage, also newly developed in this project, where one can book workshops, sign up for email newsletters, join waiting lists, submit project inquiries, and so on and so forth. To make communication with my customers smoother for me, I also wanted a customer communication page. That is, my customer receives emails where appointments are reserved for them. They can then click on the link, come to this customer area, and then, I don't know, confirm appointments, confirm an offer, or download an invoice again, or any other things that occur in my customer communication. And personally, I hate office work. I love doing this job of consulting, training, and advising, but I just don't like office work. And what you see here, the back office, the homepage, the customer area, that's a nice to have for me. Okay, the homepage isn't, it's important, of course, it's for you, the customers, but the rest of the things are a nice to have for me, because if everything runs ideally, I don't want to be actively working in the back office at all, I don't want to really work, but rather I want an AI system to run in the background, namely Hermes. Hermes is the latest agent scream, shit, scream I meant to say, running around out there, and I want this agent to do a large part of the application for me via the API in this application. Create invoices, create receipts, book receipts automatically, handle part of the customer communication, etc., etc., because my goal is to completely automate 80-90% of the things that annoy me about office work by the end of 2026, and to automate them with Hermes and, very importantly, with local AI. I naturally have relevant customer information that must never go to the cloud. I have very sensitive customers, so everything here should function locally. We'll get to that in later videos, but that's my scope. I said enterprise architecture. Enterprise architecture, of course, means a lot of architectural forms, but I decided on a microservice architecture. Do I need this microservice architecture? No, what do I need? Absolutely not. None of these aspects are relevant to me, but in the last 15 years, this back office application that I've already built has always served as an example for my architecture workshops. There's a monolithic architecture, a layered architecture, and a component architecture. And here, I thought for future workshops, a microservice example, especially one developed with AI, would be a pretty cool thing. That's why microservices. These microservices, or upstream of these microservices, is an API Gateway, handled with traffic, and additionally on top of that, these three clients you just saw, so the homepage, customer UI, and the back office. Below that, there's also a workflow engine, N8N, and this workflow engine is supposed to handle all business processes and the execution, the orchestration of different domains. We'll do a video about that in a week or two, then you'll see how that works in more detail. That means not only the communication from top to bottom, but the individual services can trigger domain events, and these are then used by workflows in N8N. The next step is that there are also external services, so Office 365, Google, Trello, all sorts of things where I send emails, make calendar entries, create team meetings, or record something in Google Sheets, create cards in Trello, so that I have a unified overview of what still needs to be done, and so on and so forth. So, also a lot of services. Then, of course, the local AI. As I said, I want a large part of the work here to be automated by AI agents. For that, I need a local AI. I've decided on Gen, Gen, Alibaba Gwen 3.6. A 36 billion parameter model, I believe, a Mixture of Experts model, and we'll do a video about that too. And the whole thing runs locally on my hardware and is offered upwards into my internal network via an MCP server. That is, it says in this accounting system, dear AI, can you call these and these operations, create invoices, I don't know, book receipts, and so on and so forth, and on top of that sits the Hermes agent, which runs on my server and is supposed to take over most of these things autonomously. So, the architecture is, of course, only one thing, because I wanted to cover the DevOps area as much as possible as it is normally done in practice projects. That means we first have two different environments, the development environment and my production environment. I don't have server infrastructure, I only have a NAS system here, on both my desktop PC and the NAS system, Docker runs. And I've always had two PCs. One is here at my desk, and then upstairs in my daughter's office, in the children's room, I've set up a second office where I could work late into the night. By the way, many thanks to the nice colleagues who wrote in the last two or three videos how bad I look, that I look like someone ran me over. Those were exactly the four months where I mostly worked late into the night. So, in any case, two development machines and the NAS environment, and first, the coding agents that I used. At the beginning, a lot of Codex, later almost exclusively Cloud, and also local AI, but we'll talk about that in the next or the video after next. On the one hand, they have access to a code repository where all the source code was stored, but of course also to another folder where specifications are. We'll get to that shortly, and it states what exactly should go into this application in terms of functionality, i.e., which modules, which functions, etc. At the same time, the coding agent also has access to a so-called harness, and this harness doesn't specify what should be implemented, i.e., the function, but how it should be implemented, which architecture, how it should be coded, how each environment should be addressed, and so on, we'll get to that shortly. And on the local developer PCs, of course, Docker for Desktop is installed, which the coding agents could access directly. That means they can not only develop the source code but also deploy it and then run it on Docker, look at the log files, etc. If there's a new version, it can be pushed to Gitty. Gitty is a lightweight alternative to Gitlab, and this Gitty then uses a build agent, and this build agent pushes to two different environments on the NAS. Once a test environment and a production environment. In the test environment, these two applications run again. In the test environment, you can see from this small mask here that there's an addition because, as you'll see later, I need to add a new component to test the workflows better. But from a setup perspective, I'd say this is about 80-90% of the customers I have who work with microservices, more or less like this. Last important part, because I want the AI agent to work as autonomously as possible, it has SSH access to the NAS. In the test environment, it has almost all possibilities to work with Docker, to pull log files, and so on, to set up new containers. In the production environment, I've restricted the whole thing. It can at least retrieve the log files. That means, if I notice there's a problem in the production environment, I can directly use the coding agent to analyze it, so it can tell me more precisely what exactly is wrong. What's important, of course, with all of this is that I don't just want the whole thing to be good in terms of quality and structure, but I also want it to demonstrably work. And that's why I've reduced myself to two things when it comes to testing, or decided on two things. First, API tests and workflow tests. You can test a lot more. I could do unit tests and UI tests, but for me, it's important that the Hermes agent can later work with a clean backend infrastructure, with a well-functioning backend infrastructure. And that's why I tested the endpoints for the individual tests. This means, if I have a service here, it has its own database instance in the test environment, and then the test can perform various operations on this service, and it can, of course, evaluate the responses on the one hand, i.e., test if it gives the correct answers. It can also go and check if the correct events, the domain events, are triggered. That is, if a new customer is created, ideally a message like "new customer message" should be placed on the message bus, for example. But it can also directly perform searches on the database and check if any manipulation operations on the database have been correctly mapped. With the workflow tests, the goal is to start up the entire environment with all services, because the workflow engine N8N runs on top of it, and the goal is now that I have written tests or had tests written that naturally execute these workflows, that work in the background with all services and the correct DAT, i.e., the correct database with test data, and then this test can again, of course, verify the response from the N8N workflow and execute it, and at the same level, of course, also assert the database, i.e., check if the workflow has correctly asserted the data in the database. And that's where this additional point from the test environment comes in, which I just mentioned. These workflow engines naturally send emails via Office 365 or create Trello cards or open Google Sheets or whatever, and these are always REST calls to systems that are not in my test environment, towards Office, towards Google, etc. That's why I opted for Wemmck, and Wemmck essentially fakes Docker so that these calls, which would actually go to Office 365, for example, are intercepted by Wiremock and then corresponding mock data is returned, which this test can then use to assert. This means that the components backend and workflows, which are extremely important for my AI agent in the future, are tested sufficiently. After we've talked about what was developed, how the whole thing should be tested, with which architecture, with which environment, we now come to the way I proceeded. I started at the beginning in a relatively naive way because I wanted to compare different approaches and, above all, the final quality. In Phase 1, I started, which I always call micromanagement. You know it yourself, when you work with LMs, you have a prompt at the beginning, and from this prompt, you generate source code. This source code is then reviewed accordingly, whether the functionality is correct, whether the code structure is correct, etc., and then you iterate over it. That is, you move on to the next prompt. This is, of course, very granular and often has the problem that the LM often outputs source code, architectural approaches, designs that are not really suitable at all. That's why I moved relatively quickly to the second phase. I called it quality-driven. That is, we still start with our query, i.e., with the prompt we send to the LM. But this time, we have defined a so-called harness beforehand. You can think of a harness like a template for the LM. The LM normally knows nothing about our application, our architecture, our intentions in the application, and in this harness, we try to give this LM more intention about how the architecture, the source code, etc., should look. So, we have different files in this harness, and from these files, a so-called system prompt is generated. That is, every time I open a new session with Cloud, with Codex, the content of the harness was prepared and sent to the LM. So the LM knew exactly how I wanted to develop, which guidelines should be followed, and so on. Then I generated these prompts again from the prompt, as before, the source code, and then came the first sub-step, because I incorporated tools into this harness. Normally, you can have the LM check if the source code complies with a specification or guideline, but now you know that this is not particularly accurate and, above all, it takes a relatively long time. I went ahead and used statistical tools. In this case, ReSharper and NDepend for the source code level. And with this combination of guidelines, static code analysis, and static structural analysis, I was able to ensure that the LM could ensure after every code generation that the code change was within my guidelines. For this experiment, I used two different guidelines: first, the architectural guideline, which describes the system architecture. I was allowed to take this from a customer project for NDepend, with 880 rules, and a very extensive coding guideline for C#, but also for Blazor. That's the frontend technology from Microsoft. So, after this tool check, after we could ensure that everything was designed according to the architecture, I then went ahead and did the control again. This time, however, not so much the control regarding functionality, because mostly the functionality was already implemented correctly. We'll get to the next step shortly, but it was much more about observing, okay, is the source code that has now come out as requested in the harness? If there were any deviations, it was clear that the harness had to be corrected. This is also called harness correction development, where you iteratively go and correct this harness. Phase 3, then Spec-driven. Spec-driven means that we no longer work primarily with prompts, but we teach our LM how requirements are broken down into individual development tasks. And the moment we communicate to our LM in this harness how we want to do it, how we envision it, we can take specs directly, i.e., direct stories that you normally know from Azure and Azure DevOps and Jira, and give them to the LM, and the LM breaks them down into different prompts, works them off, develops source code again. This source code is then checked again with tool support to see if it complies with the design and architectural guidelines. Then it's checked again to see if everything is correct, if the requirements were broken down correctly, if the architecture was implemented correctly, and if necessary, this harness is checked, corrected. Phase 4, then Test-driven. I worked with tests directly from the beginning, wrote the first tests manually, had them analyzed, but then quickly moved on to defining in the harness how we want to test. That is, a new file describes exactly how a test should be structured, what should be tested, what shouldn't, how the tests should be set up, how they should be executed, and so on and so forth. So that after this spec, from which we generate the prompt, we first have a test case generated. And importantly here, many people show tests in videos where source code is taken and a test is generated for it. That doesn't make much sense, because then the test simply tests the source code that was there. In software development, we want to generate tests that verify system behavior, requirement behavior. So, if we have 10 requirements with, say, 30, 40 acceptance criteria, the purpose of the test isn't to test something we've written in the source code, but rather tests should be written that verify the specification behavior. And here, these tests are derived directly from the requirements, from the acceptance criteria that I've written there. The further process is then as usual or as in the previous step, we generate source code. Then we execute the tests, of course, and see if the source code has achieved the desired system behavior. Then again, the check if the architecture was okay, the verification, checking if the tests are correct, if the architecture implementation is correct, and if necessary, then again this correction of the harness. Fourth step, then, is the Idea-driven phase, where you go and say, okay, we've now brought our LM to the point where we can input specifications into our LM or into our coding agent. And we're going a step further, because if we look closely at what happens in the requirements analysis phase, we have ideas from stakeholders, we collect them, process them, and generate requirements from them. But that's also more or less just a task where we transform text files at the end of the day, and therefore, we can teach this harness how ideas look, how ideas should be described. And then the harness can take this idea, which we can give it, and break it down into requirements, and these requirements, which are written at a business level, we can then take them into review, look at them, see if they are fully written, if the acceptance criteria are correct, etc., etc. And then we can, from this entire set of generated specifications, feed them into the system, and then they will be developed accordingly. And the last step, Phase 5, is then Voice-driven Development, where I, as a developer, cringe, but it's unfortunately, unfortunately good. So, we simply go ahead, and this description of an idea is nothing more than us taking in information and then creating this idea document. And I've taught the harness how ideas should be formulated in our situation, and then I simply dictated the ideas to it via voice, and importantly, you don't have to imagine that I talked to it for four hours and it diligently carved everything out. I told it: "Hey, I want to work with you on ideas, and ideas always mean I have certain ideas for the system, certain concepts. You have an overview of the entire system, of all specifications, all tests, etc. I want you to critically discuss these functionalities with me," and this, it really hurts me to say this as a software developer, is simply incredible. I once drove to Vienna to a customer, and on the way to Vienna, I discussed the largest module with it, the workshop planning, and this planning took four hours. I discussed with it for four hours in voice mode with Claude how we would implement the whole thing. It kept saying, "Yes, but remember, and in DMDUL it's done like this, and legally you would have to implement it like this and that." So, in the end, at the beginning, I started in the first phase a bit as a developer. In the second, third phase, I then slipped into the role of the architect. Then in the fourth, fifth phase, or in the fourth phase, I slipped into the role of the PO only, and now here in Phase 5, I was pushed out of the entire software development into the role of the stakeholder, and a large part of the application that you will see in the summary is all developed idea-driven. That means many of you will ask, did you look at the source code, etc., etc.? Yes, I did at the beginning, but honestly, at some point, it was simply superfluous to look at the source code because I knew that with the tools I integrated, ReSharper and NDepend, the source code was always as I would have written it. So the structure is the same, the classes are the same, the components are the same. It adhered to all specifications, and where it still had a bit of creativity, it was at the source code level, and sometimes it didn't write it the way I would have, but it didn't write it badly either. That means I can be sure that if I work with this system as you see it here, that in the end, functionality will come out that is tested, that exactly matches what I specified. And the structure that is created in the background is more or less the same as if I had written it. Every single class, every single method. With the source code, it deviates a bit, but let's be realistic, if I were programming with other colleagues, I wouldn't always find the source code exactly as I would want it. But here, in terms of quality, the worst source code I've ever seen was a 2. The rest was always between 1 and 2. What is the result of all this that you have seen? So the project has become very large, very extensive, there's no other way to say it. So, in the end, let me quickly look at the specs here, I implemented 213 user stories, 4083 acceptance criteria. There are 10 services, i.e., microservices. In the end, 420 endpoints. In my MCP server, there are 108 tools that can be called. The system alone sends 135 emails via N8N. The emails look great, by the way. 88 workflows within N8N, and there are 10 databases corresponding to the services with 113 tables and 1300 pages of documentation. And the documentation, I must say, is also made super well, super structured, and above all, I've always had one of the big problems that the documentation was relatively quickly out of sync with the implementation. I built a daily nightly task for myself at some point after two days, so that every night on the build server, with Cloud code, a completely new documentation was created. That means I always had up-to-date documentation in the morning with the rest, and I looked at the source code less often than I actually looked at the documentation, which has become extremely good. But the things that are relevant to us, of course, are the results at the code level, and if we look at it, we have written 2900 API tests. So for the endpoints, for the workflow tests, 392, and we've ended up with 420,000 lines of code, and that's quite a lot. Of course, you might have projects that are significantly larger, but we still need to talk about the scope. And that these 420,000 lines of code could be created, that it managed to do that, that was clear to me. That was not a great surprise. The surprise was rather the quality, because in terms of quality, I've already said, for me, things like test coverage count, and we have test coverage in the backend and workflows, because these are the tests that I defined, of 85%. It's not 100%, but these 85% cover the most critical paths in the application and much more. The 15% that are missing here are missing solely because I didn't mock the workflows that are triggered by Office or Outlook, i.e., where the trigger comes from externally. It would also be possible. I could also reach 100%, but I told myself afterwards, okay, the trade-off between the effort that would have to be invested and what would come out of it, it's not worth it. Next point is the analysis of structural debt. That is, how exactly does the application that was developed adhere to the architectural specifications that I made, how exactly does it adhere to the design specifications? How exactly is the source code written the way I want it? And it's 0%. That means the entire application, every class, as I said before, every method, and most of the source code lines are written as if I had written them. And that is the actual, exactly this process.
of which I was talking. I can teach the LM to develop in the same way that I want or that your team wants or anything else. This means that all these excuses that always exist, oh, slop is still being produced, the quality, the architecture is bad. Yes, that may be. But if you learn and understand how it really works, then it's not a problem. And I had just mentioned the same thing for the documentation. It's really extremely good. And here, you see, here DDOG stands for Developer Documentation. So, I had a developer handbook built more or less. And SPT stands for Structural Dept, so here we're not talking about technical debt, but about structural debt. How close is the finished application to the architectural specifications? If we now deal with the costs, then the whole thing will become even more shocking, more interesting, more shocking. See it as you wish. For the entire application, as I said, I tried to take four to six months. Taking time means for me, I'm in customer projects, I'm on the go, I'm in reviews and have trainings and workshops etc. I managed to take 45 days completely off within this half year. Mostly at night, but these 45 days at 8 hours per day, one might have to add, I reinvested in this system. And if you now consider the source code volume and think that this came out with 0% structural debt and 85% test coverage, then I already notice in which direction it's going. These 45 days, if we take them with a full cost rate of €500 for a developer, that's a relatively common rate, then we land at personnel costs of €22,500. But I didn't just have personnel costs, I also used an LM on two different computers. I started with Codex at the beginning, but switched to Cloud relatively quickly because I find Cloud much better and much more accurate, especially for this guideline thing. And on the first computer, I used a total of 20.2 billion tokens, and on the second PC, 6.8 billion tokens. This seems enormous to all those who have dealt with something like this before, but you have to look at it, there are quite a lot of cached read tokens included. So they are not generated tokens at all, but the tokens came from the KV cache, a total of 27 billion tokens. And many of you worry that token prices will soon rise. That's why I calculated with the current token prices, not with the subscription, and I would have paid €21,000 for these tokens at Tropic. For these tokens and a small scene, I would have taken a Chinese model, for example, Deepse V4 Pro, also a very good model. Then I would have only paid €2,000. We'll also make a special video about that, because the worlds also diverge enormously there. But I didn't pay for either of these things at all. I had a Cloud 20 Cloud 20 Max subscription, two of them. Um, I kept running into the weekly limit at some point, which is why I used one for 6 months and booked a second one for another 3 months. So in total, I paid €1620. So, if we summarize this, the AI took 45 days, the whole thing cost me €25,000, 0% structural debt, 100% developer documentation, 85% tests. And so that we can compare this a bit, I've now simply selected one or two development teams from my clients. One team has five developers, 350,000 lines of code, that's their current application, it's about 24 person-years in scope. And the second team has four developers, 300,000 lines of code, and about 17 person-years in scope. And I had them estimate, I said, okay, look at the specs, look at the Defop structure, architecture, coding, testing, and documentation, and estimate how long it would take if we developed it in your team. And I got the numbers back from these two teams. They were 23 person-years and 15 person-years. And when I saw the numbers for the first time, I was still coding, I thought this can't be right, it's going to diverge too much. So I had Cloud and Codex estimate, they were at 20, and then I took the whole day and made my own estimate, and I came out at 18. And now you understand roughly in which direction it's going. If we average these four, we would be at 19 person-years. And if we now extrapolate this, these 19 person-years are 4180 working days, and at a full cost rate of €500 per developer again, we are at €2 million in project costs. With the quality that will now be determined, it will of course be a bit more difficult. I've now simply taken the projects from the two teams, because in this first team we have 14% technical debt, 40% documented, and the developers themselves say that our documentation is very far from reality, and 60 to 70% test coverage. The two teams also implemented both workflow tests and API tests, by the way. That's why I selected them among others. And the second team has 11% debt, 25% documented source code, they are a bit more accurate, and 55% test coverage. This means that if we were to average this, one could say, okay, such a team has an average of 13% structural debt, 33% documented, and 60% test coverage. And if we then compare that, we arrive exactly at these results that I was talking about at the beginning. We have 4180 working days, €2 million in project costs versus 45 working days and €25 in project costs, and the AI delivered significantly better structural, documentation, and testing results than the team would have. And throughout the entire experiment, I repeatedly estimated how much faster I am with the AI than humans. And at the beginning, I was sure, I would definitely be twice as fast. Then, in between, I had the feeling, oh, four times could also be realistic. And so I thought, yo, it could also be eight times. But it was completely different, because in reality, the factor was 93, 93 times faster with the AI than the human development team. And we can discuss this factor a lot now and say what it means for us, how it can be transferred to others. We will also do that in the next videos. Big promise. But first, we have to hold on to ourselves, we simply have an incredible speed boost with a simultaneous quality boost as well. And we can now discuss here at the factor of 93, yes, the tests are included, and the whole documentation is included, we don't need all of that. Okay, granted, let's say the factor is not 93, but only 60. And you can now say, oh, but the teams, they didn't have that much time, only had one day to estimate all these things, and maybe this and that doesn't fit, and others do what you want, let's go down to 25. This means we are now at a quarter, it would still be 25 times faster than a human team. And the application I developed is huge. It's packed with features. It's super tested, it's running super stable now. It's truly enormous, it's truly astonishing what came out in the end. What have I learned from this? A lot. The video list is full of what we will be doing soon. We will also do a higher frequency video on this exact topic. And on my newsletter, the first issue with Lessons Learned will be on Wednesday. Exactly on this project. You still have time to sign up. For those who sign up later, I'll add a checkmark with Cloud right away so that you can also have the last newsletter sent to you. First of all, if you look at the factor, let's take the factor of 96 again. I'm proud of it. I'm looking forward to this factor. I'm a software developer, it feels like my whole life. I started programming at 8 years old. Since I was 13 or 14, I've been trying to do this professionally. I eventually made it my profession. I've learned all of this that we're doing from scratch with a lot of sweat, with many nights of work, with thousands of books I've read, with thousands of experiments etc. I taught myself all of this. That is my capital in this job today. Nevertheless, this thing here in this example is 93 times faster than me. The quality is significantly better, and all for a fraction of the cost I would normally have needed. And on the one hand, that's of course extremely cool. It was a lot of fun to do, even all the nights until 4 AM, and I slept very little, you can believe me. It was a lot of fun, it's highly interesting, but on the other hand, you see how much faster it goes, and if you're as involved as I am in various teams, industries, and the industry, you can imagine what that means for our industry, what it means for your company, for your teams, for you as an individual developer. Honestly, that does worry me a bit, but on the other hand, the AI is here, it's not going away. We have to come to terms with this matter. This means you don't just have to accept this truth as it is. Whether it's 96 times faster for you or 25 times or maybe 150, I don't know, the experiments shouldn't show that, but we can be incredibly fast with incredible quality, and that already today in the year 2026. That AI is fast is a given, that you can also do larger projects with AI coding is also a given, but for me the real surprise and the real sensation was that it's not just so fast, but above all with the quality. That truly left me speechless, and I'm curious, let's discuss it in the comments. I'm curious if the quality here surprises you or if you're already working that way, if you're going in that direction, because this is the way we can develop today. But I believe for me, if I take all my clients together, I might have two clients who are already developing in a similar way as here, and then we have the whole spectrum, the first few prompts, who are already doing a bit more. And this, I must say, is sensational. Does this mean everyone can develop such an application? Everyone can develop an application with this quality, with this speed? No, the quality in the end is so good because I put so much time, so much energy into the harness, because for many years, decades, I've been doing quality in architecture. I know how to describe something like this for a team, and it's not far from describing it for an AI. And therefore, if you want to deal with something like this, learn the fundamentals. As I've always said, AI is a super, super tool if you have understood the fundamentals of software development, software architecture, and that is also super helpful with AI. At this point, you saw the overlay, I've switched the workshops to AI for Developers, Requirements with AI, and AI Strategy for Companies. And there we deal with exactly how you as a developer build something like this, how you can develop requirements for systems, but also how you as a company strategically reposition yourself with this AI, because the later you do it, the more difficult it will be to catch up later. It's already a bit late, but you can still achieve it as a developer, as a company. Therefore, let's tackle this together, not only in the workshops, but also in these videos. We will start again earlier or more frequently with the videos from next week, then we will work through it. If you still have specific questions about these things, what we should discuss in which video, you can also gladly write them in the comments below, that you would like to see the harness or anything else, then we will do that too. That was my experiment. I'm still shocked, I tell you honestly, and I'm curious, really curious, what you have to say about it. I wish you a nice day, have fun working, see you in the next video, and bye.