Transcription
Hello everyone. Welcome back to day five of building a software factory. We're building memo, a Notion-style note-taking app, entirely built with AI agents on the Owner platform. Today, I'm joined by Philip, Owner CEO. Philip, it's great to have you.
Um, let me quickly maybe summarize, uh, what's happened this week, and then we'll jump a little bit into, uh, Philip's view of the project, which is a very interesting one. So, um, we are now halfway through the project. Uh, day one, we've done a setup of the concept of the stack. Day two, we've gotten into the first scaffold. Um, got the first pre-review running. On day three, we jumped in with Chris, uh, on a CTO, talking all things about the product spec, how the, uh, spec turns into a working product. Uh, merged the first 50 PRs. Uh, yesterday with Lou, we went a little bit under the hood, talked all things implementation, review loops, and today we're going to go into bug fixing.
Um, Philip, you've actually been stress-testing the app already. Uh, what's your take so far?
Yeah, of course. I jumped on the app immediately to see if it worked. It did work, but of course, there were some, some, uh, bugs that I found. Um, so obviously, some bugs like around functionality blocking things that you would just expect to work. For example, the workspace creation didn't work in the first place, or inviting others didn't work. Um, hyperlinking didn't work, and those are things that, you know, just need to be there. There were other things like UI bugs that were, were a bit weird or gnarly, but, you know, not blocking. So mostly nits, like some, some, I don't know, things were misaligned, or like the checkboxes were invisible, or like some odd drag-and-drop behavior that were not blocking, but just a bit strange. And, um, yeah, I'm really, obviously, I filed them, but it took some time, right? I like took a screenshot, shared them in a Slack channel, like, you know, gave some description of what was happening. And, um, I'm really curious to, uh, dive in today to see like how much we can limit this, and, uh, also to understand like, how can we avoid these things in the first place, that, you know, as few bugs as possible actually, like, are in the production, um, um, app, and if they occur, how can we proactively spot them, uh, and then also fix in an automated way? And, um, even if I report this manually, I'm curious like, you know, what you did with my reports in Slack, like how you fixed them, uh, in a way that hopefully didn't cost you that much time. And, um, and then in the end, like, we're building this factory, which doesn't mean that there's like, putting out the machines and the conveyor belts and then we step away. Like, you know, you need to like review them and all of that. So, how can we now iterate on this factory so that it recursively becomes better, more automated, and all of these things like diminish over time?
And one thing I will say, I did some, I did some, some, uh, prep before I came here, uh, because I was curious like how much, uh, was actually going on in a truly automated way. So I checked the, the agent executions and automations usage, and I think there were 14 automations set up by you. 11 are running, um, um, regularly, and 2,000 agent executions have run, and 98% were, were automated. So, like truly running in the background, like not triggered by you manually, but just like, you know, firing on an event or on a schedule. And 2% were of those 2,000 agent executions were manually triggered, which is probably like you setting up the harness and like, you know, iterating on the system. So it is almost like a fully automated factory line already, which is great.
That's crazy volume. It's, uh, it's insane to hear the numbers and to see what's possible in such a short amount of time. But obviously, now we are in the phase where we're going to have to improve, uh, fix the bugs. And as you said, right, you've filed quite a few, uh, and flagged quite a few. We've given the system some time to see what it catches by itself and where it got stuck. Um, from the 12 you have filed, um, I think about six or seven were already fixed by the, uh, by the system itself. So, for example, the hyperlinking was actually something where, um, the error was flagged into Sentry and then fed back into the bug reporting, uh, fixing that, ensuring that you can paste, for example, a hyperlink directly onto a word rather than having to go into the bubble menu, and then fixing it. But there were obviously some bugs as well that were not fixed. Uh, one of them I want to address right now. Uh, so let me share, uh, one of the ways of fixing a bug, uh, with the system. Um, we're going to make use of the Owner Slack integration here, and we're going to address the bug of the placeholder in the empty editor being unaligned with the with the title. So, we just go into the thread here and quickly going to tag Owner to fix the placeholder, and then we're just going to set it off. So, it can be as simple as that. And there are a couple of options in the system how we can flag bugs, uh, and how they will, uh, fix them. So you can already see that Owner is now starting an environment. And, um, I will quickly share a different screen so you can actually see how the link works. So, sorry, give me one second.
And this is really, uh, I think a nice, uh, way of fixing bugs that you want to get rid of fast or are flagged by your team members. So,
Yeah, I can, I can, um, so from, from our own product owner, we have a feedback product channel, uh, that has usually been, I mean, I don't want to say graveyard, but it has been like a place where like lots of product feedback has been shared, right? And it's like, it usually has been like very tedious, like bringing this into Linear, and then from Linear, bringing this into Owner. So what we've seen with like the Linear agent coming out was like that people tag Linear and say like, create an issue for that, and then at some point, somebody took it in Linear and then actually solved it. Right now, like this is almost skipped, and people directly tag. So our Head of Product Engineering, Matt, I think his morning routine is like going through the feedback product just to be like, add Owner, add Owner, fix this, fix this, fix this, and just like fires off like a dozen, um, agents doing these things. So it's great to see.
No, it's great. And, uh, we, we can see that right here, right? Uh, Owner spun up an environment to to address this bug. It's going through the Slack channel, reads out the, uh, the description, reads out the, uh, attached screenshots, analyzes what's going on, and is going to fix the bug. On Slack, we can see, um, the, the view of this automation as well. So, um, it's like a two-way link, which is quite nice. Obviously, as you just said, right, there's a, a level on top of what you can do is actually create an automation to automatically go through the Slack channel rather than you having to, uh, tag on different issues. So, depending on your setup, if you have a dedicated channel for bugs, uh, you can have an automation just run on it, um, whenever there is a new bug being opened and really have this autonomous flow, or you can tag Owner in this case, uh, to go ahead and file the bug. So, we'll go back later, um, hopefully see that this has been working, and review the outcome of this. Um, I think another way that's cool to to see, and we're going to quickly go into, uh, the, the second way, um, that I like to use right now to flag bugs that are, for example, you know, usage-related, are, let's go back into the GitHub view. Just raising an issue here, in this case, a bug report. One thing I've noticed that image upload is not working. So I'm going to be deliberately quite lean with my feedback here, also to show you that the automation will actually investigate, add more detail to a bug report that is reported by a user, in this case, and then, hopefully, fix this. If it runs into, uh, something that it can't fix, it will ask for human input. However, I would expect that this is something that the agent can fix itself. So, uh, give me two seconds to, uh, write this out.
So our expected behavior is obviously that we go into the editor, are able to upload an image. Currently, what happens if you do slash image, you can select your image, nothing happens. So there's something broken in the backend. And we're going to try to fix it. You can use the,
And what's the advantage of like creating this as an issue rather than like in Slack? Is there any difference?
The difference is mainly, uh, who's your team, right? Who's on your Slack? It's probably your internal team in this case. Um, this is going to be open-sourced, right? And probably after this, um, week, we're going to open up the issue list for, for everyone to add issues. Um, and even though if you're not in the Owner Slack, you will be then able to add bugs, open up feature requests, and then we will run through them automatically. Um, naturally, there is like always a trade-off, right? Now, because we are running everything autonomously, um, and it's internal only, you know, the risk is fairly low of a request being malicious. But naturally, in an open-source environment, you then want to have another loop and actually assessing whether the request is, uh, malicious, or is actually, uh, well-designed, or, you know, actually adds to the feature, or has been addressed in the past. So, let's quickly go in here. Um, go to editor, try and upload image, and then image gets uploaded and shows in editor. Priority in this case, two. And then we just create it. And again, we'll let this run now. Uh, we will go back later and see this will be picked up by our feature planner. So, we've discussed quite into detail, uh, how our automations work. Um, our feature planner currently works on a 30-minute schedule, running every 30 minutes. You can also, uh, trigger it manually. So, we will double-check if this has been picked up in a few minutes already. Otherwise, we'll just trigger this, um, and then we'll see the automations working from our feature planner. Will be, um, basically triaged, given some detail, and then the feature builder is going to pick it up and ultimately then pushed through up to the PR review, which will hopefully close this or raise some more points to, uh, improve this.
Great. So now we've talked a little bit about, uh, how we can fix bugs. And as you can see, there is a bit of manual work currently involved. But where we want to go to is, uh, that these bugs don't even come up in the first place. And, um, one way of doing this is obviously what we discussed with Chris on stream: you need to have a great foundation, right? And I think one of the issues that we had, uh, with our repo in the first place, our product spec was detailed, but not detailed enough. And this leads ultimately to the agent coming up with an outcome that is suboptimal. And one thing that we have here is the, the quality.md file. The quality.md file, um, is another file that the feature planner works on. So every time it goes through the backlog of our issues, um, it decomposes, uh, the, the current issues, but also decomposes the, the product spec to check if there's any gaps. As we build out the repo, we actually grade the different parts of our, uh, of our architecture. And you can see that currently it says everything is on an A, which is another thing to, to obviously mention, and you have to scrutinize this self-grading as well. But if we compare it to the history, so if we go back a couple of, uh, yeah, not even weeks, a couple of days, couple of hours, um, and compare to where we started from, uh, there were a lot of, uh, grades on a lot of sections, not even graded, or graded with B and C. They were then upgraded to A's and B's, and as we went along, going to A's, right? And this basically means that we have a loop that is self-improvement, self-improving on the product checks, uh, if our spec has been delivered to. And again, the more detailed you are with the product spec and the, the closer you can communicate your vision, um, the better your outcome will be, and the more efficient your factory will run. I think this is a quite nice analogy to to an actual factory, right? Your processes, uh, need to be well-defined, and, you know, your production line needs to be, uh, set up in an efficient way to get a good outcome. Um, if you're not clear about the product that you want to build, you can have a great production line, but it will not get you a good outcome.
So Zach, a quick question here. So the quality.md file, is this something that you create in the beginning, you just like test against, and your your goal is to just, uh, increase the grades, or are you also iterating, uh, on the quality.md file itself to basically improve the blueprint of what this should look like as you build out the product?
Yeah. So the quality.md file was part of the harness, uh, initially, and it will be, um, updated by the agents as we go along. Right. Um, currently, the main, uh, updating agents, and you can see, initially was intended for the, um, automation auditor to, uh, write into it. If we go into the current version, we actually see that, uh, it's mainly the, uh, the feature builder and the feature planner now, as well as the bug fixer updating this, this file. And how it works effectively is, um, there's the, the planner and the builder updating the grades as they add more features or bug fixes. Um, but also, um, we go and look at the staleness of the actual issues. So we decompose, for example, the individual grades, and occasionally, whenever the backlog is empty, look if there's any gaps and figure out whether the grade that has been given is actually representative. And this is, I think, where if we look at how we can improve this in the future, add more scrutiny about, um, the, the grade. You know, and we probably want to be a little bit more, uh, scrutinizing of giving like an A, even if we know, okay, there's probably not like an A grade to be given here. So this is a future improvement that we can add to our automation loop, which is also quite, uh, good to have.
So I'm, I'm curious on the, digging a bit on the grading scale because, um, it seems like these are the success criteria, so that the agent knows like what each grade means, which seem, if you go up, it seems fairly, um, fairly lightweight. So I'm wondering like, is, can you talk more about like why you chose like a letter-based grade versus like a score from zero to 100, for example?
Uh, and then also how to, um, how to set up these success criteria, basically, how much should you go into detail from your experience versus like keep it light?
Yeah, I think on, on the, on the grading itself, if you have a scale that is too big, um, you have a bit of randomness with the actual score. So a scale that is a bit, uh, smaller will give you a more accurate, uh, representation of where you are. And, you know, whether you use like a zero to 100 and then you give it, you know, five sections from like poor, mediocre, good, um, you know, very good, and excellent. This ultimately translates to, you know, in this case, a four or five-letter scale. Um, I agree that the actual, uh, explanation of the individual scales, uh, should be more in detail and have more examples, right? Because if we look again at how the LLM will work, the more accurate and specific you are with your examples, how you can get to an actual grade, the better your outcome will be. And this leaves a lot of room for interpretation in this case. So this will be something where we probably see, uh, some work going into the next few days and, uh, hopefully see some improvement again on the, on the outcome, which will be, yeah, an interesting one. I think another, um, very interesting, uh, example, if we look at the ultimate loops, uh, is how we work with Sentry to do monitoring. And this is something, uh, we've already touched base, um, briefly earlier, right? We've explained that Sentry is our monitoring, uh, loop for the actual app, and there are some bugs that were found directly with Sentry. One of them, for example, was the workspace creation. So the agent has actually, through testing, realized that there is a workspace creation issue. Uh, fixed the workspace creation issue. So users were able to go into the UI, see the workspace creation, uh, interface, but then couldn't create a workspace creation because there was a, a role security rule preventing new inserts. So this was flagged to Sentry, and we can see now one of our other automations, uh, working on this. And I've brought up, uh, I think you cannot see this yet. Now you can see it.
So this is our incident, uh, responder automation, and this is the log. So, uh, with all automations that are running on Owner, you can always go in and look at the conversation of the agent. Um, so you can see at the beginning, you'll have the prompt. In this case, we will connect to Sentry and get all the issues that are, uh, flagged. We then check if the issue has already been named in one of the bugs that we had. Uh, and in this case, we found some new issues that we want to fix. Um, as you can see, the agent then goes to work, um, figures out what's wrong, and then opens up, uh, issues. Uh, and this is representative of one of the issues that we have, uh, we have fixed. You can then go back into GitHub to see the logs of whatever we fixed. One of the examples that I just mentioned for, uh, was the Supabase insertion rule. And this is not what we are currently seeing on screen, but effectively, the agent realized that, okay, there's a rule on the database requiring us to change the RS rule. Uh, in a production environment, you would probably have some, uh, yeah, higher risk or lower risk appetite depending on what you're building. So, this is something we discussed yesterday with with Lou. There are some PRs that are low risk that you want to auto-approve. Uh, we currently run everything on auto-approve. In a production environment, you would probably do something that changes the database access, uh, with a manual review. But again, this is the great thing about the software factory, in your setup, you can really decide what you want to auto-approve and what you want to just, uh, yeah, have your eyes over or your team look over. I think this is also, once you start getting more users onto a platform like this, um, and you have more edge case testing where you know, um, you will get to like a really great flywheel because the more testing you have, the more coverage you'll have, the more edge cases you will encounter. And again, a lot of the tests you would want to have before you deploy, right? We're talking about like end-to-end testing, unit testing, integration testing, which is all things that you can set up. But you will probably never reach 100% coverage. But if you get to a point that even if you catch a bug that is in production and you can automatically fix it, um, you will get to, you know, higher customer satisfaction, a better user experience. And this is also something where I just think this is like a really great example with the software factory, uh, where we're getting to.
Yeah, and I think it's, it also shows like again, the goal is not to oneshot something. I know there have been like Cursor, for example, have been doing a self-driving codebase with, um, the browser that they have been building, which is more like, we give it a go, it oneshots it over a week, and then there is something. What we are trying to do here is more like, you know, we set something up and then we iterate on this factory so that it is in the end, like self-driving and like self-maintaining, like heals itself with the goal of like pulling the human out. But I think it's also really important that, um, you make sure that what you're building here in a self-driving way actually does the things you want to do in the way that you expect them to with the quality level. Because otherwise, yes, you do have a bunch of automations that's doing things, but, you know, everything looks different left and right. In the end, you probably spend more time figuring out why. So I think this iterative loop of, uh, you know, going in, like figuring out how it works, improving it, automating more and more, makes a lot of sense. And in the end, also with this, it's all about, um, yeah, spotting things more in a more automated way, um, earlier on, or like preventing them, but then also increasing the time to reaction. Right? So Sentry is nothing new, like everybody uses it, but like the way you respond to it is different here, right? Whereas the, the most annoying thing is, you know about a bug and it sits somewhere in your backlog because you do have more important things to finish, and like more and more users stumble into this thing. The Sentry, you know, logs keep coming in. And I think what the, um, incident responder, um, um, automation does is like it shortens that time to action so much that usually a user wouldn't run into the same bug. I did because I like jumped on the app immediately when it was available. Um, but before you could fix my bug, Sentry already logged it, and then Owner fixed it. Right.
Exactly. And this is, I think, the, the great thing again, not, not all of them, but, uh, and I'm, which is great. And again, this is also kind of depending on your setup. You know, how do you surface your errors? Uh, if you, you know, have a setup that swallows some of the errors, um, you will obviously not catch them. But if you set it up in the right way, uh, so that your observability tools catch everything, uh, you'll have a faster feedback loop. And, uh, naturally, this is not just limited to Sentry, you can use whatever observability tool you want to have. You can also observe your, your database directly. So we currently have a database endpoint, endpoint that we, uh, that we monitor for database health. But you can obviously, uh, you know, look into your Grafana, into your Prometheus graphs, um, and also link that into your software factory to ensure high quality. Maybe to quickly touch base on,
Let me ask one question, sorry to interrupt. Um, so, uh, we spoke about like, you know, increasing the test coverage and like, you know, using Sentry and hooking that up. Like, what are in that system to prevent bugs and to spot bugs, like what are the biggest blind spots that you're that you're still seeing?
I think coverage, right? Um, you need to ensure that everything is wired up correctly, that all your, uh, edge cases are covered. So, for example, if you think of Sentry, you can, uh, connect front-end, you can connect back-end. Ideally, you want to have everything connected, and you want to make sure that depending on where your architecture sits and where you deploy your app, that you actually have a way to capture everything and make it visible for your automations, right? Because if the, uh, automation can't see your, your bugs, then you can't really, uh, blame it for not fixing it. And I think this is what we, what we also saw in the first few days. One of the issues that we had is, um, with the setup of how we queried Sentry, we're actually not looking for all the different exceptions that were thrown out, um, because we were looking specifically just for errors and not just for warnings. And that caused us in the first two days to not see all of the Sentry errors in the dashboard. I think a good example again, how this was fixed, going into Owner, tell Owner something is not right with the automation, and it investigates and then comes up with the, the improvement and the solution itself. So while you still have this manual oversight, and again, we, we talked about it in the past, we have this on-ramp phase still where you build your factory, and this like building phase requires a little bit of oversight, and you probably want to have this manual oversight in the beginning. You then go from a phase where you are very hands-on to like less and less hands-on. But, um, yeah, if you make sure that everything is set up and you keep an eye on it, uh, you get to a point where the factory runs itself very fast.
Um, and one other related question. So in the beginning, I, I, I talked about like that there's like two broad categories of bugs that I encountered: like functional bugs that probably like Sentry throws an error for most, and then other bugs that are like just like looks weird in the UI. It doesn't look that nice, shouldn't be like that. Things are not aligned. Like, how, how do you think about automating the, um, detection of these kinds of bugs? Like, what ways are there, and like, what are the limitations that we currently still have?
Yeah, I think, um, just to not jump too much into the UX/UI because we're going to talk about that on Monday. But generally speaking, it comes back to spec, right? And it comes back to how specific are you with your vision when it comes to to the product design and product feel. Uh, you need to be able to communicate it, uh, very clearly. I mean, it's nothing, nothing different to when you work with a human team, right? If you, as a product owner, can't communicate clearly what you want for your product designer, your product designer will go out and be probably quite creative, and it might be a great design, but it's not aligned with what you want as a product owner. And the same thing happens here with the software factory, and probably is even worse, uh, with, with a software factory if you're not quite specific about what you want and what your priorities are. And for example, you know, things like an alignment of the placeholder should be easily, easily caught, but it was nowhere in the spec. And while you think it's trivial, if you're quite specific about it, you can probably catch this before it ships. So those are really the, the loops that we want to look at. And, um, yeah.
And I think,
Um, a couple of changes, um, that we're also doing at Owner is like, you know, the agent is able to also take screenshots, record a video, like, you know, when it ships something, like it can actually go through, you know, testing. It has a browser built-in, and then review that to see some of these things that maybe didn't, you know, wouldn't throw an error, you wouldn't see it in the code necessarily, but you see it visually. Um, so I, that also, um, helps a bit.
Yeah.
Um, cool. So I think we're almost at time, but one thing that I found interesting, like, um, I said in the beginning, there were 14 automations. Um, 11 were running. One that caught my eye that hasn't run yet, I think, was the automation auditor. So would love to, you know, for you to talk a little bit more about that.
100%. So, um, we touched base on the automation auditor, I think yesterday, very briefly. The idea for the automation auditor is really to be our improvement loop for the software factory. The idea for it is really to run, look at what the automations are doing, looking at the output of the automations, and then, um, suggesting improvements, right? So rather than me going in and saying, hey, somehow our incident responder is, uh, saying there are no issues in Sentry, even though there are issues in Sentry, um, ideally, I shouldn't have to manually prompt an improvement there, but the automation, uh, should do this itself, and the auditor should catch this. Now, if we look at kind of what we just mentioned about the on-ramp, and you want to have some kind of oversight initially while you build out your factory, it probably makes sense to wait for the automation loop and the automation auditor loop to come into place because we see that the automations are not perfect out of the box, right? So we take some time to get the factory right before we then switch over to complete auto-run mode. Um, if you don't, you basically have two possible outcomes: either it goes great and it actually improves itself, or it goes complete haywire and it goes into a direction where everything that you produce from there on out, um, goes kind of bad. And if we go with the approach where we have a bit of the human on the loop oversight, do the steering, and we get the software factory to a place where it produces great outcome, then we can actually go into it and say, does any automation auditor loop improve or, um, yeah, make our output worse? And this is, I think, the, the idea, and we will see that next week, uh, when we switch it on and see if, if it goes, yeah, for the better or the worse.
Cool. All right. So, I'm curious what, what, uh, what's happening to the bug that we tried to fix in the beginning?
Let's have a quick look. Right. So, now we have four issues opened. Um, two of them are regressions, which is also great, right? So this is coming back from our testing, that some of the PRs actually pushed something, um, that is not working. If we now look, for example, into the, um, image upload bug, the one that we manually fixed, or not fixed, but manually, uh, posted, we can see that the description was, um, enhanced by our automation. It was automatically, uh, put into, uh, the backlog and is now being worked on, depending obviously on the other priority bugs that are in the issues. So if we look into it, um, this is a priority one bug. So priority one will obviously take precedent over priority two, which is why this hasn't been fixed yet. If we quickly go into our pull requests and look what's been closed, um, we now have an alignment on the placeholder, and the page icon emoji picker has also been shipped. So maybe let's quickly just go into the app and look if our alignment has worked. If not, then we will obviously go back and try and fix this. So let's open a new page. As we can see, this is an interesting one because it didn't quite fix what we wanted. Uh, it did it in a different way, right? So we originally had the, uh, the drag handle to, um, overlaid by the placeholder. Now, it has moved our editor component rather than keeping the editor component in place. So we'll go back and fix this. But this is again, one of those things where you can be quite specific what you want, and we probably should have been more specific about the way it should fix it. Um, because now we have an outcome that is somewhat in line with what the original bug report was, that the drag handle shouldn't be overlaid by the, um, placeholder. But what it didn't get right is the alignment of the placeholder with the, with the title. So we'll go back, fix that. Um, you'll be able to see all the fixes that the, the system does by itself, and, uh, yeah, hopefully over the weekend, we'll have a good outcome then.
Cool.
All right. Great. So, this was week one. Um, we are halfway through. So, if you have any questions, any suggestions, what we should add, uh, if you discover any other bugs that, um, we haven't addressed yet, uh, please feel free to, uh, uh, let us know. You can find us on on X either through @owner_hq or @swactory_dev. Um, our website is software-factory.dev. You can find the product there. You can find all the links to the past videos there. Um, and I'm going to be excited to have you all back next week. Tune in for some more software factory, uh, ideas, some more updates, and hopefully a fully working product. Then, Philip, thanks so much for joining.
Um, pleasure. Thank you.
Nothing else, then I think we wrap up, and I hope you have a good weekend.
Bye-bye.
Cheers. Bye.