📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Stop Shipping Slop: The Future of AI Code Is Quality

OpenHands57:02

Transcription

Yeah, thanks Jamie and thank you all for coming. So, um, as she mentioned, I am chief architect at Open Hands. I work for a startup now, but I've worked for legacy organizations as well. Um, you know, which I, I try not to use "legacy" as a, as a pejorative term. I don't think, uh, in software engineering, it's a good thing that we, we use "legacy" to mean the bad code because the code we're so bad at keeping code good.

But, um, I am, I am excited that software quality has become so, um, important, so relevant right now with, um, with the changes we've had with coding agents, because software quality has been one of the huge focuses of my career over the last 20 years.

Now, let me zoom out for a moment to even, you know, even before my time in the industry. In the mid '90s, there was a popular consumer operating system that would notoriously frequently crash and display a screen like this. The screen was blue. It was unofficially called universally, and someone can comment it. Someone could. It was the blue screen of something. Do you remember what it was called? Was it the blue screen of mild inconvenience? Was it the blue screen of, you know, oh well, I don't mind so much? I'm assuming that you have guessed by now, it was the blue screen of death. Not because that's what it said in the manual, but because the public universally chose that term.

Now, if you're coming here, you're probably like, you know, a very software-centric person, perhaps. But I have talked to people who are, you know, grand specialists and hair stylists and people from all walks of life who are not like, you know, um, enthusiasts about programming history. I have never found a single person who did not immediately remember, "Oh yes, blue screen of death. That is what we all called it." Um, that is because software should work when you're trying to get work done, and the software explodes in front of you. It's not a good thing. Software should work.

And I am, I'm lifting that slogan actually from a, uh, a conference I was invited to speak at. If you're in the Midwest in July, I would love to, uh, to see you there. But, um, I think this is a very good message for the moment. We'll get into why.

Um, we are, I think, in the industry, we are recovering from what I've seen as some very destructive sentiments. Um, for instance, let's talk about possibly the, the worst slogan in the history of bad slogans. "Move fast and break things." Now, this was made famous by, by this clever fellow here who, um, started a popular social network you may have used. Um, and people still repeat this as though it is Meta or Facebook's slogan. They stopped saying it in 2014. It hasn't been their slogan even internally for over 12 years now.

And it's easy to think about why. Look at this, you know, consider this car accident here, this highway accident in which, fortunately, no one was, was killed. Um, are people moving fast now in this situation? No. Um, everyone's moving slowly. You know why? Because everything is broken. The mentality of "move fast and break things" just takes us to a world where everything is moving slowly amidst a bunch of broken stuff. This is not where we want to go.

Um, unsustainable, unsustainable solutions do not count. If we can't make money today and make money tomorrow, we haven't really solved the problem. If you can't give users software that works today and will work tomorrow when you make your next changes, it doesn't count. So, there isn't really a dichotomy at the end of the day over time between quality and productivity. Quality is productivity. Something is either delivered with quality, or it is not delivered at all.

Now, I, uh, I make this case in a little more detail in a recent blog on the, the Open Hands blog called "Velocity is Dead," where I, I comment around about, um, a recently viral, uh, compiler that Anthropic put out that's kind of interesting. They've generated a C compiler using, um, you know, various agent orchestration and, and, uh, compiler fuzzing testing techniques. I was doing actually a similar experiment that I, I briefly talked about, but I, I make the broader point that quantity is a solved problem at this point with coding agents being able to run many of them, generating large code bases so quickly.

You know, it's interesting. If you're still fascinated by that, get it out of your system. But in addition to these experiments in quantity that we're having, like that 100,000-line compiler, or my 40,000-line version, the, uh, million-line web browser and spreadsheet that Cursor, uh, put out experiments about, it's a, it's a staggering amount of code, but it's zero users. And it's not only incidental because it's an experiment that there are no users of these things. Like, if you boot up that spreadsheet, it is not only like no one using it, it is unusable.

And so what really is there to celebrate once we've established that we can run coding agents that long and produce a pile of code that is at least plausible, where do we go next? And that's where I'm saying we also need experiments in quality, audacious experiments in quality. And we can do this. We can have nice things. It is not the case that it's just impossible to use coding agents to produce high-quality results. But it is also non-trivial, right? It, it is an acquired skill. It won't happen by default right now.

But I would like to convince you that even on the reluctant side, where you would expect people to be the, the, the most cautious to be using judgment, not just, you know, me who's employed by an AI startup. You have, uh, you have very credible people saying there is a path forward. I'll read a, a quick quote here by Emily Bash, who I, uh, I interviewed on our YouTube channel about test-driven development recently. She says she's spoken to half a dozen different technical coaches in the past few weeks who she knows are excellent engineers, and they've all said they're now writing almost no code by hand.

Um, you've also got Brian Finster, who maintains minimumcd.org, which is, I think, the most authoritative definition online about continuous delivery and how to achieve it. He says, "Whenever someone delivers code, they're shipping their quality process. It doesn't matter if they're coding by hand or using AI." So what he's advocated, he's put out a lot of stuff in detail on this website, minimumcd.org, about how to create a pipeline that is so robust that you can take advantage of this, this output, this copious output that, uh, that coding agents are able to give you without creating just bottlenecks otherwise in the, in the place.

Just one more quote by him. So the reason he pushes so hard for continuous delivery is it's the most effective method he's ever seen for uncovering and forcing the improvement of systemic problems. So, I, I like to say continuous delivery is a forcing function. Every reason why you can't regularly ship, why you can't regularly, if it's not in front of users, it's at least, uh, as a fully integrated system where the business can, uh, can evaluate it. Every reason you can't do that points to some issue with your process. If you can't ship every day, I mean, you may be correct that you can't, but it points to a problem.

So, you know, this is where I think the, uh, the DevOps movement got so much right in terms of the focus on rapid delivery, but not necessarily for its own sake, but because it's such a reliable, uh, way to surface the dysfunction. And doubling down on those lessons is where a lot of people are seeing, "Okay, this is how we get the most out of coding agents."

Now, you have Paul Hammond, who I'm going to reference again in a moment, um, who is, uh, you know, a huge practitioner of test-driven development. He says XP, meaning extreme programming, which includes test-driven development and pair programming and some other practices, says, "This is the missing piece for AI-assisted development." Now, these are, these are highly credible people that have been, um, you know, advocating for development practices for many years and are not the type to just jump on a trend. You know, they're not being, they're not being paid off by Silicon Valley. Like, they're saying these things because they believe them.

Um, and when we see, um, when we see them reporting what they've seen, you know, uh, I, I think, I think it gains a lot of, uh, it gains a lot of, it has a lot of credence that there is a path forward for disciplined development with AI. This isn't just a matter of, you know, posting giant demos that turn out to be a bunch of slop, you know.

So, on, on Paul Hammond's, um, GitHub profile, you can find some agent skills that do test-driven development and a bunch of, um, mutually complimentary practices, and he's put out some material on, "This is how he uses them. This is how he integrates that into his feedback loop." It is, uh, it is kind of subtle to do, but I, I have found that to be very valuable. The way the, the skills he's put together, uh, work with each other and, you know, with his process. There's, I don't think any one definitive way to do something like test-driven development, uh, with coding agents. It's mutating very quickly right now, but this is one crystallized format that I, I think if you're interested in, "Oh, what's the exact prompt? What's the exact skill file I should look at applying?" This is a very good place to.

But, um, in general, you know, you probably are coming today say, not just for me to convince you that we need the anti-slop. We need, um, to focus on high quality instead of just velocity now that we have coding agents. But you want to know how. And this is where I've really challenged myself to give you the most useful message I can, because there's so much to it. And it depends a lot on your situation, because we're really just talking about how to create and sustain high-quality software. But I think I can give you six words that are going to be very high leverage for you right now: "Tell your agent how to test."

If you ask today's coding agents, "Could you make me this feature?" most likely they're not going to test it at all, or they might test it themselves. They won't necessarily give you the automated test that your codebase will require, uh, sustain that feature. So people have figured out, "Okay, I have to ask it to test it," but step beyond that, they're not necessarily going to test it by default in the way that you need. And agents, they exist. You know, we have agent loops because we have environment feedback. So you're really maximizing the better the feedback you can give them into if they've done a good job, um, the further they can go, right? The better job they can do for you. If you're shaping one thing about the implementation, it would be, "How should you validate that you've done a good job?" And so if you can get in the mindset of thinking about testing strategy very early and very actively in, um, using a coding agent to develop something, you're going to see great results from that. So that's, if you take one message away, it's, "Tell your agent how to test." We'll see some examples of that.

Recently, I did the Ralph Wigum webinar, which is more about patterns of using agents than this, you know, kind of, uh, testing point that I'm making. But the reason that Ralph Wigum works so well, aside from context management, which it does well, is also that it emphasizes for every step you take, you know, "How are you validating it?" So this, I'll go into, um, the example that I happened to run in the background while I was doing that webinar. This was a little app that I made with, um, a series of invocations. It was one-shot in that I just ran a script over and over and over in the background. And, uh, the Ralph Wigum pattern is, uh, just calling a headless agent over and over on the same prompt. That prompt says, "Do the next step using these instructions," which include how to test. And this was the app I built, you know, during that time, you know, in the background. It was a Lisp interpreter that runs inside of a web page. It had an editor using the code mirror library. It had, um, you know, it would underline things and show you the error if there was something wrong with it. And, you know, not bad, a, uh, a minimal but usable programming language and editing interface, um, in, uh, absurdly little effort. It was also well-tested. The, uh, coverage reported by this is actually undercounting because it's not including the UI, but I asked for separate tests of the UI that just aren't included in unit test coverage. So, um, and, you know, we could, if there's questions on it, I'll show you the actual code. I contend it was not only like tested to the point of coverage, but it was tested well, and it was built well. It was built about how I would have wanted it to do.

Um, I'm going to go through the prompts that I use to make just a point on how I tried to find high-leverage instructions that would, um, give me these outcomes that were the kind of outcomes I wanted. Uh, so excuse me for, uh, you know, reading a large part of a slide here, but, um, this was the seed prompt in which I'm giving it a spec to go from. And so, "Design a spec with a plan to make a Lisp interpreter for a useful subset of the Scheme language. It should be implemented in modular TypeScript and build to single-page HTML with inline JavaScript and CSS. That page should have a simple GUI with editor in CodeMirror for syntax highlighting, which is a library that does that code editing interface in your browser, and output and errors should display when 'Run' is clicked. Source errors displayed inline. Make it look cool and hacker-like."

So that's basically all I told about what I was trying to build, and I, um, a couple of notes from this, uh, you know, this kind of technical guidelines section that were particularly high-leverage. We have how we're tracking our task in the Ralph orchestration. That's not really relevant to this one, but I said, "There should be data-driven tests for all interpreter features in the form of YAML files with input and output saved in this folder." So I happen to know that just for the kind of code that we're working on here, we're doing a language engineering task, that storing, you know, inputs and outputs in a flat file, these data-driven tests are a very effective way to, to test that and make it working separately. I've got a UI. So I say, "Okay, you should test that using Playwright." And, um, if there's, uh, if you feel you need to update the spec, I'm giving it leeway to do it. You know, because it was a one-shot example. I'm, I was not very picky about the details. Your needs might differ, but I'm picking a few elements here of like, "You should use data-driven tests separately. There should be a presentation layer, and it's tested in this other way." Um, these little points about testing strategies just shaped the whole process and, and made it go well. So that's an example, and in your app you're trying to build, there will be other examples that might be relevant one way or another. "Tell your agent how to test."

We'll try to do a, a live example here. Um, and this could be one that you might find interesting to practice on. So, we'll, um, we'll see how this goes. But this is, if I search for "realworld," it's the mother of all demo apps. This is several years old. Um, it, it unintentionally is a really cool, um, useful example for trying out coding agents, but it was kind of made to showcase different, you know, web frameworks. So that's, um, it's kind of a clone of Medium, except in the just "realworld", you know, repo, you're not going to find actual implementation of it. It is just the specs. So that makes it kind of ideal for experimenting with, uh, with coding agents and different, uh, development, uh, processes.

So I'm going to fire up Open Hands Cloud here and let's say, "Okay, I would like you to, let's build an implementation of this spec, make plan.md for, uh, with, let's see, Markdown task bullets. It's a pattern that I, I like to do a lot so it can take them off as it goes. And milestone sections." And I'm going to make a few choices here. But, you know, your preferences here, regardless, the motivation is, I'm going to try to make choices that set it up for success architecturally. So I'm going to, uh, do something kind of unorthodox. I could say, "Just, you know, use React on the front end, Python on the back end." I am, you know, these days very interested in, uh, languages that provide strong type soundness. So I'm going to say, "Let's do a full-stack Gleam and Luster app," which, you know, so this is a kind of obscure, uh, programming language and web framework. And let's see, "Presentation layer should be tested in with Playwright," and let's see, "Backend use hexagonal architecture," which we'll talk about in a minute. Um, "At every milestone, make an assessment task for, uh, for tests and modularity."

So, I don't mean to like, I would be copying and pasting if I thought I had like the perfect prompt, but these are the kinds of experiments I do is like, "Okay, what are all the ways that I could set this up for high confidence?" I even misspelled "hexagonal," but, you know, like that's the beautiful thing about coding agents is they, uh, they kind of understand what you, what you mean even when you make a typo. So, um, I'm picking a bunch of quality practices that I, I think on this kind of app might be helpful. And because it's so easy to generate these things, you, you know, we, we might try it out several times and, you know, one, one of them we like more. But, um, I'm not having it implement it right now. I'm having it plan the implementation. So, we'll come back to it once it's, uh, it's kind of researched and found that plan. And if you'd like to ask me about some of those particular choices that I made, we could, uh, we could talk about it in Q&A. So, let's, uh, we'll come back to that a little later.

Let's talk about shift left. So, hopefully, a lot of you are familiar with the term "shift left," right? It means bringing feedback earlier into your development cycle. So, as, as close as possible to where the decisions that are being made, such as code being written, are happening, is when you get feedback on, "Is that going well?" In the world of AI agents, sometimes you call this back pressure. I see that as kind of the same thing as the DevOps shift left principle.

Um, back pressure is, you know, when agents hear from their environment what, um, you know, about, "Did code fail static analysis? Did tests fail?" Something to indicate they might need to correct their path. And, you know, because notoriously agents, their flexibility comes from being driven by large language models, they can get off course. The better designed the, the back pressure, the feedback of their environment, the more we can keep them on track. The more we can, you know, have a good complement to the non-deterministic elements of the system. People talk a lot about build time, by which I mean not just like the act of building some software, but I mean literally the compiler is running, that kind of build time, the build process. We talk about build time, which is when compiler type errors or linters might kick in, and test time, when you're running the code in automated tests. These get a lot of discussion, but you can go further left than that. You can prevent bugs at design time. Indeed, I would argue most bugs are prevented at design time, just like most fires are prevented by building codes, by the way, construct buildings.

So, there's so much knowledge we could talk about here between design time and build time and test time. This is not so much a list of things I'm going to go through today. It's just to present, just to present like how vast this is, even if you have developed software for a long time, even if you're passionate about quality. You will probably find some of these things that you, uh, you haven't tried before, and some of them might be very relevant to what you're doing.

Something very underappreciated, for instance, is mutation testing. Now, it might be more relevant to you now that we have coding agents, even though it was a pre-existing practice. Here's why. Um, we are justifiably worried that when you ask for well-tested code, you will get code that doesn't cover everything. But if you ask for test coverage and look at your test coverage metric, it might be what's called "gaming the metric." It might not actually be testing all of the cases. It might just be running the code. So, given you might miss meaningful cases, there's something called mutation testing where you make random changes to the code and see, "Hey, if I change the logic here, a test should fail." If it's well-tested, I should be able to create a test failure by intentionally breaking the code. That's called a mutation. However, uh, it's tedious to do mutation testing. Honestly, you know, I, I have always felt that. I have very rarely used this technique, even though I was aware of it. But actually, there are very good tools that help, and coding agents can drive the process. So you can actually just say, "Okay, now that you've established that there's code coverage, do mutation testing. Use this library or technique and so on." You'll find prompts out there to help you do this, and they will find mutants. "Oh, I've missed a case where the logic can change, the test didn't fail." Uh, so I didn't have, before, honestly, the patience to do mutation testing, and now I do. Um, now it is very easy to incorporate a lot of these, uh, practices to help quality if we know to use them. These also interrelate. You know, these aren't fully independent.

A big example of this is that, um, feedback between test time and design time. If you know what you're doing, you will realize you find yourself writing a test where you're not actually very confident that this is going to prevent a bug. This may be just testing how it happens to be implemented. It's not testing that the functionality actually works. That is, uh, that is something you, you kind of learn with experience, which tests are, are really helping you and which aren't. But the bottom line is that this feeds back into your design. As, um, so it's not when we say design time and build time and test time, like these are discrete phases. No, these are, these are always occurring. Design and redesign is always occurring.

I've given a, a name to express this idea of, of what it means when your design fights your ability to test it. I call it the octopus problem. We'll see if that sticks. Now, octopuses are cute. At least graphical representations of them are. People have different views, but they, uh, they reach out and touch a lot of things. If you have a piece of code that reaches out, it does it, you know, it has side effects, it updates a database, it, um, it calls some service, um, one block of code is, is a lot of responsibilities, is interacting with a lot of different things outside of that code. This is hard to test. Even if you were to say, "Hey person, hey coding agent, make me rigorous tests for this octopus code here," you're going to get something that is, is ultimately very brittle. There'll be very many chances for the test to pass and not really to tell you what you needed to know.

What is easy to test? Not an octopus, a pipe. Pipeline. You'll see a lot of the more plausible examples of coding agents producing these huge code bases that that seem to work pretty well are compiler-like. Like, like why is that? A compiler, it thinks deep thoughts, but at the end of the day, it's just transforming input into output. It doesn't interact with anything else. And there are lots of other, you know, applications, like, let's say you're making a compression, you know, like you're implementing a zip command of some kind that has one input, one output, is very well-defined what the properties of that input and output are, and so it's very easy to test, and so you can have pretty high confidence that you've created a pipe of the kind that you meant. If it's an octopus, not so much.

So, what do we do about this? Because we also like about the octopus code that it can have actual results. Often times, we, that's why we're making a program, is because we need it to call an API service. We need it to make that database update. We, we can't just declare war on side effects, uh, completely, and never, and never do anything, right? Even Simon Peyton Jones of Haskell said, you know, like, even if the only side effect is that the computer heats up when you run it, there was some side effect.

There's a concept called hexagonal architecture that I mentioned earlier, and we're going to see, you know, how the, uh, agent fared in this instance with that hint. But this is another one of those patterns. I use it more now than I did when I was coding manually. Before, I would think all the time about, "How am I decoupling? How am I making a testable design?" But because I was so actively involved in the design of every little class, every little file, um, I was able to integrate that in my thought process. However, I need very high-leverage things I'm going to tell an agent. I need to be able to express unambiguously, like, "This is the design guideline," and hope that it's going to get followed. And that's where phrases like "hexagonal architecture" or "use the repository pattern," which is related, "use ports and adapters," these sorts of things. Okay, that's a key word. It's able to act on that, and that can be very high-leverage, and something without going into specifically like, "You should watch a talk on hexagonal architecture and ports and adapters," you know, for sure, but it is trying to ultimately accomplish this goal of making things more pipelike, making things more full of testable units that we can then compose. That, that is ultimately what we're trying to do when we architect for testability.

What else can happen at design time? Well, this is, uh, something a little less conventional. I want to give you a hint of how deep the rabbit hole goes. There's, uh, a family of techniques called formal specification. This part of formal methods, the, uh, pursuit of bringing mathematical rigor into programming. It's a subfield of computer science. The area of formal specification has to do with not proving the code itself, but proving things about your design. Now, it's kind of counterintuitive that this could be so valuable. We've kind of been conditioned as developers to distrust documentation. We think it's going to fall out of date. We think even if you design it right, you might not actually implement the thing that the design says. These are valid concerns. However, we have, we've seen that formal specifications, they give you really valuable feedback on your designs. And if you know certain things, if you know that a design itself is at least consistent, which you don't know if you've written prose, if you've written a markdown file, if you know you have a consistent design that solves certain problems, it does just reduce the, uh, the problems you're fighting at implementation time. So, you know, think of it this way: better to be spending time, you know, implementing a sound design and be dealing with implementation bugs than to be perfectly implementing a bad design.

This tool called Quint is one of the newer, uh, executable specification languages or, uh, or formal spec, if you like. It's based on an earlier one called TLA+, which has been heavily used by people like Amazon and, and, and others. I'll, um, I've, I've used it a little bit, even at Open Hands, and I found it, like, both whether you're designing something new or you've got an existing system, it's very helpful to, to, um, to even give Quint to an agent and say, "Here, model this and tell me what's true about it." You know, like, "Find some properties that are interesting." But there is a learning curve associated with understanding what properties make sense in a formal specification context. I would say if you're inspired, by all means, go, go and learn about this. One of my favorite introductions to this area is by my friend H.L. Wayne. He has a talk called "What Isn't Your System Supposed to Do?" And he talks from first principles about what formal specification is meant to achieve, but actually doesn't, you know, start by introducing you to the, uh, these actual specialized languages that you use to create these. But again, you know, this is something I've, um, I spent a great deal of 2024, actually, I call it the "Euro Formal Methods" year, learning about this area and sharing some of what I was learning. But we now, with agents, have the ability to get past a lot of the surface-level drag of the syntax of these formal specification languages. We still have to have some judgment and know, like, "What is a useful way to apply them?" But, uh, agents are at least able to, to help us with, with the part about using it.

Let's go back to the, uh, live example and then we'll do the final leg here of the journey. And, um, I hope you're preparing questions for the end here. So, let me get out of here and go back. So, we've got, hopefully, made our plan. So, we've got Gleam and Luster, which is the web framework, recall for Gleam implementation plan. I'm making kind of a conscious trade-off here, which I've, I've made before. I wasn't so sure this was going to, uh, work, but in choosing a, a, a more obscure, you know, it's, it's stable now, but it's only recently 1.0. A more obscure language like Gleam, you know, that a, um, the training set of LLMs will be a little less aware of it. I found this actually, it worked surprisingly well before. So, I'm, I'm, I'm still, I'm willing to take a risk on it. And what I'm getting out of it is, um, I think the, the things I think set, um, us up for making correctness very ergonomic. I, I like the, the choices made in the type system of Gleam. It was influenced by things like Haskell, if you're familiar with them. Um, but definitely, like, it is a, it is a big trade-off to think of using something more obscure versus, you know, using something that, uh, has the correctness guarantees you like. Rust is, you know, a much more popular thing that has some similar influences.

So, first, it's going to do a bunch of repo setup, right? It's going to do, okay, yeah, Docker Compose, so on and so forth. But what is its actual implementation plan? So, milestone one is "Domain Core and Hexagonal Skeleton." So it's going to make a bunch of classes that are for passing the, the data in the domain around. It's going to do adapters to talk to the database. That's, um, I think also would be called like the repository pattern. There tend to be a lot of names for architectural concepts like this. Um, so it's doing independent tests of, I guess, the, the data store and, and the token round trip of some kind. Then it is implementing the services. Notice it's like got the unit test last. Here's where, uh, if I was to critique this a little bit, you'll find agents will, will tend to want to do a bunch of a huge pass of untested code and they're like, "And then I'm going to test it." Like, I don't think that's that's great. I would be trying to get very continuous validation. That's not good use of back pressure. But regardless, um, so then it says, uh, it's going to do an assessment of the service layer quality. I like that it's doing that. "Confirm each service returns, you know, in this convention that it's decided to use." And "Check the services are stateless." And that's, uh, that's good. In this, for this kind of architecture, the service should be stateless. The repository should be handling the database adapter should be handling all the state. And we've got an HTTP API and, and then we've got the front end. So this is where I wonder, like, is it really going to be able to do all this? Um, and it also is treating the front end as its own just independent slice, which is kind of how the real world, uh, like that blog, you know, example spec, it is kind of treating the front end and back end as just completely separate and static. But again, in a real situation, I would be thinking, "Actually, should I be trying to do the simplest possible vertical slice and know that I've got a well-integrated front end and back end before adding another vertical slice of of functionality?" Just to show you my thought process a little bit. Um, okay. "Previews, Feed, and Navigation." So, this is pretty full-featured. Um, so I'm going to, uh, try something else out, actually, and say this is an, something I, uh, tried out earlier is I had it create a Quint specification for for this of the properties that ought to be true about this design, and it was able to run. So, just, this is like a, a, a design doc that lives and breathes, you know, it can, it can have conditions that pass or fail, but it's not actually an implementation you can use. It's a, it's a, it's its own kind of design doc. So I'm going to say, "Incorporate this Quint spec into your design and implement a usable full-stack slice of this blog to completion. Use your browser."

Everything okay, Jamie?

>> Yes, everything's fine. Um, we have about 15 minutes left.

>> Okay. And so we'll just check back at the end. Do as much as you can on your own. So we, we have, we have needs here for the, for the demo. So, um, let me just, let me just wrap up the concepts I want to present, and then I would just love to hear from you. You know, what's, since this is such a rich area, what are you most interested in in talking about? What are the difficulties you've had? You know, um, are there techniques you've tried that have worked?

So once again, we come back to the six-word takeaway: "Tell your agent how to test." Experiment with that. See what that gets you. Here's some more resources that I would recommend depending on your interest here. If you'd like to learn about just the core of what makes software engineering tick, right? These are kind of pre-coding agent books that I think are especially still relevant: "Working Effectively with Legacy Code" by Michael Feathers, "Kill It with Fire" by Marian Blotty, "Modern Software Engineering" by Dave Farley. So, in terms of just the basic physics of large code bases working on software products, I think this knowledge is still highly relevant.

If you want things that are specifically about overcoming the challenges, you know, with building with coding agents, I like a lot of Kessler's work. It's also in the, the Simon Society, Emily Bash's organization. Um, so she has this, this repo called "Augmented Coding Patterns," where she's mapped out a lot of the trouble people are having about, you know, agent behavior and, you know, for this problem, this is the solution we're finding works. It's kind of this, this living community of resources, very helpful. And Open Hands' agent team has been very actively involved in just the science of how we understand and improve agent capabilities. A couple of things I'd like to mention there are the Open Hands Index, which is, you know, aggregates a bunch of the different benchmarks for different kinds of, of agent capabilities. And a brand new benchmark, we'll see if it becomes part of the index, but I'm very excited about this one, um, which has a guest blog on our, on our website called "EvoClaw: Evaluating AI Agents on Continuous Software Evolution."

So, the, um, first really large breakthrough, as I see it, in coding agent benchmarks was when it went from generating, uh, single files in benchmarks called HumanEval, if you remember that from 2023-ish, to this benchmark called SuBench, where we're evaluating tasks at the scope of an entire repository. Once we did that, we knew we had things that could be practically applied, you know, to real-world code bases. However, it was still just a snapshot in time. And code bases change over time. When we talk about maintaining quality, it's about making sure things stay good. They do not get worse with time. And that wasn't being captured. But we're starting to have approaches like EvoClaw to measure what is the impact of this agent over time. Is it actually making the codebase better or worse? Is it, you know, are bugs going to become more likely over time, things like that? And I think that is just such a promising area of research. So, I'm excited to, uh, to have that happening now and for us to be involved in it.

Um, so with that, I, uh, just want to, before we go to Q&A, I want to kick it off to Clark, who's going to talk about, you know, if you want to get hands-on with us, uh, you know, helping you implement Open Hands in your own situation, what that might look like. Clark?

>> Great, great. Thanks, Ray. I appreciate the, uh, the last hour and wanted to do a quick voice-over on the Open Hands pilot program. Um, I'm Clark. I'm on the, uh, GTM team, GTM team here at Open Hands. Been with the org about 6 months, um, and working with customers of all different shapes and sizes, um, around Open Hands Enterprise. And so if you want the ability to, uh, deploy the Open Hands solution in a secure way on your infrastructure, um, we'd be happy to walk through that with you. And so, um, you can see who the fit is, what you get as part of the package, and then what it typically takes, uh, to run through the 30-day evaluation.

Um, the next step for this, if you want to get in touch, would be following up at the link in the bottom there, reaching out to me at my email or my, my LinkedIn profile, or on the Open Hands website, uh, when we jump on with a forward-deployed engineer, uh, to go through the scale and scope of what you're trying to accomplish with Open Hands and back into some of the broader themes that Ray has covered today and in previous webinars. Um, so, yeah, please reach out if you're interested. Happy to field questions on that. Um, you know, as part of the question Q&A dialogue session that Jamie's going to administer in a second here.

>> Cool. Great. Excellent.

>> Uh, and with that, we'll go ahead and kick off some questions. Um, we can start. I'll do a little selection. So, uh, Muhammad had a question in regards to, "Build on Ray's point. Not every part of a software system should change at the same pace. Good architecture protects a more stable, deliberate core while letting faster changes happen at the edges through plugins, adapters, extensions, points. In that sense, 'move fast, break things' can still be valid, but only where experimentation is expected. The blast radius is limited. Failure is affordable. Would you agree?"

>> Man, I mean, what a well-made point. You know, what could we add to that? And that's why, like, "move fast and break things" is still, like, this is a bad slogan because "things" is pretty vague. Huh? But, uh, not, yeah, not all changes are equal in risk, and, uh, you know, architecture involves tradeoffs of what changes are likely to be needed in the future. Is this going to meet today's needs and tomorrow's needs? And this is why, you know, coding agents, some would argue, I, I think there's something to this, that they're just sort of inherently bad at these kinds of judgment calls, but also they're not even equipped. They, they don't know what your needs are likely to be tomorrow. This is, um, in terms of your needs of change over time. This is what you need to be injecting in the process. You know, that, that is what you're here for. So, I just, I love what, I love everything you just said.

>> Excellent, Muhammad. Thank you so much for sharing that. Uh, I'm going to go to, uh, Ragunath, who asked, "When you ask the agent to implement the solution using hexagonal architecture, how do you verify if the AI has actually implemented it properly?"

>> Yeah. So, um, there's a, a couple things here. You know, I know I'm, I'm, uh, uh, in order to get to the, the content, I'm kind of, uh, I'm not spending a ton of time on the ins and outs of, of this part. But, you know, basically, agentic review, manual review, algorithmic review are the three tools at your disposal. Uh, there, there have been, um, cases if you study something called evolutionary architecture, people have sometimes made even like automated tests that certain architectural constraints are being followed. You know, there are times when that's helpful. Uh, but I mean, I, uh, I, I said the things I said about like the, the Ralph Lisp example, like, you know, that it had made clean separation in certain places because I looked at it, you know, I know it because I checked. And there are things that, um, are pretty high leverage where it doesn't take you that long to verify, where there's not even a tradeoff to consider, like human review is not holding us back at all here. But there's also a lot of progress that we've made. Like, I was skeptical of this at first. I have become a lot more open-minded to, um, agentic review as a piece of the puzzle. And, uh, as has Paul Hammond, who's a repo I mentioned earlier. You know, he's got this area on test design reviewer, which is based on another person I mentioned, uh, Dave Farley's, uh, criteria of what makes a good test. Emily Bash has also written about this. So we thought, like, "Okay, he's got this evaluation framework of like, are the tests the way I want them?" Similarly, um, an architecture reviewer, sort of as a separate agent session coming in to say, like, "Was this really followed?" That has started to, to get pretty, pretty usable. So, yeah, um, those are, those are the three tools: algorithms, people, and, and agents. And, uh, and we can get, uh, pretty good results with a, with a fair amount of setup that's, that's being sustained over time once you've made those decisions. Great question.

>> And then I think we have time for one more from Nelson. And their question is, "In this in the age of LLM coding agents, out-of-date docs represent a much larger harm versus human engineers because they consume, they LLMs consume it naturally during development. Doc drift directly impacts code quality. From my experience, do you have any suggestions on managing this with CI/CD?"

>> Yeah. Um, so I mentioned a little bit formal specification. You've seen with behavior-driven development similar things of just executable where the, the docs interleave with the, uh, automated tests. These are areas, but I'll introduce a new, uh, word that I'm not hearing enough, which is traceability. Requirements. Traceability is a, a long-existing concept that I think is, is again high leverage right now where, for example, there are particular keywords of, "This is where these sources of truth tie together." Some, you know, if you're able to just keep the code coexisting with the docs in a way, like with Swagger docs, you know, "Okay, the API code literally compiles into the Swagger doc," you know, that they're in sync. But for other things where you can't, having these keywords, these IDs that link them together so you can easily navigate between the, you know, the real source of truth, which is probably the code at trunk branch, and all the things that's supposed to be describing that code, that gives you the, um, the kind of, you know, connections you need to easily keep them in sync. But, um, you know, like I said, the, uh, with the agentic review, an agent review of doc consistency is something that people have also been, uh, having some good luck with. Um, so those are some places to, to start, you know, maybe, uh, too many options even.

>> Thank you, Ray. Uh, and with that, we did have a couple more questions. One was about test-driven development and how to get started with that, and another one was about the SDK. I'll let Ray follow up on the first one, but before then, for those that have to leave at the top of the hour, we just want to thank you so much for attending the webinar. And like we

shared before, uh, a recording will be shared as well as the presentation to all those who RSVPd. So, please keep your eye out for that. And please make sure to join our Slack. The link is in the chat. That's the best way to continue the conversation, especially if you have questions related to Open Hands or using Agentic AI that is beyond the scope of this presentation. And, um, we also have a contact form, if you, which the link is there as well, if you want to follow up with Clark. And with that, Ray, I'll let you get started.

How would someone learn more about test-driven development?

Yeah, and what was the other, was the other one about the SDK?

The other one was about the SDK.

Yeah, I, I would need something more, more specific, but, um, you know, software agent SDK is, uh, is.

I can pretty, I can give you a little bit more background. They were saying that they were using the Minimax, uh, 2.5, and the agent was pretty conservative and wouldn't necessarily take the risk or go as far as you would need it to go.

Yeah. Oh, okay. I'll comment on, on that, uh, briefly afterward. So, I've got an interview with Emily Bash, who I've mentioned a few times, uh, in this, in this talk about the relationship between AI and trust-driven development. And there's also a section in there, I think in the, the second half, where she talks about how to get started learning TDD. Um, things like cottis, which she's got a great deal of, um, on our GitHub that are linked from that interview description, are, uh, helpful. To work with a coach, to pair program with people are some things that came up there. But I would say, yeah, uh, watching that interview is, is great. And I, I'm, I'm very happy to see the resurgence of interest in test-driven development, even though it's historically like a manual process. We're kind of figuring out how to reimagine how that works with our human judgment going into, with the agent loop. But, um, I think it is a very rich ground still for us to, to learn about.

Um, so the point that was made about Minimax and eagerness, that's, um, that's pretty interesting. Eagerness is a word that comes up. People have said that, for example, like, oh, um, this Opus model versus this GPT model, this one's more eager. What does that mean? Well, it's, um, in terms of how far it will go on its own, what liberties it will take, uh, these different models based on just quirks of their training process have these different, you know, kinds of things. So some have even said by standardizing on agents MD instead of Claude MD and, you know, um, GPTM MD, for instance, and Minimax MD, we're actually trying to use a one-prompt approach, which isn't always the case. So, I think it's, it's a, a good observation that all I can say is like, watch what it's doing and adjust your prompting based on what you're seeing. But it is very much the case that some models you will need to go out of your way and say, "Now, don't improvise at all. Don't do anything I didn't tell you to do." Whereas someone, some of them you want to actually encourage to, uh, to take a little bit more on because they have, yeah, these different, uh, predispositions based on, on training sets. And Open Hands being a, you know, a model agnostic agent, uh, yeah, we, we are in the position of, of trying to make a, a, you know, system prompts that work pretty well regardless. So you, you will have to adapt a little bit if you're changing the model.

And what would you say are best practices when using the SDK in terms of balancing memory management?

Well, I think the, uh, the Ralph Wigum webinar had a lot of good, uh, uh, you know, discussion about context management. So memory management, context management in this context mean the the same thing, I think. But, um, what ultimately what makes Ralph kind of unique is that idea of for every leg of the task, restarting with a fresh context window. And I think that's an important thing to try. Some people go back to just doing super long conversations. But a good place to start, I would say, is, uh, try doing kind of the opposite of what I did, where I'm just coming back to this one mega conversation where I'm implementing the whole app. Like, I'm doing that because it's like, it's easy to do with low attention right now. But as you're, you know, try getting more scientific with it to where you can, you can start each and do it headless. Say, I'm just going to prompt it once at the beginning and let the whole conversation play out. That will force you to get very deliberate about what context, what memory you're putting into it. So, you know, I hope that's helpful.