📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

와튼 1,372명 실험 — 구글 출신 엔지니어가 말한 "AI 시대에 사람이 남는 자리"

AgentOS11:55

Transcription

AI가 틀린 답을 줬을 때 사람들이 어떻게 했는지 조사한 실험이 있어요. 와튼 스쿨에서 1372명을 대상으로 했는데요. AI가 틀렸을 때도 73%가 그 답을 그대로 받아들였습니다. 그리고 AI 없이 혼자 판단했을 때보다 오히려 더 확신했어요.

이 숫자를 꺼낸 사람은 에디 오스만이에요. 구글에서 14년 있으면서 크롬 개발자 도구를 만드는 팀을 이끌었고 코어 웹 바이타를 공동으로 주도했던 사람입니다. 마지막엔 구글 클라우드 AI 디렉터로 재미나이 개발자 경험을 총괄했고요.

AI 엔지니어 월즈 페어 2026 클로징 키노트에서 발표되어요. 제목이 오D뒤 아우터로프 바깥꼬리를 소유하라는 뜻이에요. 에이전트가 코드를 짜고 테스트하고 보고까지 하는 시대에 사람은 정확히 어느 자리에 남아야 하는가? 이 발표는 그 경계선을 꽤 구체적으로 그립니다. 영상 중간중간 발표자 클립이 들어갑니다. 한국어 자막 같이 띄울게요. 먼저 지금 어디까지 왔는지부터요.

소나가 2026년에 낸 스테이트 오브 코드 리포트를 보면 커밋된 코드의 42%가 AI가 만들었거나 AI가 상당 부분 좋은 코드였어요. 2023년에 6%였습니다. 3년 만에 일곱배가 됐고 2027년에 65%까지 갈 걸로 봤어요. 코드를 만드는 비용은 이렇게 계속 내려가고 있습니다. 그런데 만드는게 싸졌다고 해서 확인하는 것까지 싸지지는 않았어요.

I love working in my factory. I love building my engineering loops. But the problem is still capacity. If 96% of people don't trust that code, but only about half always verify before committing, we have this danger that we've got distrust without bandwid. And so safety comes from making verification cheaper, clearer, and harder for people to skip. And if you zoom out from the individual reviewer to the organization, review and validation start becoming a bottleneck when governance isn't able to catch up and adoption is already moving way faster than any company can go and set their policies. And this means that we have some hard questions we have to deal with like did a model actually touch this file and the hard questions are also like what constraints guided that work, what evidence was produced, what risk was accepted and who owned the result. Now the agent can ship more than any of us can review, right? So what are we still good for? It's a question that's on a lot of our minds, right? And you know, if Homer Simpsons experience automating computers can teach us anything, maybe this is our future. I don't think it is, but it's one direction things could take. Now, let's try that again. If change is where humans enter the loop, if generation scales faster than comprehension, the scaress resource becomes judgment that's backed by evidence.

에이전트가 확인할 수 있는 양보다 많이 만들어내는 상황. 그럼 사람은 뭘 위해 남아 있느냐? 발표는이 질문에 답하려고 용어 두 개를 꺼냅니다. 알파와 DK. 알파는 지금 모델이 할 수 있는 것보다 내가 의미 있게 더 나은 지점이에요. 쉽게 말하면 내 강점이죠. DK는 그 격차를 모델이 따라잡는 속도고요. 내 강점에 걸린 시계라고 보시면 됩니다. 그래서 요즘 다들 취향 테이스트를 이야기하잖아요. 아무나 열 개를 만들어 낼 수 있으면 그중 뭐가 존재할 자격이 있는지 고르는게 희소해지니까요. 근데 그 취향도 결국 알파입니다. 알파이는 시계가 걸려 있고요.

So yes, taste matters when production gets cheaper. And if anyone can generate 10 options, the scar skill is really knowing which option deserves to exist. The taste is not some eternal mo it's alpha as well. Now the people with taste are still going to matter. I personally think they're still going to matter for a long time. But the best version of that skill is not my better calls and leaving behind examples that your team in the system can learn from. Now let's applied the decay test. Well, we used to have speed that decade. We used to have recall. You know, harnesses have memory. Verification is moving into harnesses, evalows, static checks and model critique. Taste, I continue to think this is going to decay much more slowly, but it still resets as models learn from examples and preferences. Even judgment in some ways is a slope rather than a wall. So, the strategy is not to cling to any one capability. It's for us to keep moving our edges up a level. So this is one of the reasons why what can the agent do is not the best strategic question anymore. The list of things that agents can't do just keeps shrinking. The better question for us is really what can only a human be answerable for? Not because you know any of us are magical any because some decisions actually require ownership require.

>> 오직 사람만 답할 수 있는 거. 여기가 남는 자리라는 얘긴데요. 그래서 발표는 엔지니어라는 단어를 조금 더 좁게 잡습니다. 지금은 누구나 컴퓨터로 뭔가를 만들 수 있는 시대죠. 만드는 사람의 수는 역사상 가장 많아졌고요. 근데 코드를 짜서 결과물을 존재하게 만드는 것과 시스템을 놓고 판단하는 건 다르다는 거예요. 제약을 따지고 트레이드 오프를 방어하고 리스크를 관리하고 무너졌을 때 연락이 가는 사람. 그 자리를 지키려면 잘하는 걸 늘리기 전에 피해할게 먼저 있습니다.

AI to solve our problems. For code, it's the gap between how much code exists in your repo and how much any human on your team genuinely understands. And this is why things like delegation depth end up mattering. You can have a build that passes, you know, your tests, a PR that you can merge, but your team can still end up losing its ability to actually explain the system that they are shipping to production. Now a very real pressure is much is also how much we delegate. So agents can now stay inside the system long enough for the human to lose the thread. So a 30 second run, right, can feel like an interaction but an hour or a day scale task, so something long horizon, that's a workstream. And when tasks can end up, you know, lasting that long, especially when you begin running many of them in parallel, review can't just be a glance at the end. It has to become a whole control system. The second thing to avoid is cognitive surrender. Now this is when you blindly accept AI um responses like delegation is important. Delegation says do the work then show me enough evidence that I can judge it. I still make a judgment in that situation. Surrender is really saying hey your answer is now my answer before I have formed any opinions myself. Now, uh, did a study that kind of offers us a warning light here. When AI was wrong, 73% of people still thought that, you know, they picked the wrong answer and they felt more sure. So, the failure mode is not using AI, but it's borrowed confidence. The third thing to avoid is orchestration tax. Now, if you've been in the Bay Area, you will see people who for better or worse are still walking around with their laptops open or are talking to you about cloud agents. And we're increasingly trying to run more and more and more in parallel or telling each other that we're shipping with hundreds of agents or thousands of agents. More AI agents running does not mean that there is more of you available. Your cognitive bandwidth does not parallelize. So every loop that you create ends up causing more decisions to route, merge, verify, and integrate. And the fix is not necessarily few agents but it's about your attention like a system like where you enter what you require what you reuse in it.

인지 >> 부채, 인지적 항복, 오케스트레이션 세금 세 가지 다 공통점이 있어요. 결과물은 남는데 내 판단력이 깎인다는 거죠. 그러면 이렇게 되죠. 책임이라는 단어가 무거워집니다. 에이전트한테 떠넘기고 숨고 싶어지고요. 근데 발표는 여기서 방향을 뒤집습니다. 책임은 에이전트가 충분히 좋아진 다음에 남는 찌꺼기가 아니라 나머지 전부를 확장시키는 조건이라는 거예요. 여기서 커리어 계산이 하나 나와요. 내가 가진 엣지의 반감기는 모델 릴리스 한 번이래요. 속도든 기억력이든 검증이든 프론티어가 움직이면 같이 움직이니까요. 근데 서명의 반감기는 커리어 전체입니다. 서명은 결과물에 붙는 이름이에요. 내가 내보낸 것 뒤에서 있는 사람 팀 조직. 그러니까 스킬은 레버리지를 벌고 책임은 그 레버리지를 신뢰로 바꿉니다.

And this is one of the lines that I want to draw pretty clearly. Ag, they can route, they can merge, they can escalate, they can operate inside policy. And in many systems, you know, they can, they should. But execution and responsibility are very different things. The agent can follow your runbook, but it can't inherit the consequences. When something fails, the question is who understood the policy? Who accepted the risk and owns the blast radius? High agency is something that a lot of us talk about these days as being like this thing that we're looking for when we're hiring. High agency is actively taking ownership of your outcomes. So knowing when to delegate, when to inspect, when to stop, and when to put your name on the result. High agency in this world is not I personally do everything. You know that version doesn't really scale. It's not just hustle theater but it's ownership with judgment attached. This agency lad tries to make that a little bit more concrete. At the bottom you've got someone that flags a problem and leaves it for the system. High up. They execute, diagnose, propose, recommend and resolve. And the rare top movement is discernment. You know, maybe you find a problem and you decide whether or not it's worth investing in. Maybe it's not and maybe you move on. But when agents make more paths possible, agency is not chasing every single path. It's really just deciding which paths deserve your ownership and attention. So translate that into an operating model. Agents can run much more of the inner execution loop. They can investigate, implement, test, and report. I think that there's leverage in that but that outer loop is still engineering. So deciding, verifying, approving, owning. That inner loop is capability. The outer loop is agency. And this is a boundary that I really care about. Your agent returns evidence. It returns diffs, tests, logs, ration, traces, trajectories, screenshots, whatever the work itself requires. But then the engineering really begins. We decide whether the work was worth doing. We verify whether the evidence is enough and we approve or redirect or own what reaches production. It doesn't matter if you're someone that's just working with a small number of agents, whether you're working with thousands of agents. I still very much think that these ideas apply. The boundary is not human looks at AI output the boundary is evidence and responsibility.

정리하면 이렇습니다. 하나, 만드는 비용은 계속 내려가는데 확인하는 비용을 그대로 해요. 둘, 영향이면 사라집니다. 속도도 기억력도 검증도 이미 넘어갔고 취향도 예외가 아니에요. 셋, 에이전트를 쓸수록 세 가지가 조용히 셉니다. 인지 부채, 인지적 항복, 오케스트레이션 세금. 넷, 안쪽 꼬리는 영향이고 바깥 꼬리는 판단과 책임이에요. 그 경계는 사람이 화면을 들여다 보는게 아니라 증거와 책임입니다. 다섯, 규칙 한 줄 설명할 수 없으면 내보내지 않는다.

마지막으로이 발표에서 제일 좋았던 부분을 하나 더 말씀드릴게요. 소프트웨어를 만들기 쉬워질 때마다 사람들은 세상이 소프트웨어를 덜 필요로 할 거라고 예측했어요. 고급 언어가 나왔을 때, 프레임워크가 나왔을 때, 클라우드가 로우코드가 나왔을 때. 근데 매번 반대로 갔습니다. 비용이 내려가니까 그동안 만들 엄두를 못 냈던 것들이 쏟아져 나왔어요. 에이전트도 같은 일을 할 거라는 거예요. 엔지니어링을 없애는게 아니라 병목을 옮기는 거죠. 이걸 만들 수 있나 해서 이게 존재해야 하나 그리고 우리가 그걸 책임질 수 있나로 구독과 좋아요는 영상을 만드는데 큰 힘이