📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Salesforce Killed The Browser. Every Agent Runs Your CRM Now.

AI News & Strategy Daily | Nate B Jones23:08

Transcription

Every week, another AI agent launches. Just in the last couple of weeks, OpenAI shipped workspace agents. Anthropic put clawed managed agents into beta. Salesforce turned the entire platform into headless 360 for agents. Perplexity put personal computer on the Mac. Moonshot dropped Kimmy K 2.6 with a 300 agent swarm.

Another company says, "This is the one that changes everything. You've heard it before." Another benchmark chart goes around. You've seen it before. Another founder posted demo where the agent does the entire job while everyone in the replies argues about whether it was real. And if you're leading a team right now, the reaction I keep hearing from folks I know is it's not excitement, it's exhaustion.

The question is not what launched this week. The question is now which of these millions of things actually deserves an afternoon of my team's attention? My answer is that you need a filter because the agent conversation has quietly moved from model quality to infrastructure. The launches that matter are not always the ones with the best benchmark score or the loudest demo. The launches that matter are the ones that change what your existing tools can reach, what your agents can do, and how easily your team can stack those systems all together.

So, I'm going to give you the five question filter I use for every agent launch. And then I'm going to run five representative releases through it. Chad GPT Workspace agents, Salesforce Headless 360, Microsoft Copilot Wave 3, Kimmy K 2.6, and Perplexity Personal Computer. The Claude managed agents piece is going to come back near the end because it helps explain the bigger shift underneath all of this. Claude is no longer just a product you switch to or away from. It's becoming a direct product, an embedded engine inside other people's products and now a managed infrastructure layer for teams building their own agents. And at the end, I want to answer the question underneath all of this, which is usually phrased as when should I switch from Claude or when should I switch from Chad GPT or do I switch from co-pilot? I think that question is framed wrong. This is not a switching question anymore. It's a layering question.

Okay, let's get to that filter. The filter I use has five questions.

First, does this plug into the tools my team already uses or does it expect me to move my work to a new environment? That sounds simple, but it eliminates a huge number of launches right away. The best agent news is infrastructure news. It gives the agents you already use a better way to reach the places where your work lives. The worst agent news is a new destination your team is supposed to migrate to. We've already lived through a decade of SAS proving that migration is the most expensive thing you can ask a team to do. People do not want another place to move their work. They want the work to become easier where it happens.

Second, does this let other agents build on top of it or is it a closed product? If I can point Claude code at it or codeex or cursor or a custom internal agent, that's infrastructure. If it only works as its own standalone experience, it's a feature. Features commoditize, infrastructure compounds.

Third, does it own or access data I care about? Agent quality is downstream of data access. A mediocre agent with your full customer history is often more useful than a brilliant agent staring at an empty context window, which is why C-pilot matters inside Microsoft 365 companies. It's also why Salesforce matters inside revenue organizations. It's why agents that look boring from the outside can be extremely valuable inside the systems where the work actually happens.

Fourth, is there an ecosystem forming around this? One-off launches will fade and ecosystems, they compound, right? Marketplaces, SDKs, partner programs, developer tools, consistent shipping cadences, those are all things to watch for. Those are things that tell you whether a release is going to stick around 6 months from now. So, a product with a growing marketplace is very different from a product with a press release.

Fifth, can I stack my agents on the top? This is one that people often forget because a release that lets me compose with other agents is much more valuable than a release that simply adds an additional agent I need to evaluate against all the others. The first one, it multiplies. The second one just adds one more agent to check.

Run any launch through those five questions and the noise gets quiet because most launches fail those tests. The ones that pass, they're worth an afternoon of looking into. And anything else, it can just go into another pile to look at when you're bored on a Friday afternoon.

Now, let's run the current agent news through exactly that filter and see what passes the test.

Start with chat GPT workspace agents because this is the one that put the question in front of a lot of teams this past couple weeks. OpenAI's move here is very clear. Workspace agents are not just another version of custom GPTs. They are shared codeexpowered agents for teams. They run in the cloud. They can work across connected tools. They can be used in chat GPT or Slack. They can be scheduled and they're built for repeatable business workflows rather than one-off sessions. That matters. The interesting shift is we're moving from an I have a helpful personal assistant model of agents to our team has a reusable work unit model of agents, a product feedback routing agent, a weekly metrics reporting agent, a risk screening agent, a software request triage agent. Those aren't magic examples. Those are exactly the kinds of repetitive workflows companies keep trying to automate and then failing to maintain because the process lives across Slack and email and documents and spreadsheets and ticketing systems and somebody's memory.

Workspace agents pass the filter for a specific category of work. Recurring team workflows where the conversational builder, the Slack surface, the cloud execution, the permissions, the approval, and the shared agent directory matter more than having native control of a system of record. That's a real category, but it's not every category. If your workflow, for example, is deeply native to Salesforce, then Salesforce has that data advantage. If your workflow is deeply native to Microsoft 365, Copilot might have the data advantage. If your workflow is Frontier Coding, you probably still care more about the coding agent than the workspace wrapper. So, I think the right way to think about workspace agents is not this replaces every other agent. The right way to think about it is this is OpenAI's strongest answer for shared repeatable cross tool work that teams want to run from chat, GPT, or Slack. And that's just one of the five cases.

The second one is less flashy, but it may matter more. Salesforce Headless 360 is the launch most people are probably going to forget about to be honest. The name makes it sound like platform plumbing. And it it is platform plumbing. That's the point. Salesforce announced Headless 360 at Trailblazer DX. And the important part is this. Every major capability across the Salesforce platform is being exposed as an API, an MCP tool, or a CLI command. That means the browser user interface is no longer the only way to use Salesforce. Agents can reach into Salesforce directly. Coding agents can work with live org context. External tools can call Salesforce workflows without a human clicking through the interface. Parker Harris, Salesforce's co-founder, apparently asked the question this way. Why should you ever log into Salesforce again? Headless 360 is the answer.

The numbers matter here because they show the shape of the strategy Salesforce is going after. They're building more than 60 new MCP tools. They're building more than 30 preconfigured coding skills. Support for tools like Claude Code, Cursor, Codeex, and Windsurf. An experience layer that separates what an agent does from where its output appears. So the same underlying agent can render across Slack and mobile and teams and chat GPT and Claude and Gemini and any other MCP compatible client. And then you have agent exchange which pulls together the Salesforce app ecosystem, right? Slack apps, agent force agents, tools, and MCP servers in one marketplace. You also have the builder fund behind the ecosystem. And that's why this matters more than the headlines might suggest. Salesforce is not launching an agent. Salesforce is trying to become infrastructure under the agent economy. So every company runs on Salesforce. If your company runs RevOps on Salesforce, the question is no longer should we use agent force or workspace agents for CRM work. The question is which of the agents we already use do CRM work because Salesforce finally expose their data properly? And the answer is maybe all of them. A coding agent can build Salesforce apps with live data. Now, a workspace agent can update opportunities through the right permissions. A custom internal workflow can trigger Salesforce flows. A Slack native agent can act on CRM data without asking a human to copy and paste. So, if you run headless 360 through my five question filter, it scores extremely well. It plugs into an existing system where enterprises already live. It's explicitly open to other agent frameworks. It owns the data revenue teams care about. It has a real ecosystem and it is designed for other agents to stack on top of it. This is what infrastructure looks like in the Asian era. And there is one sleeper detail inside the Salesforce announcement that connects to a much larger trend. So As agent force fives, the Salesforce development environment uses clawed sonnet 4.5 as its default coding model with GPT5 available as an option through multimodel support. That's part of a much larger pattern. Anthropic's enterprise strategy increasingly looks less like build a standalone product that replaces everything and more like be the agent layer inside other people's products. You see it in Salesforce. You see it in Microsoft. You see it in Perplexity.

And that brings us to the third case we'll look at today. Microsoft Copilot Wave 3. I know a bunch of you think I don't cover Copilot, but I cover Copilot. That's still the story a lot of people are missing. It predates the last week of announcements, but it belongs in the conversation because it shows the other version of the same infrastructure shift. The two pieces that matter are copilot co-work and work IQ. C-Pilot co-work brings longunning multi-step agent execution into 365. Microsoft built it by working closely with guess who? Enthropic bringing claude style agent tech into the copilot surface. So work IQ is the data layer. It gives co-pilot access to the full context of work inside Microsoft 365. So it has email, meetings, chats, files, shareepoint pages, identity, permissions, and organizational context. The data access is the point and it's also the moat. A chat GPT workspace agent reaching into SharePoint through a connector absolutely can do useful work. But co-pilot sitting natively inside SharePoint and inside Outlook and Teams and Excel and PowerPoint and the Microsoft identity system is different for Microsoft native enterprises. This is not a small difference. The agent is not just connecting to a file. it is operating inside the organizational graph.

This is where the filter helps because copilot is not equally strong on every axis. It is very strong on data access for Microsoft 365 shops. It is very strong on native permissions and enterprise governance. It is strong when the work lives in Excel, Outlook, SharePoint, Teams, PowerPoint, and the surrounding Microsoft stack. But it is weaker on openness to external agents. It is harder to point outside agent frameworks at Copilot than it is to point them at something like Salesforce's MCP layer. The ecosystem energy is very different. It's closed. And for coding heavy workflows, most serious engineering teams are not going to touch co-pilot. They're going to go for codeex or cloud code. So copilot wave 3 passes the filter for a specific audience and fails it for others. If your team's work mostly lives in Microsoft 365 and mostly is not production engineering, then co-work can be a significant agentic product you should evaluate. But if your team's work crosses ecosystems or depends heavily on coding or needs external Asian composability, co-pilot's native data advantage is just not worth it. So the question is not is co-pilot good. The question is really for which shape of work is copilot the right layer.

Now the fourth case is almost the opposite. It gets a lot of press attention because the model's very impressive but for most enterprise teams it matters less than the infrastructure launches. Kim K 2.6 which just launched is very technically impressive. Moonshot released 2.6 as an open weights model under a modified MIT license. The model card frames it as a native multimodal agentic model built for long horizon coding design autonomous execution and swarm-based orchestration. The headline capability is the swarm architecture. 300 sub aents coordinating across up to 4,000 steps. The published evaluations show very strong coding and agentic benchmark results. And because the weights are available, serious teams can self-host it instead of sending work to a closed provider. And that's a big deal.

But this is where the filter keeps you from being hypnotized by a bunch of new benchmarks. For enterprise buyers, Kimmy does not pass the filter the same way Salesforce, Microsoft, OpenAI, or Perplexity pass. It does not own your workcraft. It does not sit natively inside 365 or Salesforce. It does not have the same Western Enterprise connector story. It is not primarily solving the problem of how does my existing team route recurring work through tools we use. It's solving a different problem for dev teams. Building their own agent infrastructure. 2.6 is a real option. Open weights under a modified MIT license means you can run it on your own hardware, keep data inside your environment, fine-tune or inspect, and avoid being locked into a closed lab for every longrunning agent workflow. And for that person, for that team, Kimmy matters a lot. But for the western proumer or business team asking, should I use a hosted Kimmy product? The answer is definitely no. Not because the model's weak, the model is strong. The issue is the product and infrastructure context around it. If you're typing sensitive company work into a hosted product where you are not comfortable with the data path, the model's benchmark score is not the point. The deciding variable is trust and governance and data boundaries and connectors and whether the product fits the environment your team operates in. Those are different cases. The self-hosted developer team is just evaluating Kimmyy's infrastructure. The casual hosted user is just evaluating Kimmy as a workplace product. Totally different worlds. And that's why Kimmy matters strategically, but is often not the default answer for most teams that are watching this video. It tells you that openweight agent models are moving quickly. It tells you that long horizon agent architecture is advancing outside the Western Labs. It tells you that teams with a capacity to run their own stack have more credible options than they did a few months ago. But if your question is, "What should my sales ops finance or product team use next week?" Kimmy is almost certainly not the product to evaluate.

And that brings us to the fifth case, which is almost the mirror image here. It's less model ccentric. It's more workflowcentric. Perplexity personal computer on the Mac quietly closed one of the biggest gaps in the Perplexity Computer ecosystem. They've been launching really smartly recently. Perplexity Computer already existed as a digital worker in the in this vein of OpenClaw, right? But but it had a problem. It could feel like a cloud assistant that knew a lot and searched well and produced work but did not fully live on your machine and did not have local file access. The Mac rollout changes that personal computer adds local file editing, local computer use, local browsing through comet, voice orchestration and deeper control over work happening in background. Perplexity also made Claude Opus 4.7 the default orchestrator model for computer with other model options available. That matters because the product category becomes easier to understand. Perplexity is not just giving you a chatbot that searches. It's trying to give you a digital worker that researches and reasons and browses and edits files and creates artifacts and moves across connected apps.

Run it through our five question filter and the score is mixed but useful. It's strong on connectivity. Perplexity has a broad connector surface and computer is built around chaining capabilities together like research and analysis and docs and slides and emails and code and schedules and follow-ups. It's stronger now in local access because the Mac app can touch files and apps directly. It's moderate on ecosystem and team level workflow structure. It's not the same thing as a shared workspace agent in Slack. It's not the same thing as C-pilot sitting natively in Office 365. It's not the same as Salesforce exposing its trust layer and business logic through an MCP. So, perplexity passes for a very specific category, research heavy work that produces a deliverable like competitive intelligence or market research or sales prospecting or financial analysis or document review or weekly ops reports. anything where the work starts with gathering context from many places and then ends with a polished artifact. Now, it fails when the job is a team shared recurring process that needs governance and ownership and repeatability across an org. It fails when the work is deeply native to Microsoft 365 or Salesforce and the native graph matters more than the research layer.

This is not a good or bad question, right? It's it's a routing question. And that's the pattern across all five of these examples. Workspace agents are for shared recurring workflows in chat GPT and Slack. Salesforce headless 360 is for agent access to CRM data and business logic and revops infrastructure. Copilot co-work is for Microsoft native work where work IQ is a data advantage. Kimmy K 2.6 is for teams that can use openweight agent models as infrastructure. Perplexity personal computer is for researchheavy work that needs to become an artifact. Don't force one product to do every job. Assign the work to the tool. This is the part most teams skip and it's where a lot of wasted license spend happens. Someone buys the license and then the company tries to make that one product cover every job class because adopting a second tool is expensive. But the expensive thing is not having multiple tools. The expensive thing is routing the work incorrectly.

If you have a researchheavy deliverable, perplexity is a great candidate. Ask it to compare competitors and synthesize news and build a market map and draft a report and turn it into something a person can review. If your team lives in Microsoft 365, co-pilot can be a good candidate. Ask it to operate inside the emails and meetings and docs and spreadsheets and shareepoint sites where your org lives. If your work is coding or model centered reasoning, direct claude or claude code or codecs or cursor or specialist coding environments are probably still your home. The surrounding product matters, but sometimes the model and the developer workflow are the center of the job. If your revops run on Salesforce, headless 360 is a no-brainer. Not because everyone needs to become a Salesforce dev, but because the agents you already use may now be able to act inside the system that already contains your customer and your pipeline and your workflow and your permissions data. If your team has repeatable crosstool workflows that live naturally in chat GPT or Slack, then workspace agents is worth testing, especially when the workflow needs to be shared and scheduled and improved over time and owned by the team rather than a power user.

That is a practical answer and it leads directly to the question I keep hearing in different forms. When should I switch? When should I switch from Claude, from Chad GPT, from Copilot, from Gemini? The reason that question is so tempting is that it feels very clean to ask. Pick a default, move the team, standardize the workflow. But the agent market is not moving toward one default agent for everything. It is moving toward layers.

The Cloud specific version here is worth unpacking because it is one I hear often as Claude gets into more and more teams. You may already be using Claude without switching to Claude directly. If your company uses Microsoft Copilot Co-work, Anthropic's Tech is part of that product. If you use Perplexity Computer, Claude Opus 4.7 is now the default orchestrator. If your Salesforce team uses Agent Force 5's, Claude Sonnet 4.5 is the default model. That's the point. Enthropic enterprise strategy increasingly looks like sitting inside other companies agent stacks. The model becomes a layer inside a product that owns the data, the workflow, the interface, the permissions, or the marketplace. Claude managed agents adds a third shape to this. That's not claude as a chat product. It's not claude hidden inside Microsoft or Salesforce or perplexity. It's anthropic saying if you want longunning agents on claude, but you don't want to build all the infrastructure, let us give it to you as a managed layer. So claude now shows up in at least three ways, right? Direct claude where the model's the product, embedded claude where another vendor owns the workflow and the data layer, and managed claude infrastructure where anthropic gives teams a place to run agent systems without treating chat as the main interface. So the right question is now not really should I switch from Claude or to Claude. Often it is which rapper around Claude fits the job and the framework here is bigger than Claude. It applies whether your default is chat GPT or co-pilot or Gemini or something else.

There there are really just three questions.

Question one, when should you stay in your default agents direct product? Stay there when the model is the center of the work and the surrounding integrations are secondary. Coding, long context reasoning, novel research, custom agents where you are building your own workflow logic, tasks where you want direct control over the model's behavior, and the wrapper is not adding a lot. If you're a Claude user, it's why claude code remains such an important product for engineering teams and why co-work is valuable. If you're a chat GPT user, direct chat GPT or direct codec still makes sense for many open-ended reasoning tasks that don't need to become a recurring team workflow. If you're a C-pilot user, this is often the bucket where you should look outside co-pilot because copilot's strongest advantage is the integration, not the model.

Question two, when should you use a different product that happens to run the same underlying model? Use the wrapper when the wrapper gives you data access or workflow integration you cannot realistically reproduce yourself. Copilot with work IQ gives anthropic style agent execution against the Microsoft 365 graph in a way that direct cla with a connector cannot replicate. Salesforce with Claude inside inherits Salesforce permissions and metadata and business logic and trust boundaries. Perplexity with Claude as an orchestrator gives you a research and artifact workflow that is different from a blank claude chat. In those cases, you're not really switching from your default model. You're moving the model into the product layer that has the right data fabric.

Question three, when should you use a product that runs a different underlying model entirely? Use the different model when the surrounding product matters more than the marginal difference in model quality. So chat GBT workspace agents may be the right choice for a Slack native recurring team workflow even if another model would write a slightly better paragraph in isolation. Google Gemini and Workspace may be the right choice for teams that live in Google because it inherits the graph. Self-hosted Kimmy K 2.6 may be the right choice for a dev team that wants openweight agent capability without a closed lab dependency.

So the model matters but the model is no longer the only thing that matters. The wrapper, the graph, the permissions, the connectors, the workflow surface, the ecosystem, and stackability often matter more. That's the reframe your team needs to get on board with. The switching costs are real. Prompts do not transfer cleanly. Memory and context don't port cleanly. Skills are more portable than they used to be, but they're not magically plugandplay. Team habits matter a lot. If your team has spent six months learning how to specify work to one agent, moving them to another product can restart a lot of that learning. So, don't switch casually. Layer deliberately. Keep your default where it works best. Add specialists where the specialist wins really clearly and build the judgment to route work based on the shape of the task. That judgment is the new literacy of the agent era.

The agent layer is not one product category. That's the trap we fall into. People see AI agent and they assume all these launches belong on one big comparison chart and they try and build it. I've seen them try. They they don't belong on one chart. Some agents are model products. Some are workflow builders. Some are enterprise data layers. Some are openw weightight infrastructure. Some are research workers. Some are wrappers around other models. Some are control planes for systems your company already uses. The launches aren't equal because they're not trying to do the same job.

So that five question filter is a way to simplify that. Does it plug into the tools your team uses? Does it let other agents build on top? Does it own or access data you care about? Is there an ecosystem forming around it? Can you stack agents on top? Now, if the answer is yes across all those questions, the launch probably deserves a lot of attention from you. If it's no, it may be impressive, but it's probably not for you right now. And that is how I would think about the rest of this year. Filter for infrastructure over features. Filter for ecosystems over demos. Filter for stackability over walled gardens. Filter for data access over benchmark charts. And then match the shape of the work to the shape of the tool. That's the whole game right now. And the teams that learn to route work across those layers are going to compound faster than the teams that are chasing whichever model or agent had the loudest launch.

If this was useful, give me a subscribe. I also wrote the full version of this argument with the filter and the product breakdown over on Substack and a lot of specific guides to help you pick agents and get started. I'll link that below. Cheers.