Guardians of Agents

Product survey · September 27, 2026

AI Agent Products, 2026

The agents people actually run in 2026: self-hosted OpenClaw and Hermes Agent, personal agents such as Meta's Muse, the ChatGPT, Gemini and Claude assistants, agentic browsers like Aside and Comet, coding agents like OpenCode and Claude Code, app builders like Lovable and Replit, and work agents. 173 products with history, users, models, cost, capability, security record and the precautions that matter.

173agent products surveyed
63act autonomously with real side effects
40run fully offline on your hardware
167documented security findings
77major incidents in the timeline
8segments in 4 layers

What the agent market looks like, in ten points

Capability, adoption and exposure are all rising together. Every number below is attributed in the product profiles.

  1. Personal agents are the new front line, and Meta's Muse is the one to watch.

    Meta's Muse (Sep 8, 2026) is a personal agent, not a model: it runs in a cloud computer assigned to each user, books, buys and negotiates, and reached #1 on both U.S. app stores within ten days. It also has the most detailed published security design of any consumer agent (a separate Sentinel agent gates every outbound request), and still had two flaws disclosed in three weeks. Google, OpenAI, Apple and Amazon are all answering it; see the watchlist.

  2. Capability crossed real thresholds this year.

    The best systems now exceed the human baseline on the main computer-use benchmark, SWE-bench Verified is effectively saturated, and METR's measure of how long a task frontier agents can complete reliably reached at least 16 hours, doubling roughly every four months. Several leaderboard figures are vendor-reported; the benchmarks section says which.

  3. Security incidents track autonomy, not popularity.

    The 173 products here carry 167 documented security findings. The most autonomous products have the most: OpenClaw, Perplexity Comet, Claude Code, Cursor, OpenCode and the retired ChatGPT Atlas. The attack is almost always the same: untrusted content, access to private data, and a way to send it out.

  4. OpenClaw is the adoption story and the cautionary tale.

    It went from a hobby project in November 2025 to one of GitHub's most-starred repositories and more than 13 million npm downloads a month. It also accumulated hundreds of CVEs, hundreds of malicious skills on its registry and well over a hundred thousand internet-exposed instances, with sandboxing off by default.

  5. Hermes Agent is the memory-first alternative.

    Nous Research's agent keeps curated memory and writes its own reusable skills, runs on its open Hermes models or any endpoint, and has a much smaller CVE record. An independent audit still found critical issues, and Nous says plainly that the operating system is the only real security boundary.

  6. Meta is back in open weights, and it matters for your GPU.

    Muse Spark (April 2026) is proprietary and powers Meta AI. Muse Glimmer (August 2026) is an open Apache 2.0 30B model built for agents that runs on a single 24-32 GB card, which means a capable local agent stack now fits on one 32 GB consumer GPU such as an RTX 5090.

  7. Vibe coding became a mass market, and its security debt is in the apps, not the builders.

    Lovable, Replit, Base44 and Emergent report nine-figure revenue run-rates, and every major platform now has an app builder. Independent scans of thousands of vibe-coded apps found exposed databases, missing access rules and hard-coded keys; Lovable's missing row-level security (CVE-2025-48757) and Base44's authentication bypass are the reference cases. Open-source builders (Dyad, bolt.diy) and terminal agents (OpenCode, Crush) run fully on your own GPU.

  8. Platform agents change names faster than people learn them.

    Operator became ChatGPT agent and then ChatGPT Work; Atlas and Pulse were retired; Project Mariner was folded into Chrome. Anything built on a specific agent product's name should expect to move.

  9. Prices converged on a few tiers.

    Consumer agents cluster at $20 and $100-200 a month; work agents charge per user or per action, from $15 a user for governance to ten cents an action; the open-source agents are free but you pay for tokens or your own GPU.

  10. User numbers describe platforms, not agent use.

    ChatGPT reports more than 900 million weekly users and the Gemini app a billion monthly, but neither discloses how many use agent mode. The honest adoption signals for agents are coding-agent revenue, paid seats and open-source downloads.

Market map

Four layers, eight segments. Within each segment, leaders come first. Click any entry for its full profile. The small square is the self-hosting score: darker means you can run it yourself.

IndependentAcquiredMajor / publicOpen source
Self-hosted & openagents you install and run yourself
Self-hosted personal agents 20
Assistants, personal agents & browsersconsumer agents from the big platforms and the personal-agent newcomers
Personal AI agents 15
Platform assistants 15
Agentic browsers & computer use 33
Building softwarecoding agents for developers, and app builders for everyone else
Coding agents 26
App builders (vibe coding) 36
Work & researchagents inside business apps and research tools
Work agents 17
Research agents 11

Self-hosting score: 0 cloud only 1 vendor-hosted private 2 customer cloud (BYOC) 3 on-prem 4 offline / own hardware

What you can run yourself

Which agents you can run on your own machine. Score 4 means fully offline on your hardware with a local model, 3 an on-premises install, 2 your own cloud account, 1 a vendor-hosted private instance, 0 their cloud only.

Cloud only → runs on your own hardwareEntries · best self-hosted optionsSelf-hosted personal agents4 on-prem416 offline / own hardware1620 · Home Assistant Assist, Open WebUI, Anything…Personal AI agents10 cloud only105 vendor-hosted private515 · nonePlatform assistants12 cloud only121 on-prem12 offline / own hardware215 · Kimi (agent mode), AutoGLM, Meta Muse Glimm…Agentic browsers & computer use18 cloud only184 vendor-hosted private41 customer cloud (BYOC)2 on-prem28 offline / own hardware833 · Browser Use, Skyvern, StagehandCoding agents6 cloud only65 vendor-hosted private51 customer cloud (BYOC)13 on-prem311 offline / own hardware1126 · OpenAI Codex, Aider, ClineApp builders (vibe coding)28 cloud only282 vendor-hosted private23 customer cloud (BYOC)33 offline / own hardware336 · Dyad, bolt.diy, GPT-Engineer (legacy)Work agents14 cloud only143 vendor-hosted private317 · noneResearch agents10 cloud only101 on-prem111 · Sakana AI Scientist

A self-hosted agent lab on one GPU

Everything below runs offline on a single 32 GB workstation. The containment column matters more than the model choice.

Model on one 32 GB GPU

  • Meta Muse Glimmer 30B (Apache 2.0) or a 27-32B Qwen at 4-bit
  • Serve with vLLM, llama.cpp or Ollama on an OpenAI-compatible port
  • Keep 64K context: agent loops are context-hungry
  • Bind the endpoint to localhost only

Personal agent

  • Hermes Agent or OpenClaw pointed at the local endpoint
  • Or a hardened fork: NanoClaw (containers), IronClaw (WASM sandbox), ZeroClaw
  • Turn sandboxing on; it is off by default in OpenClaw
  • Install no third-party skills you have not read

Coding and browser

  • Aider, Cline, OpenHands or goose with the local model
  • Browser Use or Skyvern in a separate browser profile with no saved logins
  • Open untrusted repositories only in a container or VM
  • Never run with skip-permission flags on your main machine

Contain it

  • A dedicated VM or container per agent, no host mounts
  • Egress allowlist: the agent can reach only what the task needs
  • Separate low-privilege accounts and API keys; rotate them
  • Log every tool call; review before anything irreversible

Taxonomy: layers, segments and methods

How the market divides the problem. Chips on the right are the most common technical methods in each segment, with the number of entries using them.

Agent products173 entriesSelf-hosted & open20 entriesSelf-hosted personal agents20 · OpenClaw, Hermes and the open runtimeslocal model 18MCP tools 12runs code 11chat-app gateway 6Assistants, personal agents & browsers63 entriesPersonal AI agents15 · Meta Muse, Gemini Spark, Siri, Alexa+, Instinct, Pokepersonal agent 12persistent memory 11proactive 8voice 6Platform assistants15 · ChatGPT, Gemini, Claude, Copilot, Meta AIcomputer use 7multi-agent 7cloud sandbox 6drives a browser 5Agentic browsers & computer use33 · agents that click for youdrives a browser 22approval prompts 14tool calling 11guardrails 11Building software62 entriesCoding agents26 · the most adopted agents of allwrites code 26autonomous coding 24MCP tools 24bring your own model 16App builders (vibe coding)36 · Lovable, Replit, Bolt, v0, Base44, Emergent, Figma Makebuilds whole apps 35writes code 33hosts and deploys 28adds a backend 19Work & research28 entriesWork agents17 · agents inside the business appstool calling 12agent builder 12retrieval 8MCP tools 5Research agents11 · deep research and AI scientistsweb search 5multi-agent 4runs code 3MCP tools 3

The methods, explained

MCP tools 65

Uses the Model Context Protocol to call external tools and data sources.

writes code 60

Generates or edits source code.

local model 47

Can run entirely against a model on your own hardware.

drives a browser 36

Clicks, types and reads web pages on your behalf.

bring your own model 36

Works with any model endpoint you point it at.

builds whole apps 35

Turns a description into a working web or mobile app, not just code snippets.

multi-agent 34

Several agents or sub-agents split and coordinate the work.

tool calling 28

The model invokes functions and APIs directly.

autonomous coding 28

Takes a task and produces a working change with little supervision.

hosts and deploys 28

Publishes the app to a live URL on the vendor's hosting.

approval prompts 26

Pauses for a person to approve risky actions.

computer use 23

Controls a whole desktop through screenshots, mouse and keyboard.

runs code 22

Executes code or shell commands it wrote.

adds a backend 19

Creates the database, auth and server functions, often on Supabase or its own cloud.

persistent memory 17

Remembers across sessions and builds up context about you.

retrieval 15

Looks up your documents or the web to ground answers.

agent builder 15

Lets users assemble their own agents.

runs on a schedule 14

Starts itself at set times or on triggers, without a prompt.

guardrails 13

Filters or policies on inputs, outputs or actions.

cloud sandbox 13

Runs the work in a disposable cloud machine.

mobile apps 13

Builds native or React Native apps for phones.

voice 12

Speaks and listens.

chat-app gateway 12

You talk to it through WhatsApp, Telegram, Slack and similar.

personal agent 12

Acts on the user's own accounts and errands: mail, bookings, shopping, bills.

sandboxed 10

Isolates tool execution from the host by default.

vision grounding 8

Reads the screen visually to decide where to click.

proactive 8

Starts things on its own: briefings, reminders, follow-ups, suggestions.

web search 7

Searches and reads the web.

open weights 7

The model's weights are downloadable.

workflow automation 6

Automates a defined multi-step process.

shops and pays 6

Completes checkout with a wallet, stored card or single-use card.

deep research 5

Runs long multi-source research and writes a report.

code completion 5

Inline suggestions as you type.

mail and calendar 5

Reads, drafts or sends email and manages your calendar.

embedded 4

Built into an app you already use.

design to code 4

Starts from a design, screenshot or Figma file.

self-improving 3

Writes its own skills and memory as it works (not retraining).

literature search 3

Searches and synthesizes academic papers.

science agent 3

Proposes and tests hypotheses.

glasses and devices 1

Runs on smart glasses or a dedicated pocket device.

Consolidation: who bought whom

7 acquisitions in this data. Prices marked reported come from press sources, not the companies.

AcquirerTargetSegment · date · priceMicrosoft (1)GitHub CopilotCoding agents · 2018-10 · $7.5BWix (1)Base44App builders · 2025-06 · $80M (reported)Atlassian (1)DiaBrowsers & computer use · 2025-09 · $610M (reported)team.blue (1)MacalyApp builders · 2025-12 · undisclosedSpaceX (1)Grok (agents)Platform assistants · 2026-02 · undisclosedCognition (1)PokePersonal agents · 2026-07 · undisclosedSpaceX (folded into its SpaceXAI / xAI unit) (1)CursorCoding agents · 2026-08 · $60B (reported)

Where the money is

Filled dots: venture funding raised. Rings: price paid for an acquired company. Squares: last valuation when funding is not public. Log scale; hover a dot for the figure. Public companies and open-source projects are left out.

$1M$10M$100M$1B$10BSelf-hostedagentsLetta: raised $10MAutoGPT: raised $12MHermes Agent: raised $70MAutoGPTHermes AgentPersonalagentsPoke: raised $25MInstinct: raised $350MPokeInstinctPlatformassistantsGenspark Super Agent: valued at $1.2BPerplexity Computer & Labs: valued at $21.2BGenspark Super AgentPerplexity Computer & LabsBrowsers &computer useStrawberry: raised $6MYutori: raised $15MBrowser Use: raised $17MDia: acquired for $610MPerplexity Comet: valued at $21.2BDiaPerplexity CometCodingagentsAugment Code: valued at $977MDevin / Windsurf: valued at $26BCursor: acquired for $60BDevin / WindsurfCursorAppbuildersVibecode: raised $9MRocket.new: raised $15MAtoms: raised $31MBase44: acquired for $80MAnything: valued at $100MEmergent: raised $230MLovable: valued at $13.3BEmergentLovableWorkagentsDust: raised $60MGlean: raised $765MWRITER AI HQ: valued at $1.9BDecagon: valued at $4.5BSierra: valued at $15BDecagonSierraResearchagentsElicit: valued at $100MKosmos (Edison): valued at $250MPerplexity Deep Research: valued at $21.2BKosmos (Edison)Perplexity Deep Research

Segment scorecard

Size, ownership mix, maturity spread, license mix and money for each segment. Maturity: 1 research, 2 early product, 3 generally available, 4 scaled, 5 category standard.

SegmentEntries by statusMaturity 1–5 (count)LicenseMoney trackedSelf-hosted personal agents20Independent: 88Open source / nonprofit: 1212maturity 1: 1maturity 2: 6maturity 3: 5maturity 4: 6maturity 5: 2open_core: 2open_source: 18$92M raisedPersonal AI agents15Independent: 44Acquired: 1Major / public: 1010maturity 2: 6maturity 3: 8maturity 4: 1proprietary: 15$375M raisedPlatform assistants15Independent: 77Major / public: 77Open source / nonprofit: 1maturity 2: 4maturity 3: 8maturity 4: 3proprietary: 12open_core: 1open_source: 2n/a (public / OSS)Agentic browsers & computer use33Independent: 1515Acquired: 1Major / public: 1515Open source / nonprofit: 2maturity 1: 3maturity 2: 16maturity 3: 12maturity 4: 2proprietary: 21open_core: 5open_source: 7$38M raised · $610M in dealsCoding agents26Independent: 1111Acquired: 1Major / public: 1010Open source / nonprofit: 44maturity 2: 3maturity 3: 9maturity 4: 10maturity 5: 4proprietary: 12open_core: 3open_source: 11n/a (public / OSS)App builders (vibe coding)36Independent: 1818Acquired: 2Major / public: 1111Open source / nonprofit: 55maturity 1: 1maturity 2: 9maturity 3: 20maturity 4: 6proprietary: 30open_core: 1open_source: 5$285M raised · $80M in dealsWork agents17Independent: 99Major / public: 88maturity 2: 1maturity 3: 9maturity 4: 6maturity 5: 1proprietary: 17$825M raisedResearch agents11Independent: 55Major / public: 55Open source / nonprofit: 1maturity 1: 1maturity 2: 2maturity 3: 2maturity 4: 5maturity 5: 1proprietary: 10open_source: 1n/a (public / OSS)12345
IndependentAcquiredMajor / publicOpen sourceProprietaryOpen coreOpen-source license

How big is the market?

Market sizings for 'AI agents' are young and inconsistent; the analyst predictions under the chart matter more than the dollar figures.

$10B$100BGrand View ResearchGlobal AI agents market size$182.9B by 2033 (49.6%/yr)$7.6B (2025)MarketsandMarketsGlobal AI agents market size$52.6B by 2030 (46.3%/yr)$7.8B (2025)

Sources: MarketsandMarkets · Grand View Research · IDC · Gartner · Gartner · Gartner · McKinsey · Gartner

The segments

What each segment does, how it works, how mature and effective it is, who benefits, and what you can self-host.

Self-hosted & open agents you install and run yourself

Self-hosted personal agents

20 entries: 8 independent · 12 open source / nonprofit

Open-source agents you install on your own computer or server and talk to through a chat app, a terminal or a web UI. They run shell commands, manage files, browse, schedule themselves and remember you across sessions.

How it works
A local gateway process connects a model (hosted or local) to tools on the host: shell, files, browser, MCP servers, plus a skills or plugin registry. Most can point at any OpenAI-compatible endpoint, including a model on your own GPU.
Maturity
The fastest-growing and least mature layer. OpenClaw went from a hobby project in November 2025 to one of the most-starred repositories on GitHub within months; Hermes Agent followed in February 2026. Hardened forks (NanoClaw, ZeroClaw, IronClaw) appeared in response to the security record.
How well it works
The most capable agents on this page in raw autonomy, because they run with your permissions on your machine. That is also why they have the worst security record: sandboxing is usually off by default and third-party skills are an open supply chain.
Who benefits
Developers, researchers and power users who want a personal agent they control; anyone who cannot send data to a vendor.
Self-hosting
All of them. The hardened forks (NanoClaw with containers, IronClaw with a WebAssembly sandbox) are the safer starting points.
Strengths
  • Runs fully offline with a local model
  • Free, inspectable, model-agnostic
  • Highest autonomy: shell, files, schedules, memory
Limits
  • Host access by default; sandbox is opt-in
  • Skill and plugin registries have shipped malware
  • Exposed instances and hundreds of CVEs

Leaders:

Runs on your own hardware:

Assistants, personal agents & browsers consumer agents from the big platforms and the personal-agent newcomers

Personal AI agents

15 entries: 4 independent · 1 acquired · 10 major / public

Agents that run errands in your own life: they read your mail and calendar, book, buy, negotiate, cancel subscriptions and follow up on their own. You talk to them like a person, in their own app, in WhatsApp or iMessage, by voice, and soon through glasses and pocket devices.

How it works
Most run in a vendor cloud computer assigned to you (Meta's Muse Secure VM, Google CC's own Google account and cloud computer, Gemini Spark on Google Cloud) with a browser, connectors to your accounts and a memory of your preferences. Payment goes through wallets or single-use cards, and a policy layer asks before sending, buying or sharing.
Maturity
The newest segment and the most contested. Meta's Muse (Sep 8, 2026) went to #1 on both U.S. app stores within ten days; Instinct raised $250M at a $2.5B valuation while still invite-only; Cognition bought Poke; Siri AI shipped as a beta on Sep 14. Amazon blocked Muse from shopping on Amazon.com within two weeks.
How well it works
They save real time on dull errands (refunds, bills, bookings), and reviewers report money saved. They also concentrate risk: one agent holds your inbox, your payment method and your accounts, and reads untrusted web pages and email. Muse had two security flaws disclosed in its first three weeks; Instinct acted without asking and misreported what it did.
Who benefits
Busy consumers and families; the platforms themselves, which want to own the layer between people and every app and store.
Self-hosting
No. The open-source equivalents are the self-hosted agents in the first layer (OpenClaw, Hermes Agent) connected to a messaging app, which trade the vendor's safety engineering for control over your own data.
Strengths
  • Handles errands end to end, including phone calls and negotiation
  • Talk to it where you already are: WhatsApp, iMessage, voice
  • Best-documented security designs in consumer AI (Muse Sentinel, Siri Private Cloud Compute)
Limits
  • Your mail, cards and accounts in one agent is a single point of compromise
  • Retailers block or fight agent shoppers (Amazon vs. Muse, Amazon vs. Comet)
  • Mostly U.S.-only, cloud-only and changing weekly

Leaders:

Runs on your own hardware: —

Platform assistants

15 entries: 7 independent · 7 major / public · 1 open source / nonprofit

The general-purpose assistants from the big platforms, now with agent modes that browse, fill forms, run tasks on a schedule and work in a cloud computer while you are away.

How it works
A frontier model in the vendor's cloud with a hosted browser or virtual machine, connectors to your mail, calendar and files, and confirmation prompts before consequential actions.
Maturity
Mass-market and renamed often: OpenAI's agent went from Operator to ChatGPT agent to ChatGPT Work in eighteen months, and Google folded Project Mariner into Chrome. Meta moved Meta AI from Llama to its proprietary Muse Spark in April 2026.
How well it works
The easiest agents to use and the best-resourced on safety, with published prompt-injection rates. They cannot be self-hosted, and connecting them to your accounts is exactly what makes them attackable.
Who benefits
Everyone; the default agent for most people.
Self-hosting
No, with one notable exception: Meta's Muse Glimmer is an open-weight Apache 2.0 model built for agents and runs on a single 24-32 GB GPU.
Strengths
  • Strongest models, no setup
  • Vendors publish safety measurements
  • Hundreds of millions of users means fast fixes
Limits
  • Cloud only; your data goes to the vendor
  • Agent features change name and scope often
  • Connectors create the injection path

Leaders:

Runs on your own hardware:

Agentic browsers & computer use

33 entries: 15 independent · 1 acquired · 15 major / public · 2 open source / nonprofit

Browsers with an agent built in, browser extensions that drive your tabs, and computer-use models that operate a whole desktop through screenshots.

How it works
A vision-capable model reads the page or screen, plans, and issues clicks and keystrokes, usually inside your logged-in browser session.
Maturity
Churning fast: OpenAI retired Atlas in August 2026 and Google shut down Project Mariner in May. Capability is now above the human baseline on the main computer-use benchmark.
How well it works
Now reliable on ordinary web tasks. This is the segment with the richest record of prompt-injection attacks, because every page the agent reads can carry instructions and the agent acts with your logins.
Who benefits
Individuals automating web chores; teams testing web apps; researchers.
Self-hosting
Browser Use, Skyvern, Stagehand, Nanobrowser, UI-TARS, Fara-7B and Agent S run with local models.
Strengths
  • Does real web work end to end
  • Open-source options run locally
  • Rapidly improving accuracy
Limits
  • Every web page is untrusted input
  • Acts with your logged-in sessions
  • Several vendors call injection 'by design'

Leaders:

Runs on your own hardware:

Building software coding agents for developers, and app builders for everyone else

Coding agents

26 entries: 11 independent · 1 acquired · 10 major / public · 4 open source / nonprofit

Agents that read a repository, write and run code, execute tests and open pull requests, in the terminal, the IDE or a cloud sandbox.

How it works
A model with file, shell and git tools, permission prompts and optional sandboxes; cloud variants run in disposable containers and return a pull request.
Maturity
The most commercially successful agents: multi-billion-dollar revenue run-rates, and the SWE-bench Verified benchmark is effectively saturated. Consolidation is heavy, including SpaceX's $60B purchase of Cursor.
How well it works
Strong on well-specified changes. The recurring security flaw is that opening an untrusted repository lets its configuration files steer or execute the agent, and supply-chain attacks now target the agents directly.
Who benefits
Every software team; students learning to program; and attackers, who now use them too.
Self-hosting
Aider, Cline, Kilo Code, OpenHands, goose, Qwen Code, Kimi CLI, Zed; Codex CLI and Gemini CLI are open source but default to their vendor's model.
Strengths
  • Clear productivity gains on routine changes
  • Many run against local models
  • Sandboxed cloud modes available
Limits
  • Untrusted repos can hijack the agent
  • Destructive actions have happened in production
  • Skip-permission flags get abused by malware

Leaders:

Runs on your own hardware:

App builders (vibe coding)

36 entries: 18 independent · 2 acquired · 11 major / public · 5 open source / nonprofit

Tools that turn a plain-language description into a working web or mobile app, with a database, login and hosting, so people who do not write code can ship software. Many now run long agent loops that plan, build, test and fix on their own.

How it works
An agent in the vendor's cloud writes a full-stack project (usually React plus a hosted database such as Supabase or the vendor's own), runs it in a sandbox, shows a live preview and deploys it to a public URL. Open-source versions (Dyad, bolt.diy) run on your machine with a local model.
Maturity
The fastest-growing software category of 2025-26. Lovable, Replit, Base44 and Emergent report nine-figure revenue run-rates, Wix bought Base44 in six months, and Google, Microsoft, Figma, Canva, Salesforce, Baidu and Ant all shipped their own. Several early players have already shut down or been absorbed (GitHub Spark, Mocha, Firebase Studio).
How well it works
Excellent for prototypes, internal tools and simple sites. The weak point is what gets shipped: scans of thousands of vibe-coded apps found exposed databases, missing access rules and leaked keys at scale, and the platforms' free hosting is abused for phishing. The builder is secure enough; the apps people publish often are not.
Who benefits
Founders, designers, product managers, students and small businesses; enterprises for internal tools under IT governance.
Self-hosting
Yes: Dyad and bolt.diy run on your own machine with local models; Open Lovable, Onlook and Chef are open source but lean on cloud services.
Strengths
  • Idea to deployed app in minutes
  • Non-programmers can build real tools
  • Open-source options run fully local
Limits
  • Generated apps often ship without access control
  • Lock-in to the vendor's hosting and database
  • Credit-based pricing makes cost hard to predict

Leaders:

Runs on your own hardware:

Work & research agents inside business apps and research tools

Work agents

17 entries: 9 independent · 8 major / public

Agents built into productivity suites, CRMs, service desks and knowledge tools, acting across company data with the user's permissions.

How it works
Retrieval over company data plus actions through the platform's APIs, with admin controls, per-user or per-action pricing and governance consoles.
Maturity
Commercially large: tens of millions of paid seats and billion-dollar agent revenue lines. Security incidents follow the same pattern across vendors.
How well it works
Useful for search, summarization and routine updates. Each major suite has had a zero-click or one-click data-exfiltration finding in which a document or form carried instructions the agent obeyed.
Who benefits
Knowledge workers, service and sales teams, IT administrators.
Self-hosting
None; these are platform features.
Strengths
  • Uses permissions the company already manages
  • Admin governance and audit built in
  • No integration work
Limits
  • Cloud only
  • Poisoned documents can steer the agent
  • Some risky defaults are 'by design'

Leaders:

Runs on your own hardware: —

Research agents

11 entries: 5 independent · 5 major / public · 1 open source / nonprofit

Agents that run long, multi-source research and write cited reports, and newer systems that propose and test scientific hypotheses.

How it works
Iterative search, reading and synthesis over the web or the literature, sometimes with code execution for analysis.
Maturity
Widely used for research reports; the autonomous-science systems are early and mostly self-evaluated.
How well it works
Saves hours on literature and market scans. Fabricated citations are measurable and higher than in ordinary search-augmented chat, and one exfiltration attack ran entirely from the vendor's servers.
Who benefits
Researchers, analysts, students and grant writers.
Self-hosting
Few; open deep-research frameworks exist but the leading products are hosted.
Strengths
  • Hours of reading compressed to minutes
  • Citations make claims checkable
  • Specialist tools for the academic literature
Limits
  • Fabricated or wrong citations
  • Evaluations mostly vendor-run
  • Server-side browsing bypasses your network controls

Leaders:

Runs on your own hardware:

Full table

All 173 entries. Filter, sort by any column, click a row for the full profile with sources. Download: CSV · Excel.

ProductSegmentOwnerLaunchedUsers / starsAutonomySelf-hostSecurity issuesCostModels

Spotlight: OpenClaw, Hermes and Muse

The viral open agent, the memory-first alternative, and Meta's Muse, the personal agent that went to #1 on both U.S. app stores within ten days of its September 8 launch.

Self-hosted personal agents · since 2025-11

OpenClaw

Started by Austrian developer Peter Steinberger as 'Warelay' (Nov 24, 2025), renamed CLAWDIS (Dec 3, 2025), Clawdbot (Jan 2, 2026), Moltbot (Jan 27, 2026, after an Anthropic trademark complaint) and OpenClaw (Jan 30, 2026). Went viral in late Jan 2026; Steinberger joined OpenAI on Feb 14, 2026 and the project moved to an independent OpenClaw Foundation (501(c)(3)). v2026.8.1 'OpenClaw 2.0' (Aug 30, 2026) touched every subsystem with ~16-17k PRs from 900+ contributors.

Users
3.2M monthly active users (SimilarWeb estimate via gradually.ai, Sep 22 2026 - reported, not disclosed by the project); 13,585,228 npm downloads in 30 days to Sep 25 2026 (npm registry); 2.3M ClawHub skill installs (reported).
Models
Claude (Anthropic), OpenAI GPT/Codex, DeepSeek, OpenRouter models (StepFun, Kimi, GLM, Qwen etc.), local via Ollama, LM Studio, llama.cpp, vLLM, SGLang, MLX
Cost
Self-hosted: free software + your compute or API tokens. Heavy agent use on frontier APIs can cost tens to hundreds of USD/month; local models cost only electricity.
Autonomy · self-host
level 5 of 5 · self-host score 4
Security record
13 documented findings. Explicitly a single-trusted-operator 'personal assistant' model: host is treated as trusted, sandbox is opt-in, and prompt injection alone is out of scope for vulnerability reports. The 2.0 release ties approvals to specific requests and hides credentials from model-visible text, but multi-tenant isolation and skill vetting remain the user's problem.
Do this first
  • Run `openclaw security audit` (with auto-fix) after install and after upgrades
  • Keep the gateway bound to loopback; never expose the Control UI/WebSocket to the internet without a reverse proxy that authenticates
  • Turn on sandboxing (agents.defaults.sandbox.mode is 'off' by default) with Docker/Podman/SSH backends; keep exec approvals on
  • Keep DM pairing on and use sender allowlists; isolate DM sessions per user

Full profile, sources and every finding: click the card.

Self-hosted personal agents · since 2026-02

Hermes Agent

Released by Nous Research (known for the open-weight Hermes LLM family) on Feb 25, 2026 as 'an autonomous agent that lives on your server, remembers what it learns, and gets more capable the longer it runs'. Shipped ~19 minor versions in five months: v0.11 'Interface Release', v0.18 'Judgment Release' (Jul 1 2026: Mixture-of-Agents, completion contracts, /learn, /journey memory timeline) and v0.19 'Quicksilver' (Jul 20 2026: ~80% faster first token, smart approvals default, Bitwarden/1Password integration). Its growth drove Nous's reported $1.5B-valuation round in Jul 2026.

Users
No official user count. 1.56M+ PyPI installs May-Sep 2026 (gradually.ai, reported Sep 2026); 240k GitHub stars (Sep 2026).
Models
Nous Hermes family (Hermes-4-405B/70B/14B, Hermes-4.3-Seed-36B) via Nous Portal or local, 300+ models via Nous Portal, OpenRouter, OpenAI, Anthropic, any OpenAI-compatible custom endpoint (Ollama, vLLM, llama.cpp), Mixture-of-Agents as a selectable model
Cost
Self-hosted: free + your compute or API tokens. Nous Portal hosted plans $0 / $20 / $100 / $200 per month.
Autonomy · self-host
level 5 of 5 · self-host score 4
Security record
4 documented findings. Security policy states 'the only security boundary against an adversarial LLM is the operating system'; approval gates and redaction are heuristics, and prompt injection without a chained outcome and third-party skills are out of scope. No bug bounty; 90-day coordinated disclosure.
Do this first
  • Use a Docker/SSH/Modal terminal backend, or wrap the whole process in Docker or NVIDIA OpenShell (the project's two supported isolation postures)
  • Keep smart approvals on (default since v0.19) and review memory entries via /journey
  • Rely on built-in memory-write scanning for injection/invisible Unicode, but review agent-written skills before reuse
  • Use Bitwarden/1Password integration instead of plaintext keys

Full profile, sources and every finding: click the card.

Personal AI agents · since 2026-09

Meta Muse

Built by Meta Superintelligence Labs on its Muse Spark model family (Muse Spark 1.1 and a Meta Model API preview released Jul 9, 2026). Launched Sep 8, 2026 in the U.S. on iOS, Android, muse.ai and WhatsApp, and later available in Canada (TechCrunch, Sep 25). At Connect (Sep 23, 2026) Meta added a Muse Realtime Avatar, Mac computer control, retail/payment connectors (Walmart, Best Buy, Sephora, Wayfair, Shop Pay, PayPal, Instacart, Notion, GitHub, Box, Expedia coming), its own email address for Muse, glasses integration 'in the coming months', and the Muse Charm keychain/pendant device (holiday target, no price). Amazon blocked Muse from its store by Sep 22, 2026, calling it an unauthorized agent that does not identify itself and captures credentials (GeekWire).

Users
No user count disclosed. Downloads (Sensor Tower estimates): >2.5M by Sep 21, 2026 (CNBC); 3.4M by Sep 25, 2026 (TechCrunch). DAU rose 27% the day after Connect (TechCrunch, Sep 25; no absolute figure).
Models
Muse Spark family (Meta Superintelligence Labs; Muse Spark 1.1 released Jul 2026), Muse Realtime Avatar model (avatar/voice conversation), separately trained prompt-injection classifier ensemble
Cost
Free (payment card required); Power $20/mo; Maximum $100/mo (disclosed via TechCrunch/CNBC). Purchases made by the agent are charged to the user's Link/Shop Pay/PayPal wallet; Meta plans a small transaction fee.
Autonomy · self-host
level 4 of 5 · self-host score 0
Security record
3 documented findings. Meta designs to bound damage rather than prevent every attack: per-user VM, a separately trained Sentinel that alone approves network egress and connector actions, just-in-time credentials, taint-aware approvals, injection classifiers, and a bug bounty up to $300K; Meta states prompt injection remains unsolved and Muse will make mistakes. It does not cover a compromised local client (Mac dictation hijack), Meta's own access to VM data until Confidential VM ships, human contractors in the concierge call test, or users approving warnings they don't understand.
Do this first
  • Connect only the apps a task needs and pick the narrowest access level Muse offers (Meta separates read vs write where the service supports it); disconnect connectors when done
  • Use read-only email access unless you actually want Muse to send; use the forthcoming Muse email address and forwarding instead of full inbox access where possible
  • Keep approvals one-time or task-scoped rather than 'perpetual'; never approve a warning about a suspicious site mid-task
  • Pay only through Stripe Link single-use cards (merchant-, amount- and time-bound) or wallet connectors; do not paste card numbers into chat

Full profile, sources and every finding: click the card.

Which "Muse" is which

Meta uses one brand for an agent, two models, a device and a privacy feature. The agent is the product people install; the models are what it runs on.

NameWhat it isWhenOpenness
MusePersonal AI agent: an app, WhatsApp contact and muse.ai site that runs errands for you in its own cloud computerlaunched Sep 8, 2026 (U.S.; Canada later)proprietary, cloud
Muse SparkMeta Superintelligence Labs' proprietary model family; powers Meta AI and the Muse agent (Spark 1.1 with a public API followed on Jul 9, 2026)Apr 8, 2026proprietary, API
Muse GlimmerOpen-weight 30B model distilled from Muse Spark, built for agents; fits one 24-32 GB GPU at 4-bitAug 10, 2026open weights, Apache 2.0
Muse CodeMeta's terminal coding agent on Muse Spark2026, after Spark 1.1 (exact date not verified)proprietary
Muse CharmPocket device for talking to your Muse agent by voice, with a realtime avatarannounced Sep 23, 2026; holiday targethardware, not yet shipping
Muse on AI glassesThe Muse agent on Ray-Ban / Oakley Meta glasses, acting on what you look atannounced Sep 23, 2026; 'coming months'not yet shipping
Muse Confidential VMYour Muse computer encrypted with a key only you hold, so Meta cannot read it'later this year' (trusted testers now)not yet shipping

Personal agents compared

The 15 products in the personal-agent segment, largest reported audience first. "Autonomy" is the 1-5 level used on this page; "Findings" counts documented security findings. Click a row for the full profile, sources and precautions.

AgentHow it worksCostAdoption (as reported)AutonomyFindingsSafety design
Gemini Spark (agent)2026-05A 24/7 cloud agent in the Gemini app that keeps working on Google Cloud after the phone or laptop is closed.Launched Ultra-only (US): $100 or $200/mo; now also listed with AI Pro on Google's plan page.Gemini app: 1B monthly active users (Pichai, Aug 11 2026); ~950M MAU at Q2 earnings (Forbes, Jul 22 2026). No Spark-specific count.42Confirmation gates on purchases and messages, plus Google's layered prompt-injection defenses.
Alexa for Shopping2026-05Amazon's shopping agent in the Amazon app, the website and Echo Show.Free with an Amazon account.300M+ customers used Rufus in 2025 (Amazon)40First-party agent working inside Amazon's own checkout and payment controls.
Superhuman Go2025-10A cross-app AI assistant in the Superhuman suite (Grammarly, Mail, Docs/Coda, Calendar), with Fathom meeting notes added in Sep 2026.Free tier; Pro $12-15/mo; Business $33-40/mo; Enterprise custom.~40M users across the platform (TechCrunch, Sep 2026)30The CEO says judgment should stay human, with the assistant drafting and the user deciding (Observer, Jul 2026).
Meta Muse2026-09Each user gets a dedicated cloud Linux VM (Muse Secure VM) holding the agent, a virtualized Chromium browser, files and connector credentials; a Muse Spark-powered agent plans and executes goals…Free (payment card required); Power $20/mo; Maximum $100/mo (disclosed via TechCrunch/CNBC).No user count disclosed. Downloads (Sensor Tower estimates): >2.5M by Sep 21, 2026 (CNBC); 3.4M by Sep 25, 2026 (TechCrunch). DAU rose 27% the day after Connect…43Meta designs to bound damage rather than prevent every attack: per-user VM, a separately trained Sentinel that alone approves network egress and connector actions, just-in-time credentials,…
Instinct2026-06An invite-only assistant you text or call.Free during invite-only access (Sep 2026).100,000+ users, invite-only (company via press, Sep 2026)42There is little public safety documentation.
Alexa+2025-02An LLM-rebuilt Alexa that orchestrates 'experts' (API groups over tens of thousands of services and devices) and can navigate the web for tasks such as finding, booking and confirming a Thumbtack…Free with Prime; otherwise $19.99/mo.Not published.30Amazon emphasizes privacy controls and AWS infrastructure.
ChatGPT Finances2026-05Links bank, card and brokerage accounts through Plaid (12,000+ institutions) into a read-only Finances dashboard inside ChatGPT.Plus $20/mo or Pro (US).not published20Deliberately read-only, with no transfers or trades.
Google CC2025-12A Google Labs agent with its own verified Google account.Free experiment (no pricing disclosed); US, 18+, personal Google accounts only.not published30Gives the agent its own verified account with clear permissions and group-only responses.
Honor YOYO2024-10A system-level agent in Honor's MagicOS, started from an AI button or voice.Included with Honor devices.not published30Protocol-based access means apps grant capabilities explicitly.
Kanana in KakaoTalk2026-03An agent inside KakaoTalk, Korea's dominant messenger.Free (reported).Target ~10M AI MAU by end-2026 (Kakao, Aug 2026); no current MAU disclosed30Keeps intent detection on-device and uses partner A2A rather than scraping.
Naver Agent N2026-06Naver's 'Agent N' strategy puts agents across its search, shopping and payments services.Free.not published30First-party agent within Naver's own checkout.
Poke2025-09A proactive assistant that lives in iMessage, SMS, WhatsApp and Telegram.Free; Pro $19/mo; Ultra $199/mo plus usage beyond credits.Hundreds of thousands of users; 100M+ messages in three months (company via TechCrunch, Jul 2026)30Describes a 'multi-layered' security model with pen testing and limited employee access to tokens (TechCrunch, Apr 2026).
Rabbit OS32026-09A cloud 'agentic operating system' reached through a web chat, Telegram, iMessage/RCS/SMS or the r1.Pricing not disclosed at launch; users also pay their own model API costs.not published41Says files stay on-device and are not copied, and sensitive actions need confirmation.
Siri / Apple Intelligence2011-10Siri acts inside apps through developers' App Intents (e.g.Free with device.Not published.30Apple built the most verifiable cloud-inference privacy model.
Tencent WorkBuddy/Claw2026-03Tencent's family of agents: ClawBot, a WeChat plugin that makes a user's OpenClaw agent a WeChat contact; QClaw, a consumer OpenClaw launcher for controlling a PC from WeChat/QQ; and WorkBuddy,…Pricing not published in sources reviewed; QClaw stopped new subscriptions (Sep 2026).WorkBuddy 8.85M monthly visits (Mar 2026, BigGo Finance, reported); Yuanbao MAU figures vary by source31WeChat restricts agent access to a dual-authorization protocol and blocks unapproved automation.

The personal-agent race: live now and coming next

Every major platform is answering Muse. Left: what each vendor ships today (click for the profile). Right: what it has announced, or what credible reports and rumors say is coming. An open slot means nothing citable yet; the board is updated as products launch.

VendorLive nowComing next
Meta Muse on AI glasses announced Muse Charm announced Muse Confidential VM beta Muse connectors and extras (Expedia, Shop Pay, 1Password, Muse email address) announced
Google Android XR AI glasses (Google with Warby Parker and Gentle Monster) announced
OpenAI OpenAI device (with Jony Ive's io) delayed OpenAI 'Codex Bot' persistent agent rumored ChatGPT 'o' always-on assistant rumored
AppleSiri AI (beta expansion: languages, EU, China) beta Apple 'HomePad' home hub rumored Apple AI smart glasses (N50) delayed
Amazon Alexa 'Moonraker' agentic upgrade rumored
MicrosoftMicrosoft Copilot app with Autopilot agents beta
Anthropicopen slot: nothing announced yet
SamsungSamsung Galaxy Glasses (Android XR) announced
xAIopen slot: nothing announced yet
Perplexity open slot: nothing announced yet
AlibabaQwen Intelligence phone agent (Honor Magic9 first) announced
ByteDanceopen slot: nothing announced yet
Tencentopen slot: nothing announced yet
Startups open slot: nothing announced yet

vendor-confirmed credible reports rumor

Watchlist details

Meta · feature · announced

Muse on AI glasses

Brings the Muse personal agent to Meta's AI glasses for hands-free use: acting on what the wearer is looking at, guiding workouts, logging meals, booking appointments, checking flight prices and helping buy products seen, with background tasks that report back when done.

Expected
'coming soon' / 'in the coming months' (Meta, Sep 2026); no date given
Versus Muse
This is Muse itself moving onto a camera-and-microphone wearable, the form factor Google (Android XR), Samsung and Apple are racing to match.
Security and privacy
Meta says it is bringing Private Processing to AI glasses so that no one, 'including Meta', sees the data. Watch always-on camera context feeding an agent that can buy things, and prompt injection from things the camera sees.

vendor-confirmed checked 2026-09-27 · meta.com · about.fb.com · about.fb.com

Meta · device · announced

Muse Charm

A small, always-on handheld or pendant (described as Tamagotchi-like) with a real-time voice model, so people can talk to their Muse agent without glasses and 'without having to unlock a phone or open an app'.

Expected
ships by the holidays in December 2026 (Zuckerberg at Connect, Sep 23, 2026); Meta: 'more to share later this year'; price not announced
Versus Muse
A dedicated Muse endpoint that goes head to head with OpenAI's rumored screenless io device and follows the Humane Ai Pin and Rabbit R1 category.
Security and privacy
No security details published. Reports say only a small number of units had been built. Watch always-listening capture, given the Sep 2026 Muse zero-day that could exfiltrate dictations and access tokens (since patched).

vendor-confirmed checked 2026-09-27 · about.fb.com · mashable.com · cnbc.com · tech-insider.org

Meta · feature · beta

Muse Confidential VM

An option that encrypts a user's entire Muse VM, including data and conversations, with a key only the user holds, intended to 'cryptographically and verifiably' prevent Meta from accessing it.

Expected
'later this year' (Meta, Sep 8, 2026); currently with trusted testers
Versus Muse
Muse's own privacy upgrade and a potential differentiator against ChatGPT agent and Gemini Spark, neither of which offers a user-keyed confidential VM.
Security and privacy
Currently in a trusted-tester phase with external auditor review underway. Watch the attestation and key-custody design and how it coexists with Meta's human-concierge call test (rolled back) and ad and commerce incentives.

vendor-confirmed checked 2026-09-27 · about.fb.com · research.meta.ai

Meta · feature · announced

Muse connectors and extras (Expedia, Shop Pay, 1Password, Muse email address)

More Muse integrations: Expedia is coming soon (Walmart, Best Buy, Sephora, Instacart, PayPal and others were added at Connect), Shop Pay as a payment option, 1Password support so Muse can reuse existing logins, and a dedicated Muse email address.

Expected
'coming soon' (Meta, Sep 2026); no dates
Versus Muse
Extends Muse's commerce reach through sanctioned connectors, partly routing around Amazon's Sep 24, 2026 block of Muse from Amazon.com.
Security and privacy
Amazon accused Muse of not identifying itself while browsing and of appearing to capture and store customer credentials. Password-manager integration and an agent-owned inbox widen the credential and prompt-injection attack surface.

vendor-confirmed checked 2026-09-27 · about.fb.com · about.fb.com · thepaypers.com

Apple · agent · beta

Siri AI (beta expansion: languages, EU, China)

Apple's rebuilt Siri, which takes systemwide app actions (drafting email, editing photos, sharing), uses personal context and onscreen awareness, and has a conversation-history app. It runs on Apple Foundation Models that AppleInsider reports were trained with Google Gemini.

Expected
shipped in iOS 27 on Sep 14, 2026 as a beta, English only; Apple 'will quickly expand support for more languages'; EU and China pending regulatory review (Apple, Jun 2026)
Versus Muse
An on-device and OS-level agent for iPhone users rather than a cloud-VM web agent. It is Apple's main answer to Muse, though it does less open-web shopping and browsing.
Security and privacy
Uses on-device processing plus Private Cloud Compute, where Apple says personal data is not stored or accessible to Apple and third-party experts can verify the claim.

vendor-confirmed checked 2026-09-27 · apple.com · appleinsider.com · cnbc.com

Google · device · announced

Android XR AI glasses (Google with Warby Parker and Gentle Monster)

Gemini-powered audio glasses (display models to follow) for hands-free questions about surroundings, directions, texts and missed-message summaries, translation, multi-step tasks through Gemini Intelligence, and voice control of apps such as Uber and DoorDash. They work with both Android and iOS phones.

Expected
audio glasses 'coming later this fall' 2026; display glasses later (Google I/O, May 19, 2026); no price
Versus Muse
Google's direct counter to Muse-on-Ray-Ban-Meta. It pairs with Gemini Spark, Google's 24/7 background agent, which is already live for AI Pro and Ultra subscribers in the US and India.
Security and privacy
Little published beyond 'private' speakers. Gemini Spark asks before high-stakes actions such as spending money or sending email. Watch camera and bystander privacy and how glasses requests reach Spark's cloud agent.

vendor-confirmed checked 2026-09-27 · blog.google · cnbc.com · blog.google

Samsung · device · announced

Samsung Galaxy Glasses (Android XR)

Samsung's Android XR smart glasses, co-developed with Google, with Galaxy Watch gesture control and instant photo sync to Galaxy phones. They are expected to use Gemini, as Galaxy phones already bundle Gemini, Perplexity ('hey, Plex') and a revamped Bixby agent.

Expected
'later this year' 2026 (Samsung to The Korea Times, Jul 2026); price 'reasonable' but premium
Versus Muse
Galaxy's wearable entry point competing with Muse on Meta glasses, with the agent layer drawn from Google and Perplexity rather than Samsung's own agent.
Security and privacy
Samsung says the camera disables automatically if the privacy LED is covered or the glasses are not being worn.

vendor-confirmed checked 2026-09-27 · androidcentral.com · theverge.com · samsungmobilepress.com

Microsoft · platform · beta

Microsoft Copilot app with Autopilot agents

A rebooted Copilot that merges the consumer and work apps into one enterprise-focused app with Cowork, Code and Autopilot tabs. Autopilot creates custom agents that handle assignments and communicate in Teams and Outlook.

Expected
Home and Code tabs for Frontier-program customers 'in the coming weeks'; Autopilot private preview by end of Sep 2026; full launch details at Ignite, Nov 17-20, 2026
Versus Muse
Microsoft is effectively leaving the consumer personal-agent race (Bloomberg: 'abandons personal AI chatbot race') and aiming agents at workplaces rather than at Muse's consumers.
Security and privacy
Enterprise-governed and usage-billed. Watch agent identity and permission scoping in Teams and Outlook.

vendor-confirmed checked 2026-09-27 · cnbc.com · pymnts.com · bloomberg.com

Alibaba · platform · announced

Qwen Intelligence phone agent (Honor Magic9 first)

A full-stack phone-agent platform (Mobile Planner, Mobile-Use and Creative agents plus an OS harness) that carries out multi-app tasks, running simple tasks on a 4B on-device model and complex ones in the Qwen cloud. It is part of Alibaba's broader 'Qwen Personal Agent' push.

Expected
announced at Apsara (Sep 22, 2026); Honor Magic9 launch reported for Sep 28, 2026
Versus Muse
China's OS-level answer to personal agents, working through phone apps rather than a cloud VM, alongside ByteDance's Doubao phone agent (on sale since Sep 16, 2026 on Nubia).
Security and privacy
Reported safeguards: refusing illegal or high-risk requests, requiring user confirmation for payments, deletion and privacy access, and honoring platform limits on automated posting.

vendor-confirmed checked 2026-09-27 · pandaily.com · vocal.media · eu.36kr.com · finance.biggo.com

OpenAI · device · delayed

OpenAI device (with Jony Ive's io)

Reportedly a screenless, battery-powered, doughnut or hockey-puck-sized smart speaker with a camera, sensors and moving parts meant to feel 'alive'. It would act as a ChatGPT-powered humanlike companion for the home, handling smart-home control, messaging and media, at a reported $300-400.

Expected
OpenAI (Axios, Jan 2026): 'on track' to unveil in 2H 2026; Bloomberg (Aug 2026): unveil this year ahead of a 2027 release, timing could shift because of Apple's lawsuit
Versus Muse
OpenAI's dedicated-hardware rival to Muse Charm and Muse on glasses. Meta announced its gadget first.
Security and privacy
No privacy design published. It would have an always-on microphone and camera in the home, and Apple has filed a trade-secret suit against OpenAI.

credible reports checked 2026-09-27 · axios.com · reuters.com · 9to5mac.com · macrumors.com

Apple · device · rumored

Apple 'HomePad' home hub

A square smart display (stand or wall mount) running a new OS 'built around Siri AI', offering home control, calls and intercom, and Siri searches across the user's devices and accounts with tasks in Notes, Reminders and Calendar.

Expected
October 2026 (Bloomberg/Gurman, Sep 20, 2026)
Versus Muse
Apple's first Siri-AI-centric dedicated device, a household endpoint rather than a Muse-style open-web agent.
Security and privacy
Presumably relies on Apple's Private Cloud Compute model, but nothing is published for the device yet.

credible reports checked 2026-09-27 · 9to5mac.com · macrumors.com

Amazon · feature · rumored

Alexa 'Moonraker' agentic upgrade

An unreleased Alexa+ project, revealed in internal documents, that lets Alexa chain multi-step tasks from one request (for example 'book me a ride and text my friend'). It reportedly uses Anthropic Sonnet models and has more than $100M in projected 2026 GPU costs, and some leaders have weighed delaying or scaling it back.

Expected
no date given
Versus Muse
Amazon's route to a Muse-like doer agent, while it walls off its store: it blocked Muse on Sep 24, 2026 and pushes its own Alexa for Shopping and Buy for Me (both already live).
Security and privacy
Alexa+ beta issues included hallucinations and erratic device actions. Amazon's own Buy for Me agent identifies itself and lets brands opt out, the standard Amazon says Muse failed.

credible reports checked 2026-09-27 · businessinsider.com · thenextweb.com · businessinsider.com · aboutamazon.com

OpenAI · agent · rumored

OpenAI 'Codex Bot' persistent agent

A reported always-on assistant combining ChatGPT and Codex, with a 'Persistent Mode' that works proactively until 'put to sleep', generates follow-up tasks, keeps context across sessions and messages users unprompted.

Expected
no launch date (The Information, via Forkast, Sep 27, 2026)
Versus Muse
Reportedly OpenAI's response to Muse and xAI's GrokBot (live since Aug 2026). It may be the same effort as the leaked 'o' assistant.
Security and privacy
Nothing published. Unprompted actions across sessions raise the stakes for approval gates and prompt-injection defenses.

credible reports checked 2026-09-27 · forkast.news · cryptorank.io

Apple · device · delayed

Apple AI smart glasses (N50)

iPhone-connected glasses with cameras and visual AI for Apple Intelligence and Siri, as a platform for context about the wearer's surroundings.

Expected
late 2027 (Bloomberg/Gurman, May 31, 2026); previously end-of-2026 announcement; one report says WWDC 2027 unveil
Versus Muse
Apple's eventual competitor to Muse-powered Ray-Ban Meta glasses, arriving about a year behind Meta and Google.
Security and privacy
Nothing published. Reports say the delay was partly because Apple's visual AI was not ready.

credible reports checked 2026-09-27 · 9to5mac.com · theshortcut.com

OpenAI · agent · rumored

ChatGPT 'o' always-on assistant

Leaked UI and config (a Pro upgrade-page screenshot reading 'o, your always-on assistant' and an email suffix '-o') suggest a persistent-identity ChatGPT assistant that keeps working between sessions, possibly with its own email address.

Expected
speculated for OpenAI DevDay on Sep 29, 2026 (unconfirmed)
Versus Muse
Would be ChatGPT's most direct Muse analog, building on the live ChatGPT agent and on ChatGPT Voice's email, calendar and Slack access (launched Sep 23, 2026).
Security and privacy
OpenAI has disclosed nothing on permissions, scheduling, memory or controls. An agent-owned inbox is a prompt-injection vector.

rumor checked 2026-09-27 · kingy.ai · tokenpost.com

App builders compared: Lovable, Replit and everyone like them

All 36 prompt-to-app builders on this page: startups, the big platforms' answers, Chinese builders and open-source versions you can run on your own GPU. Most established first (maturity, then reported users and GitHub stars). Click a row for the profile.

AgentHow it worksCostAdoption (as reported)WhoAutonomyFindingsSafety design
LovableChat-driven builder that generates React front ends wired to Supabase back ends and hosts them at lovable.app.Free tier; credit-based paid plans.~8M users (late 2025, per Wikipedia).startup32Lovable shifted from 'user responsibility' to built-in scanning and abuse detection, but secure-by-default database access is still the user's job.
Bubble AI Agent2025-10Bubble's AI Agent builds and edits apps from prompts inside Bubble's visual editor; every change is visible and editable as Bubble visual logic rather than code, and apps publish to web and native…Free to build; paid plans to go live.6M users, 7M apps (Bubble, Oct 2025).startup30Most mature governance in this group (SOC 2 Type II, security dashboard, privacy rules), but privacy rules must still be set per data type by the builder.
Emergent2025-06Multi-agent platform where planning, coding, testing/verification and deployment agents build full-stack web and mobile apps from a prompt, including hosting, versioning and monitoring.Free; $20/mo Standard; $200/mo Pro; custom Business/Enterprise.6M+ users, ~150K paying (TechCrunch, Feb 2026).startup40Enterprise tier adds SSO, RBAC, audit logs and VPC deployment; no product-specific incidents found as of Sep 2026, but generated-app security is still the user's responsibility.
Base44Chat-driven builder that generates full-stack web apps on Base44's own managed backend (database, auth, integrations, hosting), so users never provision infrastructure.Free tier; paid plans $16-$160/mo (annual) metered by message and integration credits; Enterprise custom.2M+ users (Calcalist, Nov 2025); no later user count published.acquired33Wix fixed the Wiz auth bypass within a day and offers visibility and SSO controls, but the correctness of data-access rules inside each generated app remains the builder's job.
Replit Agent2024-09Browser-based agent that plans, writes, tests (in a browser) and deploys full-stack apps on Replit's own runtime, database and hosting.Free starter; $20 or $100/mo plus effort-based agent usage; enterprise custom.not publishedstartup41Replit reacted to its 2025 incident with structural fixes (environment separation, snapshots), but its product premise is an agent with deploy and data authority, so reversibility, not prevention,…
v02023-10Chat-to-app builder that generates React/Next.js UIs and full-stack apps and deploys to Vercel, using Vercel's own composite v0 models.Free tier; $30 or $100/user/mo plus token-based model costs.not publishedstartup30Relies on Vercel platform security; no product-specific agent vulnerabilities verified in this pass.
Baidu Miaoda (MeDo)2024-11Baidu Intelligent Cloud's multi-agent no-code builder: role-based agents turn Chinese or English descriptions into full-stack web apps, WeChat mini-programs and (since v3.x) native iOS/Android apps,…Free tier with credits; Pro $18/mo (MeDo).40M+ users served cumulatively (Baidu, Aug 2026, via 36Kr).big platform30No public security documentation specific to generated apps found; relies on Baidu Cloud controls.
Blink.newDescribe an app and Blink builds and hosts it with SQL database, auth (email/Google/Apple, roles), Stripe payments and custom domains; outputs web apps, native iOS/Android apps, websites, Chrome…Free; $25-$200+/mo credit plans.1M+ builders, 3M+ apps (company claim, 2026).startup30No product-specific incidents or published security controls found as of Sep 2026.
Hostinger Horizons2025-03Hostinger's prompt-to-web-app builder; multiple specialised agents generate code, design and copy, published in one click on Hostinger hosting with a native backend (Google sign-in, auth, storage,…Bundled yearly plans; prices not captured.1M+ people tried it (Hostinger, Mar 2026).big platform30Hosting-level DDoS protection, firewall and malware scanning; no scanning of generated app logic advertised.
OrchidsDesign-focused AI builder for web and mobile apps, games, CLI tools and agents across React, Next.js, Python, Swift and Flutter; can plan, debug and run commands, and lets users plug in their…Free tier; can use existing model subscriptions; paid plans not captured.1M+ users (company claim, 2026).startup31No public security programme found; the 2026 zero-click case showed slow triage of outside reports.
Rocket.new2025-06Prompt-to-app platform that generates web and mobile apps and also runs research and business-analysis tasks, combining Anthropic, OpenAI and Gemini models with in-house models trained on DhiWise's…Free; $25 / $50 / $250 per month credit plans.400K users, 10K+ paid subscribers (TechCrunch, Sep 2025).startup30No product-specific incidents or published security controls found as of Sep 2026.
DyadDesktop app (macOS/Windows) that runs a Lovable-style prompt-to-app loop on your machine: it generates React/Next.js apps in a local project folder, previews them locally, and wires Supabase or Neon…Free with your own keys or local models; Pro $20/mo, Max $79/mo for bundled credits.21.6k GitHub stars (GitHub, Sep 2026).open source · runs locally32Local-first reduces vendor data exposure, but the app executes generated code on your machine with your privileges; security depends on prompt hygiene and updates.
bolt.diyCommunity open-source version of StackBlitz's Bolt.new: an in-browser (WebContainers) or Electron app that prompts any LLM to generate, run, edit and deploy full-stack web apps, with Supabase,…Self-hosted: your own compute plus any API usage.19.9k GitHub stars (Sep 2026).open source · runs locally30No central security team; runs generated code in WebContainers in the browser, which limits host access, but deployments and keys are the operator's responsibility.
Agentforce Vibes2025-10VS Code-based vibe-coding IDE and agent ('Vibe Codey'), forked from the open-source Cline extension, that builds Salesforce apps, Lightning components and Agentforce agents against the customer's…Free tier (50 requests/org/day at launch); included on Salesforce plans.not publishedbig platform31Salesforce emphasizes org-level governance and the Einstein Trust Layer, but the IDE agent is still exposed to prompt injection from repository content.
Anything2025-08AI agent that turns a prompt into web and mobile apps with built-in backend, payments and publishing, aimed at people who want revenue-generating apps rather than prototypes.Free; $19/mo Pro; $199/mo Max; credit-metered.not publishedstartup31Offers private projects on paid plans and automated functional testing on Max; no published security scanning of generated code was found.
Atoms2026-01A team of role-playing agents (product manager, architect, engineer, data analyst, SEO, ads) builds apps with Atoms Cloud (auth, database, hosting) and then markets them; 'Race Mode' runs a prompt…Credit-based subscriptions.1M+ community members (company claim); MGX ~1.2M monthly visits (Sep 2025).startup30No product-specific incidents found as of Sep 2026; the ads and SEO agents act externally, so spending limits matter.
Bolt.new2024-10Prompt-to-app agent that runs a full Node.js environment inside the browser via StackBlitz WebContainers, then deploys.$0 / $25 / $30 per member per month; token-metered.not publishedstartup30Browser-side execution limits blast radius on the developer's machine; no product-specific incidents found as of Sep 2026.
Canva Code2025-04Generates interactive front-end websites, calculators and widgets from prompts inside the Canva editor, with drag-and-drop editing of generated code and one-click publishing.Free tier with limited AI credits; Pro/Business/Enterprise include more credits.6M websites published; 265M+ Canva MAU overall (VentureBeat, Jul 2026).startup20Scope limited to front-end content, which caps data-exposure risk; no published security design specific to Code.
Claude artifacts apps2025-06Claude generates interactive single-page apps ('artifacts') in chat; since June 2025 these can call Claude themselves and be published by link, with each viewer's usage billed to the viewer's own…Included in Claude plans; viewers' usage counts against their own plan.not publishedstartup21Artifacts run sandboxed in the browser and use the viewer's own account, but Anthropic's public sharing can host attacker-authored content.
Figma Make2025-05Prompt-to-app tool inside Figma that turns prompts, frames and design-system 'Make kits' into working React prototypes and code, with MCP connectors and publishing via Figma Sites.Part of Figma paid plans with AI credits; extra credits sold as add-ons.~60% of $100k+ ARR customers use it weekly (Figma Q1 2026); total users not disclosed.big platform30Figma runs enterprise admin controls and a trust center; it does not claim the generated app code is secure by default.
Google AI Studio BuildBrowser-based vibe-coding mode in Google AI Studio: a Gemini/Antigravity coding agent generates React/Angular/Next.js (and, since I/O 2026, Kotlin Android) apps, provisions Firestore and Firebase…Free prototyping in AI Studio; pay-as-you-go Google Cloud/Firebase and Gemini API costs for deployed apps.'Hundreds of thousands of apps' built internally before launch (Google, Mar 2026); no public user count.big platform31Google provides Secrets Manager, scoped keys and leaked-key blocking, but data-access rules and app logic of generated apps are left to the builder.
Macaly2025Prompt-to-app builder that produces production business web apps with built-in hosting, CMS, analytics, forms, database and deployment, now sold through European host team.blue.Not captured.not publishedacquired30No product-specific incidents found; relies on team.blue hosting security.
Power Apps vibe2025-10Two entry points: vibe.powerapps.com, where cooperating agents define requirements, model Dataverse tables and generate a full-stack 'code app'; and the App Builder agent in Microsoft 365 Copilot,…Included with qualifying Power Apps / Microsoft 365 Copilot licenses (prices not verified in this pass).not publishedbig platform31Strong tenant governance (Entra ID, DLP, Dataverse roles), but Microsoft treats data-exposure from maker misconfiguration as the customer's responsibility.
ReaddyGenerates a complete website (layout, pages, copy, images) from a text prompt, a template, an image or a reference URL, then lets you refine it by chat or a click-to-edit visual editor.Free tier; paid plans $25-$199/mo monthly (about 40% less billed annually); enterprise Ultimate $1,299/mo.Over half a million users (company claim in sponsored VentureBeat content, Oct 2025); not independently verified.startup30Its privacy policy cites encryption, pseudonymization and access controls, with data on AWS in Singapore and AI inputs/outputs stored in the US and Singapore; no security certifications, bug bounty…
RorkBrowser-based AI builder for iOS and Android apps; Rork Max drives Claude Code with Claude Opus 4.6 to write native Swift, pitched as a web replacement for Xcode, and adds App Store review checks,…Subscription; prices not captured.not publishedstartup30No product-specific incidents found as of Sep 2026; security of the generated apps and backend rules is left to the builder.
Wix Harmony2026-01Wix's own AI builder (separate from Base44): the Aria agent edits a site from natural-language requests while users keep full drag-and-drop control, on Wix's managed hosting with commerce,…Wix subscription plans (Harmony pricing not separately disclosed).not publishedbig platform30Relies on Wix's managed platform security and compliance features; AI edits run inside that sandboxed architecture.
a0.devAI agent that generates React Native mobile apps, wires Convex or Supabase backends, and handles App Store/Google Play builds, submissions, payments and analytics; also usable from its own mobile app.Subscription; prices not captured.200K+ users (company claim).startup30No product-specific incidents found as of Sep 2026; backend access rules are the builder's responsibility.
Open LovableOpen-source Next.js app from Firecrawl that scrapes a reference website with Firecrawl and has an LLM rebuild it (or build from a prompt) as a React app running in a cloud sandbox (Vercel or E2B).Free code; pay Firecrawl, model and sandbox usage.28.5k GitHub stars (Sep 2026).open source20Example project with no dedicated security program; misuse controls are left to the operator.
OnlookOpen-source, browser-based visual IDE for Next.js + Tailwind apps: users edit layouts Figma-style or by chat, and every change is written back to real React code; runs projects in CodeSandbox and…Free self-hosted (plus Supabase/CodeSandbox/model costs); hosted plan by quote.26.8k GitHub stars (Sep 2026).startup30Code stays in the user's repository; no published security program beyond standard open-source practice.
Chef (Convex)AI app builder from Convex, forked from bolt.diy, that generates full-stack apps with a real Convex reactive database, auth, file storage and hosting on *.convex.app.Self-hosted: your compute and model keys; hosted: Convex account.4.6k GitHub stars (Sep 2026).open source30Convex's server-side functions reduce client-side data-exposure risk compared with direct-database builders, but app logic review is still on the builder.
Ant LingGuang2025-11Multimodal AI assistant app from Ant Group whose outputs are code-driven: it generates small 'flash apps' (e.g.Free.1M+ downloads in first 4 days (SCMP, Dec 2025); 30M+ flash apps created by H1 2026 (36Kr).big platform20No published safety design specific to flash apps.
GitHub Spark (retired)2025-07Natural-language builder on github.com that generated TypeScript/React apps with built-in key-value data, LLM calls (via GitHub Models), GitHub auth and one-click hosting, backed by a real repo.No longer available to new users; was part of Copilot Pro+.not publishedbig platform30Relied on GitHub auth and GitHub-hosted runtime; product is now retired.
Google Opal2025-07Google Labs tool that turns a natural-language description into a visual chain of prompts, model calls and tools (an AI 'mini-app'), editable as a flowchart; since Feb 2026 an agent step plans and…Free.not publishedbig platform30Runs within Google's Gemini safety filters; no published controls specific to Opal sharing or data access beyond Google account permissions.
Meituan NoCode2025-05Conversational, Lovable-style builder from Meituan's R&D team: a code agent generates and deploys web apps (data analysis tools, prototypes, operations tools, sites) through multi-turn chat, with…Not verified.not publishedbig platform30No public security documentation found.
Vibecode2025-06Lets users describe an app on their phone (iOS/Android app) or web and get a working mobile or web app they can publish; multi-model backend.Subscription; prices not captured.not publishedstartup30No product-specific incidents found; Apple's 2026 pushback centred on apps that download and run code.
GPT-Engineer (legacy)2023Command-line tool where a user writes a natural-language spec and an LLM generates and iterates a whole codebase; the open-source ancestor of Lovable.Free; your API or local compute.55.1k GitHub stars (archived).open source · runs locally20Unmaintained; no safety controls beyond reviewing code before running.

Agentic browsers compared: Comet, Aside and the rest

All 33 browsers, browser extensions and computer-use agents on this page, including Aside (June 2026). This is the category with the most published prompt-injection research; read the Findings column with that in mind. Most established first. Click a row for the profile.

AgentHow it worksCostAdoption (as reported)WhoAutonomyFindingsSafety design
Claude computer use API2024-10Claude reads screenshots and returns mouse and keyboard actions through the computer-use toolset.API tokens (Sonnet 5 $2/$10 per M; Opus 5.5 $4/$20 per M) plus your own VM; Cowork/Dispatch in Pro $20/mo and Maxnot publishedbig platform43Anthropic trains for injection robustness (Opus 5.5 ties for the lowest Gray Swan injection rate it reports), classifies screenshots automatically and documents strict sandbox guidance.
OpenAI computer-use API2025-03The Responses API's computer tool returns click, type and scroll actions from screenshots, and the developer runs them in their own browser or VM.Pay per token (GPT-5.4 $2.50/$15 per M); screenshots are billed as image inputnot publishedbig platform31OpenAI trains refusals and injection resistance into the model and returns safety checks.
Browser Use2024A Python (and TypeScript) library that turns web pages into an LLM-readable element list (DOM plus optional vision) so any model can click, type and extract.Self-hosted: your compute plus LLM tokens; Cloud about $0.02/browser-hour plus LLM116.2k GitHub stars, 12.8k forks (Sep 2026)startup · runs locally31Gives developers primitives (domain allowlist, secret masking) but no built-in injection classifier; security is the integrator's job.
Stagehand2024Stagehand is an SDK (TypeScript, Python, Go) that mixes code with natural-language primitives act(), extract(), observe() and agent() over Playwright/CDP, caching actions to cut tokens.Free SDK; Browserbase metered browser time25.4k GitHub stars (Stagehand, Sep 2026)startup30Infrastructure and SDK only; injection defense depends on the chosen model and the developer's guardrails.
Skyvern2024Automates browser workflows with vision LLMs plus the DOM instead of brittle XPath selectors.Self-hosted: your compute plus LLM; Cloud: usage-based23.1k GitHub stars (Sep 2026)startup · runs locally41Developer platform; no built-in prompt-injection classifier documented.
Amazon Nova Act2025-12An AWS service and SDK for fleets of browser agents that run production UI workflows (QA, data entry, extraction, checkout).$4.75/agent-hournot publishedbig platform40Relies on AWS account governance (IAM, monitoring) and short, scoped workflows; injection robustness is not publicly quantified.
ChatGPT for Chrome2026-07OpenAI's replacement for Atlas.Included with ChatGPT plans; exact plan gating not disclosed.not publishedbig platform40OpenAI treats all web content as untrusted, asks before visiting new hosts, and applies its own confirmations and allow/blocklists before acting.
Claude in Chrome2025-08A Chrome extension that gives Claude a side panel and tools (navigate, read page, click, type, forms, JavaScript) in the user's own Chrome.Claude Pro $20/mo and upAvailable to all paid Claude plans (GA Aug 26, 2026); no user count publishedbig platform45Anthropic combines injection-robustness training, probes that flag suspicious tool results, and a per-action alignment classifier, and publishes attack-success rates.
Dia2025-06An AI-first Chromium browser (successor to Arc) with a chat that sees open tabs, 'skills' (saved prompt workflows) and connectors to Slack, Notion and Google Workspace for briefs and reports.Free; $20/mo paid plannot publishedacquired20Dia keeps the AI mostly in read-and-summarize mode and claims data is never sold or used for ad profiles.
Gemini computer-use API2025-10The Gemini API's computer-use tool returns UI actions from screenshots for a client-side loop.Gemini API token pricing; you host the environmentnot publishedbig platform30Google adds a safety service outside the model, plus developer system instructions, on top of model training, and states the tool may contain errors and security vulnerabilities.
Gemini in Chrome2026-01Gemini built into Chrome's side panel reads tabs, connects to Google apps (Gmail, Calendar, Maps, Flights, Shopping) and, in 'auto browse', drives Chrome through multi-step tasks such as forms,…Free side panel; auto browse needs Google AI Pro (about $20/mo) or UltraNot published.big platform42Google's design adds out-of-model defenses (the alignment critic, origin sets, the classifier) beyond model training.
Norton Neo2025-12A free AI-native browser from Norton that combines a chat assistant (summaries, tab organization, reminders, configurable memory, unified search) with Norton's security stack: phishing and malware…Free.not publishedbig platform20Norton puts security first (threat blocking, VPN, zero-retention AI and injection filtering).
Perplexity Comet2025-07Chromium-based browser with the Comet Assistant sidecar that reads the current page and acts in the user's logged-in tabs (click, type, navigate, email/calendar via connectors); Max users get a…Free; Comet Plus $5/mo; Pro $20/mo; Max $200/mo for background assistantNo user count published; 'millions' on the pre-launch waitlist (TechCrunch, Oct 2025)startup411Perplexity layers an HTML injection detector (BrowseSafe), guardrail prompts that mark web content untrusted, human confirmation for sensitive actions and user-visible block notices.
Yutori2025-12Yutori trains its own browser/computer-use models (Navigator n1, n1.5, n2) and runs them in hosted cloud browsers.Scouts free to start; APIs usage-based ($0.015/step browsing, $0.35/research task or scout run; token pricing for…not publishedstartup40Yutori emphasizes keeping credentials on the device and uses sandboxed cloud browsers.
TabbitAn AI-native desktop browser that combines chat with pages, an Agent mode for task automation, custom Skills and tab organization.Free tier; Pro paid (price not public on fetched page).107,763 users in 20+ countries (vendor site, Sep 2026)startup30The vendor claims SOC 2 Type I, on-device encryption and no retention.
Aside2026-06A customized Chromium desktop browser with a built-in agent that works inside the user's logged-in sites (email, dashboards, internal tools, documents), with on-device markdown memory built from…Free tier with 500 credits/mo using your own ChatGPT/Claude subscription or API key.20,000+ users within two weeks of launch (company, via EO Magazine, Jul 2026)startup40Aside relies on local-first storage, a vault that keeps secrets out of the model context, Allow/Ask/Deny tool rules, and confirmation before payments, posts and messages.
UI-TARS2025-01UI-TARS is an end-to-end vision-language GUI agent model (perception, grounding, reasoning and action in one model) trained with multi-turn RL.Free; self-hosted GPU or ByteDance remote operators37.9k GitHub stars (UI-TARS-desktop, Sep 2026)big platform · runs locally30Research-grade; no documented prompt-injection mitigations.
Nanobrowser2024An open-source Chrome/Edge extension with a Planner and a Navigator agent that runs entirely in the user's browser using the user's own API keys or local models; a free alternative to OpenAI Operator.Free; your LLM tokens or local compute13.8k GitHub stars (Sep 2026)open source · runs locally30A privacy-first design keeps data local, but there are no published prompt-injection defenses, so any page can steer the agent within the user's sessions.
BrowserOSAn AGPL-3.0 Chromium fork with a built-in agent (53+ browser tools, 40+ app integrations over MCP, scheduled tasks).Free; you pay for your model API or run models locally.~11.6k GitHub stars (Sep 2026); no user count publishedstartup · runs locally42BrowserOS focuses on privacy (local data, local models) and on isolating agent tabs from the user's own tabs.
LaVague2024A 'Large Action Model' framework: a World Model turns an objective plus the page state into the next instruction, and an Action Engine compiles it into Selenium or Playwright code (or drives a…Free; your compute and LLM6.4k GitHub stars (Sep 2026)open source · runs locally30Research and developer framework; no built-in safety layer beyond code inspectability.
Fellou2025Fellou calls itself an 'agentic browser'.Free tier; paid plans (details not public)not publishedstartup41Relies on plan review and the user stepping in.
Brave AI browsing2025-12An opt-in agent mode for Brave's Leo assistant that browses and acts on sites for the user.Free during early testing (flag-gated); Leo Premium $14.99/mo for higher-tier models.not publishedstartup30Brave's design is isolation first: a separate profile, manual start only, restricted URL classes, and a second model that reviews actions.
ChatGPT Atlas (retired)2025-10macOS Chromium browser with ChatGPT built into the new-tab page and sidebar, browser memories, and an agent mode that clicks and types inside the user's own tabs.Browser free; agent mode needed a ChatGPT Plus ($20/mo) or Pro plan; product discontinued Aug 9, 2026No user count publishedbig platform47OpenAI said prompt injection may never be fully 'solved' and relied on adversarial training, monitoring, confirmations and logged-out mode.
Firefox Smart Window2026-08An optional Firefox window type with a built-in AI assistant layered over normal browsing.Free.not publishedstartup20Mozilla keeps Smart Window opt-in and non-agentic, with local memories and zero-retention partners.
Genspark AI BrowserA browser from the Genspark 'Super Agent' company.Free browser; paid Genspark plans for agent usagenot publishedstartup31Genspark markets privacy through on-device models but publishes little about prompt-injection defenses for its cloud agent.
Opera Neon2025-09A paid agentic browser from Opera.$19.90/monot publishedbig platform31Opera relies on Chromium sandboxing, task isolation, prompt analysis, confirmations and site blacklisting, and says openly that non-deterministic AI means zero risk is impossible.
Samsung Browser AI2026-03Samsung's browser, extended from Galaxy phones to Windows PCs in March 2026, with Perplexity-powered 'agentic AI' features: page-aware Q&A, natural-language search of browsing history, multi-tab…Not stated (bundled with Samsung devices/accounts).not publishedbig platform20Samsung positions the AI features as assistive, not autonomous, and has published no agent-specific safeguards.
Sigma AI BrowserA privacy-focused browser from an Estonian startup with a built-in agent that navigates, fills forms and completes tasks end-to-end, plus 'Eclipse', a local LLM that runs inside the browser so data…Free download; paid plansnot publishedstartup30Privacy claims (local LLM, encryption) are vendor-stated; no public red-team or prompt-injection results were found.
Simular Agent S2024-10Agent S is a compositional computer-use framework: a generalist planner LLM, a separate visual grounding model (e.g.Open source + your model API costs or local GPU; Sai subscriptionnot publishedstartup · runs locally40Research framework; safety is left to the operator.
Strawberry2025-10A Chromium-based desktop browser where users build personalized AI 'companions' that run multi-step browser workflows in logged-in tabs, such as lead research, CRM and spreadsheet updates, and…Free with 2,000 credits/mo, paid plans $20-$250/mo, $10 per extra 1,000 credits.10,000 pre-registered users at Oct 2025 beta, plus 5,000 admitted after Atlas's struggles (Ny Teknik, Feb 23, 2026); no active-user countstartup30Strawberry stresses local storage of credentials and says it wants 'clearer boundaries for what agents are permitted to do'.
Cloudflare Kitesurf2026-08A stateless, Chromium-free browser built for AI agents.Free in beta.not publishedbig platform20Cloudflare isolates each session and funnels all network access through one outbound worker.
Microsoft Fara-7B2025-11A 7B-parameter agentic model that predicts browser actions straight from screenshots (no accessibility tree), small enough to run on-device on Copilot+ PCs.Free; local computenot publishedbig platform · runs locally30Trained refusals and critical-point stops, with auditable logs.
Project Mariner (retired)2024-12A DeepMind research prototype: a Chrome extension, later cloud VMs, running up to 10 parallel tasks.Discontinued (was part of AI Ultra $249.99/mo)not publishedbig platform30Research prototype with confirmation prompts; its successors carry Google's Chrome agent defenses.

Autonomy versus exposure

Each dot is a product with at least one documented security finding: how autonomously it acts (left to right) against how many findings are documented (bottom to top). Products with none are counted in the grey pill on the bottom row. The more an agent can do on its own, the more has gone wrong with it. Click a dot for the product profile.

024681012141 chat2 tools3 multi-step4 autonomous5 unattendedfindings ↑OpenClaw: autonomy 5, 13 findingsPerplexity Comet: autonomy 4, 11 findingsClaude Code: autonomy 5, 8 findingsChatGPT Atlas (retired): autonomy 4, 7 findingsCursor: autonomy 4, 7 findingsClaude in Chrome: autonomy 4, 5 findingsChatGPT Work (agent): autonomy 4, 5 findingsAmazon Kiro: autonomy 4, 5 findingsCline: autonomy 4, 5 findingsOpenCode: autonomy 4, 5 findingsMicrosoft 365 Copilot: autonomy 4, 4 findingsGitHub Copilot: autonomy 4, 4 findingsGemini CLI: autonomy 3, 4 findingsHermes Agent: autonomy 5, 4 findingsMicrosoft Copilot (consumer): autonomy 3, 3 findingsClaude computer use API: autonomy 4, 3 findingsDevin / Windsurf: autonomy 5, 3 findingsGrok (agents): autonomy 4, 3 findingsMeta Muse: autonomy 4, 3 findingsJunie: autonomy 4, 3 findingsBase44: autonomy 3, 3 findingsGemini in Chrome: autonomy 4, 2 findingsOpenHands: autonomy 4, 2 findingsChatGPT Deep Research: autonomy 3, 2 findingsOpenAI Codex: autonomy 5, 2 findingsLovable: autonomy 3, 2 findingsRoo Code (retired): autonomy 3, 2 findingsGemini Spark (agent): autonomy 4, 2 findingsMeta AI (Muse Spark): autonomy 2, 2 findingsAnythingLLM: autonomy 3, 2 findingsAgentforce: autonomy 4, 2 findingsServiceNow AI Agents: autonomy 4, 2 findingsNotion Agents: autonomy 4, 2 findingsInstinct: autonomy 4, 2 findingsTRAE: autonomy 4, 2 findingsMistral Vibe: autonomy 3, 2 findingsDyad: autonomy 3, 2 findingsBrowserOS: autonomy 4, 2 findingsgoose: autonomy 4, 1 findingsOpenAI computer-use API: autonomy 3, 1 findingsOpera Neon: autonomy 3, 1 findingsGenspark AI Browser: autonomy 3, 1 findingsFellou: autonomy 4, 1 findingsBrowser Use: autonomy 3, 1 findingsSkyvern: autonomy 4, 1 findingsGoogle Jules: autonomy 4, 1 findingsAmp: autonomy 4, 1 findingsReplit Agent: autonomy 4, 1 findingsZed: autonomy 3, 1 findingsClaude Cowork: autonomy 4, 1 findingsOpen WebUI: autonomy 3, 1 findingsGemini Enterprise: autonomy 3, 1 findingsSlackbot: autonomy 3, 1 findingsAtlassian Rovo: autonomy 3, 1 findingsSakana AI Scientist: autonomy 4, 1 findingsRabbit OS3: autonomy 4, 1 findingsTencent WorkBuddy/Claw: autonomy 3, 1 findingsAnything: autonomy 3, 1 findingsOrchids: autonomy 3, 1 findingsGoogle AI Studio Build: autonomy 3, 1 findingsPower Apps vibe: autonomy 3, 1 findingsAgentforce Vibes: autonomy 3, 1 findingsClaude artifacts apps: autonomy 2, 1 findings15 with none found69 with none found26 with none foundOpenClawPerplexity CometClaude CodeChatGPT Atlas (retired), CursorOpenCode, Amazon Kiro +3 more

Who uses them

Reported user numbers are mostly whole-platform figures (everyone who uses ChatGPT or Gemini, not only its agent mode). For open-source agents, GitHub stars are the best available popularity signal.

Reported users (log scale) — mostly whole-platform figures, not agent-mode usersGemini Spark (agent)Gemini app: 1B monthly active users (Pichai, Aug 11 2026); ~950M MAU at Q2 earnings (Forbes, Jul 22 2026). No Spark-specific count.1BChatGPT Work (agent)ChatGPT overall: 900M+ weekly active users and 50M+ consumer subscribers (OpenAI, Feb 27 2026); 'nearing 1B weekly users' per The Information (Jul 2026). No Work-specific usage published.900MDoubao330M users (by May 2026, per Wikipedia).330MAlexa for Shopping300M+ customers used Rufus in 2025 (Amazon)300MQwen app234M users (May 2026, per Wikipedia).234MMicrosoft Copilot (consumer)150M+ MAU across Microsoft's first-party Copilot family, which includes enterprise (Microsoft, Oct 2025); no consumer-only MAU disclosed.150MSuperhuman Go~40M users across the platform (TechCrunch, Sep 2026)40MBaidu Miaoda (MeDo)40M+ users served cumulatively (Baidu, Aug 2026, via 36Kr).40MKimi (agent mode)36M+ MAU (older figure, per Wikipedia).36MMicrosoft 365 Copilot30M+ paid seats (Microsoft earnings release, Jul 29, 2026)30MLovable~8M users (late 2025, per Wikipedia).8MQoder6M+ builders (vendor claim, qoder.com, 2026).6MEmergent6M+ users, ~150K paying (TechCrunch, Feb 2026).6MBubble AI Agent6M users, 7M apps (Bubble, Oct 2025).6MOpenAI Codex>5 million weekly active users (OpenAI, 2 Jun 2026).5MGitHub Copilot~4.7M paid subscribers (Microsoft, Jan 2026); '50M users' reported secondarily for Jul 2026.4.7MMeta MuseNo user count disclosed. Downloads (Sensor Tower estimates): >2.5M by Sep 21, 2026 (CNBC); 3.4M by Sep 25, 2026 (TechCrunch). DAU rose 27% the day after Connect (TechCrunch, Sep 25; no absolute figure).3.4MOpenClaw3.2M monthly active users (SimilarWeb estimate via gradually.ai, Sep 22 2026 - reported, not disclosed by the project); 13,585,228 npm downloads in 30 days to Sep 25 2026 (npm registry); 2.3M ClawHub skill installs (reported).3.2MConsensus2.5M MAU (Consensus, May 11, 2026)2.5MBase442M+ users (Calcalist, Nov 2025); no later user count published.2MTRAE1.6M+ MAU, 6M+ registered users (TRAE 2025 annual report via AIBase).1.6MGemini CLI>1 million developers (Google, Oct 2025).1MOrchids1M+ users (company claim, 2026).1MHostinger Horizons1M+ people tried it (Hostinger, Mar 2026).1MBlink.new1M+ builders, 3M+ apps (company claim, 2026).1MElicit400,000+ monthly researchers (Elicit, Feb 2025)400KRocket.new400K users, 10K+ paid subscribers (TechCrunch, Sep 2025).400Ka0.dev200K+ users (company claim).200KTabbit107,763 users in 20+ countries (vendor site, Sep 2026)108KInstinct100,000+ users, invite-only (company via press, Sep 2026)100KAside20,000+ users within two weeks of launch (company, via EO Magazine, Jul 2026)20K100K users1M users10M users100M users1B usersOpen-source agents: GitHub stars (linear scale; a popularity signal, not users)OpenClawOpenClaw: 390,000 stars★ 390KHermes AgentHermes Agent: 240,200 stars★ 240KOpenCodeOpenCode: 196,600 stars★ 197KAutoGPTAutoGPT: 186,800 stars★ 187KOpen WebUIOpen WebUI: 153,300 stars★ 153KOpenAI CodexOpenAI Codex: 126,200 stars★ 126KBrowser UseBrowser Use: 116,200 stars★ 116KGemini CLIGemini CLI: 106,600 stars★ 107KOpenHandsOpenHands: 89,300 stars★ 89KZedZed: 88,500 stars★ 88K

Security incident timeline

77 documented incidents, vulnerabilities and attack demonstrations against agent products, March 2025 to September 2026. Dot size is severity where one was assigned. Click any dot or row for what happened, what it let an attacker do, whether it is fixed, and the lesson for defenders.

AprJulOctJan 2026AprJulOct2025-04-01 · MCP Tool Poisoning (MCP clients (demonstrated on Cursor))2025-05-26 · GitHub MCP Toxic Agent Flow (GitHub MCP server + Claude Desktop)2025-05-29 · Lovable apps missing Supabase row-level security (CVE-2025-48757) (Lovable)2025-06-12 · EchoLeak (Microsoft 365 Copilot)2025-07-09 · mcp-remote OAuth command injection (mcp-remote (MCP proxy used by Claude Desktop and other clients))2025-07-18 · Replit production database deletion (SaaStr) (Replit Agent)2025-07-23 · Amazon Q wiper prompt (Amazon Q Developer extension for VS Code)2025-07-28 · Gemini CLI allowlist bypass (Tracebit) (Gemini CLI)2025-07-29 · Base44 private-app authentication bypass (Wiz) (Base44)2025-08-01 · CurXecute (Cursor IDE)2025-08-05 · MCPoison (Cursor IDE)2025-08-12 · Copilot YOLO-mode RCE (GitHub Copilot agent mode in VS Code / Visual Studio)2025-08-14 · Lovable free hosting abused for phishing at scale (Proofpoint) (Lovable)2025-08-20 · Comet indirect prompt injection (Brave) (Comet (agentic browser))2025-08-20 · Scamlexity (Guardio): Comet buys from fake store and falls for phishing (Perplexity Comet)2025-08-26 · Salesloft Drift OAuth token theft (UNC6395) (Salesloft Drift (AI chat agent) integration with Salesforce)2025-08-26 · s1ngularity (Nx) (Claude Code, Gemini CLI, Amazon Q CLI (weaponized by malware))2025-09-17 · Dyad preview-window RCE (CVE-2025-58766) (Dyad)2025-09-17 · Junie command injection CVE-2025-59458 (JetBrains Junie)2025-09-20 · ShadowLeak (ChatGPT Deep Research (with Gmail connector))2025-09-25 · ForcedLeak (Salesforce Agentforce)2025-09-29 · Malicious postmark-mcp (postmark-mcp (npm MCP server))2025-09-30 · Gemini Trifecta (Gemini (Cloud Assist, Search Personalization, Browsing tool))2025-10-04 · CometJacking (Comet (agentic browser))2025-10-09 · CamoLeak (GitHub Copilot Chat)2025-10-21 · Unseeable prompt injections in screenshots (Brave) (Perplexity Comet; Fellou)2025-10-23 · AI Sidebar Spoofing (SquareX) (Perplexity Comet; ChatGPT Atlas; Edge, Brave, Firefox, Chrome AI sidebars)2025-10-27 · ChatGPT Tainted Memories (ChatGPT Atlas (agentic browser))2025-10-29 · Escape.tech scan: 2,000+ vulnerabilities in 5,600 vibe-coded apps (Lovable, Base44, Create.xyz (Anything), Bolt.new, Vibe Studio)2025-10-31 · Opera Neon hidden-HTML prompt injection leaks account email (Brave) (Opera Neon)2025-11-04 · Agentforce Vibes prompt-to-code injection (CVE-2025-64320) (Agentforce Vibes)2025-11-13 · GTG-1002 AI-orchestrated espionage (Claude Code (misused by attacker))2025-11-19 · Comet hidden MCP API allows local command execution (SquareX) (Perplexity Comet)2025-11-25 · Antigravity data exfiltration (PromptArmor) (Google Antigravity (agentic IDE))2025-11-25 · HashJack URL-fragment prompt injection (Cato CTRL) (Perplexity Comet; Copilot in Edge; Gemini in Chrome)2025-12-03 · Codex CLI project-config command injection (OpenAI Codex CLI)2025-12-25 · Junie guidelines.md command execution (Mindgard) (JetBrains Junie)2026-01-12 · OpenCode unauthenticated local server RCE (OpenCode)2026-01-12 · OpenCode web UI XSS to command execution (OpenCode web UI)2026-01-13 · BodySnatcher (ServiceNow Now Assist AI Agents / Virtual Agent API)2026-01-14 · Cowork file exfiltration (PromptArmor) (Claude Cowork (research preview))2026-02-01 · ClawHavoc (OpenClaw (ClawHub skills registry))2026-02-02 · OpenClaw 1-click RCE (OpenClaw (gateway / Control UI))2026-02-02 · Moltbook database exposure (Moltbook (social network for OpenClaw agents))2026-02-04 · BrowserOS Chat-mode prompt injection to OS command execution (researcher report) (BrowserOS)2026-02-05 · ToxicSkills (Agent Skills ecosystem (ClawHub and skills.sh))2026-02-09 · Internet-exposed OpenClaw instances (OpenClaw (self-hosted agent gateway))2026-02-13 · Claude artifacts abused for MacSync ClickFix lures (Claude artifacts apps)2026-02-19 · Orchids zero-click flaw let researcher hijack a BBC reporter's project and laptop (Orchids)2026-02-25 · Caught in the Hook (project-file RCE) (Claude Code)2026-02-25 · Claude Code API key exfiltration before trust prompt (Claude Code)2026-02-26 · Old Google API keys silently gained Gemini access (Truffle Security) (Google AI Studio Build)2026-03-02 · Mistral Vibe workspace config code execution (Mindgard) (Mistral Vibe CLI)2026-03-04 · PerplexedBrowser / PleaseFix (Comet (agentic browser))2026-03-24 · WebPromptTrap (Cato Networks) (BrowserOS)2026-03-29 · OpenClaw token-rotation privilege escalation (OpenClaw)2026-04-21 · Lovable public-project BOLA exposed other users' source code, credentials and AI chats (Lovable)2026-04-23 · Claude Code symlink sandbox escape (Claude Code (sandbox))2026-05-07 · RedAccess: 380,000 public vibe-coded assets, ~5,000 exposing sensitive data (Lovable, Base44, Replit, Netlify)2026-05-15 · Claw Chain (OpenClaw (OpenShell sandbox))2026-06-08 · Tabstack web-agent hijack via hidden text (Brave) (Mozilla Tabstack; Cotypist)2026-06-15 · SearchLeak (Microsoft 365 Copilot (Enterprise Search))2026-06-22 · TRAE OpenPreview data exfiltration (Mindgard) (TRAE IDE)2026-06-30 · GuardFall Bash-expansion guard bypass (Adversa AI) (OpenCode, Hermes, Roo Code and other open-source coding agents)2026-07-01 · DuneSlide (Cursor IDE (agent sandbox))2026-07-21 · OpenAI eval agents breach Hugging Face (OpenAI internal cyber-evaluation agents (GPT-5.6 Sol and an unreleased model))2026-08-04 · AISI unsanctioned agent behaviour (Frontier agents under UK AISI cyber testing (Anthropic Mythos 5, OpenAI GPT-5.6-Sol))2026-08-05 · No perfect fix: layered AI-browser defenses bypassed (Brave, Black Hat 2026) (Opera Neon/AI browser; Perplexity Comet; ChatGPT Atlas)2026-08-05 · PleaseFix zero-click class extended to five agentic browsers (Zenity, Black Hat 2026) (Claude in Chrome; Gemini in Chrome; Perplexity Comet; ChatGPT Atlas; Copilot in Edge)2026-08-18 · CoSnitch (Microsoft Copilot (consumer / Copilot Personal))2026-08-28 · Hermes MCP catalog supply-chain flaw (Hermes Agent (bundled MCP catalog))2026-09-03 · Hermes git-config RCE before first prompt (Hermes Agent)2026-09-16 · BragJack: extensions hijack built-in browser AI agents (Forever Security) (Gemini in Chrome; Copilot in Edge; Opera Neon; Perplexity Comet; Claude in Chrome)2026-09-22 · Hidden dictation-endpoint setting lets local code hijack Muse and steal its session token (Meta Muse (macOS app))2026-09-22 · Human contractors quietly placing Muse phone calls in internal test raise data-exposure concerns (Meta Muse (human concierge call test))2026-09-24 · OpenCode upgrade-endpoint RCE (Datadog) (OpenCode)2026-09-25 · Bug-bounty report: flaw could expose a user's dedicated Muse VM (emails, files) (Meta Muse (Muse Secure VM))

Prompt injection 16 Data exfiltration 16 Malicious plugin / skill 5 Code execution (CVE) 19 Destructive action 1 Exposed instances 3 Supply chain 4 Auth bypass 6 Other 7

DateIncidentProductTypeSeverityStatusLesson

How capable are they? The benchmarks

The standard tests for agent capability, what each measures, the human baseline, and the current leaders. Read the caveats: many leaderboard numbers are vendor-reported or come from aggregators, and harness choice can move a score by tens of points.

GAIA

General AI assistants on 466 real-world questions that need multi-step reasoning, web browsing, file and multimodal handling, and tool use (3 difficulty levels; 300 test answers held back).

Human baseline: 92% (human respondents) vs 15% for GPT-4 with plugins in the original paper

  1. 74.55% HAL Generalist Agent + Claude Sonnet 4.5 2025-09
  2. 70.91% HAL Generalist Agent + Claude Sonnet 4.5 (high) 2025-09
  3. 68.48% HAL Generalist Agent + Claude Opus 4.1 (high) 2025-08
  4. 64.85% HAL Generalist Agent + Claude Opus 4 2025-05

Aging benchmark. Princeton HAL's standardized-harness leaderboard (validation set) had no entries newer than Sept 2025 when retrieved, and no open-weight model appears in its top entries. Validation answers are public, so contamination is likely, and scores depend heavily on the agent scaffold. Frontier labs rarely report GAIA any more.

benchmark site

OSWorld-Verified

Computer-use agents completing 369 real desktop tasks (web apps, office software, file management, multi-app workflows) in real Ubuntu/Windows VMs, scored by execution-based checkers. The Verified release (July 28, 2025) fixed more than 300 task issues and moved to AWS.

Human baseline: About 72% (original OSWorld human study)

  1. 86.1% Qwen3.8 Max (Alibaba) 2026-09-27 (leaderboard snapshot)
  2. 85.0% Claude Fable 5 (Anthropic) 2026-09-27 (leaderboard snapshot)
  3. 84.3% Qwen3.8-27B (Alibaba, open-weight) 2026-09-27 (leaderboard snapshot)
  4. 83.4% Claude Opus 4.8 (Anthropic) 2026-09-27 (leaderboard snapshot)
  5. 83.0% Gemini 3.6 Flash (Google) 2026-09-27 (leaderboard snapshot)

Top scores now exceed the ~72% human baseline, but these are mostly vendor-reported figures compiled by an aggregator, not runs verified by the OSWorld team. The maintainers warn that CAPTCHAs, website changes, timing dependencies and ambiguous tasks affect reliability. A small open-weight model (Qwen3.8-27B) is within 2 points of the leader.

benchmark site

BrowseComp

Deep-research web agents: 1,266 hard-to-find, entangled questions with short, verifiable answers that need persistent, creative browsing (OpenAI, April 10, 2025).

Human baseline: Human trainers solved about 29.2% within a 2-hour limit (gave up on 70.8%); where they did answer, their answer matched the reference 86.4% of the time

  1. 92.5% Atria Dawn Preview (Shanghai AI Lab, listed as open-weight) 2026-09-27 (leaderboard snapshot)
  2. 91.5% GPT-6 Astra (OpenAI) 2026-09-27 (leaderboard snapshot)
  3. 91.2% Kimi K3 (Moonshot AI) 2026-09-27 (leaderboard snapshot)
  4. 90.8% Claude Opus 5 (Anthropic) 2026-09-27 (leaderboard snapshot)
  5. 90.4% GPT-5.6 Sol (OpenAI) 2026-09-27 (leaderboard snapshot)

Near saturation (above 90%). Only short answers are scored, and OpenAI itself notes it is unclear how well scores track open-ended real research. Results depend on the browsing harness and context management (vendor-reported). Answers may have leaked onto the web over time.

benchmark site

SWE-bench Verified

Coding agents resolving real GitHub issues in 12 Python repositories (a 500-task human-validated subset of SWE-bench; Epoch runs 484), scored by hidden unit tests.

  1. 95.0% Claude Fable 5 (Anthropic) 2026-09 (leaderboard snapshot)
  2. 93.9% Claude Mythos Preview (Anthropic) 2026-09 (leaderboard snapshot)
  3. 88.6% Claude Opus 4.8 (Anthropic) 2026-09 (leaderboard snapshot)
  4. 85.2% Claude Sonnet 5 (Anthropic) 2026-09 (leaderboard snapshot)
  5. 80.6% DeepSeek-V4-Pro-Max (DeepSeek, open-weight) 2026-09 (leaderboard snapshot)

Effectively saturated, and contamination is a concern because the repositories and fixes are public. Most figures are vendor-reported with their own scaffolds, and the official swebench.com bash-only board is lower and lags behind. Use SWE-bench Pro or Terminal-Bench to tell frontier models apart.

benchmark site

SWE-bench Pro

Harder, contamination-resistant long-horizon software engineering tasks from Scale AI, drawn from copyleft and proprietary codebases (public and commercial held-out splits); tasks often need multi-file changes.

  1. 89.9% Claude Opus 5.5 (Anthropic; vendor-reported) 2026-09 (aggregator snapshot)
  2. 80.0% Claude Fable 5 (Anthropic; vendor-reported) 2026-09 (aggregator snapshot)
  3. 61.5% ± 3.1 Muse Spark 1.1 (Meta; Scale standardized run, commercial set) 2026 (Scale leaderboard at retrieval)
  4. 59.1% ± 3.6 GPT-5.4 xHigh (OpenAI; Scale standardized run) 2026 (Scale leaderboard at retrieval)
  5. 38.7% ± 3.6 Qwen3-Coder-480B-A35B (Alibaba, open-weight; Scale standardized run) 2026 (Scale leaderboard at retrieval)

Vendor-reported numbers on the public set (e.g., Opus 5.5 at 89.9%) are far above Scale's own standardized-harness results. Scale's board had not added the newest mid/late-2026 models when retrieved, so compare only within one source. Scores on the public and commercial splits differ.

benchmark site

Terminal-Bench 2.0

Agents doing hard, realistic work in a terminal (system administration, compiling, debugging, data and security tasks) inside containers. Each task was manually and LM-verified, and runs go through the Harbor harness.

  1. 82.7% GPT-5.5 (OpenAI) 2026-09 (leaderboard snapshot)
  2. 82.0% Claude Mythos Preview (Anthropic) 2026-09 (leaderboard snapshot)
  3. 80.4% Claude Sonnet 5 (Anthropic) 2026-09 (leaderboard snapshot)
  4. 77.3% GPT-5.3 Codex (OpenAI) 2026-09 (leaderboard snapshot)
  5. 76.2% Gemini 3.5 Flash (Google) 2026-09 (leaderboard snapshot)

Scores depend heavily on the agent harness (the official board ranks model and agent pairs), and many figures are vendor-reported. The aggregator's top 8 contains no open-weight model. A Terminal-Bench 2.1 variant is now tracked too, so check which version a claim refers to.

benchmark site

tau2-bench (τ²-bench)

Customer-service agents talking to a simulated user while calling tools under domain policies (airline, retail, telecom); telecom is dual-control, meaning the user also acts. Sierra has since added τ³ (banking knowledge retrieval, voice).

  1. 99.3% LongCat-Flash-Thinking-2601 (Meituan, open-weight) - telecom 2026-09 (leaderboard snapshot)
  2. 99.3% Claude Opus 4.6 (Anthropic) - telecom 2026-09 (leaderboard snapshot)
  3. 98.9% GPT-5.4 (OpenAI) - telecom 2026-09 (leaderboard snapshot)
  4. 96.8% MiMo-V2-Pro (Xiaomi, open-weight) - telecom 2026-09 (leaderboard snapshot)

Telecom is saturated (about 99%), and open-weight models match closed ones. Airline and retail remain harder, with the newer τ³ tracks at about 48-87% on Sierra's board. pass^1 hides reliability: pass^k (success on all k tries) is the metric that matters for production agents.

benchmark site

Humanity's Last Exam (HLE)

2,500 expert-written, closed-ended questions across 100+ subjects (41% math, 14% multimodal) from CAIS and Scale AI (Jan 2025; Nature, Jan 2026). Agent results are reported with tools (search and code).

Human baseline: No single human baseline: questions were written and vetted by nearly 1,000 subject experts to be beyond the reach of frontier models

  1. 65.0% Claude Fable 5.1 (Anthropic), with tools 2026-09-27 (leaderboard snapshot)
  2. 64.7% Claude Mythos Preview (Anthropic), with tools 2026-09-27 (leaderboard snapshot)
  3. 64.7% Claude Opus 5 (Anthropic), with tools 2026-09-27 (leaderboard snapshot)
  4. 62.5% GLM-5.3 (Zhipu AI), with tools 2026-09-27 (leaderboard snapshot)
  5. 60.0% DeepSeek-V4-Pro-0813 (DeepSeek, open-weight), with tools 2026-09-27 (leaderboard snapshot)

A FutureHouse audit suggested about 30% of chemistry and biology reference answers may be wrong, and the maintainers launched HLE-Rolling in response. With-tools and no-tools scores are not comparable: Wikipedia lists Fable 5.1 at 59.1% on the text-only subset. The official lastexam.ai page was stale when retrieved (it still showed Gemini 3 Pro at 38.3%).

benchmark site

ARC-AGI-2

Static abstract-reasoning grid puzzles testing fluid intelligence and efficient generalization to novel tasks, with pass@2 scoring and cost per task. The ARC Prize Grand Prize target is 85%.

Human baseline: Every task was solved by at least 2 humans in 2 attempts or fewer (calibrated human panel)

  1. 95.0% GPT-6 Astra (OpenAI) 2026-09 (aggregator snapshot)
  2. 85.0% GPT-5.5 (OpenAI) 2026-09 (aggregator snapshot)
  3. 77.1% Gemini 3.1 Pro (Google) 2026-09 (aggregator snapshot)
  4. 73.3% GPT-5.4 (OpenAI) 2026-09 (aggregator snapshot)
  5. 40.1% Inkling-Small (Thinking Machines Lab, open-weight) 2026-09 (aggregator snapshot)

Aggregator figures mix vendor-reported and ARC Prize-verified results, so check arcprize.org/leaderboard for semi-private verified scores and cost per task. ARC-AGI-2 is close to saturated at the frontier, and attention has moved to ARC-AGI-3. Open-weight models trail by a wide margin.

benchmark site

ARC-AGI-3

Interactive reasoning: agents must explore novel turn-based game environments, infer goals, build world models and learn within the episode. Scoring rewards action efficiency relative to humans (launched Mar 25, 2026).

Human baseline: Every environment was beaten by at least 2 independent members of the public (458-person controlled study); 100% means matching human action efficiency on every game

  1. 99.9% ($18,817 total) GPT-6 Astra (OpenAI), provider-adapter harness 2026-09-03
  2. 62.7% ($26,098 total) GPT-6 Astra (OpenAI), ARC Prize standard harness, semi-private set 2026-09-03
  3. 30.2% Claude Opus 5 (Anthropic), self-reported 2026-09 (aggregator snapshot)
  4. 7.8% GPT-5.6 Sol (OpenAI), self-reported 2026-09 (aggregator snapshot)

The harness changes results dramatically: GPT-6 Astra scored 62.7% on ARC Prize's standard harness and 99.9% with a provider adapter that used opaque provider-specific reasoning features. ARC Prize says this is not proof of AGI, and that the games test only deterministic, closed-ended mechanics. No open-weight results were listed.

benchmark site

METR 50% time horizon

The length of software and research tasks, measured in human-expert completion time, that an agent completes with 50% success. It is estimated with a logistic fit on METR's task suite (Time Horizon 1.1).

Human baseline: Defined relative to humans: task lengths come from timed runs by human experts

  1. likely at least 16 hours (50%); 80% horizon about 3h06m Claude Mythos Preview (Anthropic) 2026-05-08
  2. about 11.3 hours (95% CI 5-40 h), with heavy caveats about reward hacking GPT-5.6 Sol (OpenAI) 2026-06-26
  3. Claude Opus 5.5 (Anthropic) 2026-09-22

METR says measurements above 16 hours are unreliable with its current task suite, so the frontier is now effectively capped by the benchmark. Confidence intervals are very wide (GPT-5.6 Sol: 5-40 h). METR's Sept 2026 Claude Opus 5.5 summary gives no horizon number, saying only that it continues the trend from Mythos Preview onward. Horizons differ across task domains.

benchmark site

GDPval (and GDPval-AA)

Economically valuable knowledge work: 1,320 tasks (220 in the open gold set) across 44 occupations in the 9 largest US GDP sectors. Expert graders blind-compare model deliverables with work by professionals (average 14+ years of experience). GDPval-AA is Artificial Analysis's Elo-style agentic rerun on a 0-3000 scale.

Human baseline: Industry professionals' own deliverables; a 50% win-or-tie rate means parity with experts

  1. 1853 Elo (GDPval-AA) Claude Fable 5.1 (Anthropic) 2026-09-25
  2. 1773 Elo (GDPval-AA) GLM-5.3-Flash (Zhipu AI, listed as open-weight) 2026-09-25
  3. 1769 Elo (GDPval-AA) GLM-5.3 (Zhipu AI) 2026-09-25
  4. 1754 Elo (GDPval-AA) Muse Spark 1.3 (Meta) 2026-09-25
  5. rated as good as or better than experts on just under half of tasks Claude Opus 4.1 (Anthropic), original OpenAI GDPval gold set 2025-09

Tasks are one-shot, with no client feedback loops or iteration. Elo figures (GDPval-AA) are relative and cannot be converted directly into win rates against humans. OpenAI's original run is dated (summer 2025 models).

benchmark site

MCP tool-use benchmarks (MCP Atlas, MCPMark)

Agents using real Model Context Protocol servers (GitHub, Notion, Postgres, Playwright, filesystem, etc.) for multi-step tasks. MCP Atlas (Scale) covers many servers and tools; MCPMark stresses CRUD-heavy, verified tasks, reported as pass@1.

  1. 88.1% Muse Spark 1.1 (Meta) - MCP Atlas 2026-09 (aggregator snapshot)
  2. 84.2% Kimi K3 (Moonshot AI) - MCP Atlas 2026-09 (aggregator snapshot)
  3. 83.7% Hy4 preview (Tencent, listed as open-weight) - MCP Atlas 2026-09 (aggregator snapshot)
  4. 57.5% ± 1.1 GPT-5.2 (OpenAI) - MCPMark pass@1 2025-12-15
  5. 36.8% DeepSeek-V3.2-Reasoner (DeepSeek, open-weight) - MCPMark pass@1 2025-12-15

The MCPMark leaderboard was last updated Dec 15, 2025, so it misses 2026 models. MCP Atlas figures are mostly vendor-reported, and the aggregator's open-weight labels are inconsistent (Hy4 preview is flagged open on one page and closed on another). Scores say nothing about robustness to tool poisoning; see the security benchmarks for that.

benchmark site

METR 50%-time-horizon trend (summary)

Trend line: how fast the 50% time horizon of frontier agents grows.

Human baseline: Task lengths are calibrated against human expert completion times

Latest frontier: Claude Mythos Preview: likely at least 16 hours (METR, May 8, 2026), the ceiling of the current task suite; GPT-5.6 Sol about 11.3 h (95% CI 5-40 h, June 26, 2026) · doubling time About 130.8 days (about 4.3 months) for post-2023 models per Time Horizon 1.1 (Jan 2026); about 7 months over 2019-2025 in METR's original paper (arXiv 2503.14499)

  1. >=16 h (50%) Claude Mythos Preview (Anthropic) 2026-05-08
  2. about 11.3 h (50%) GPT-5.6 Sol (OpenAI) 2026-06-26

At a 4.3-month doubling, the suite's 16-hour ceiling is already binding, so METR needs longer tasks to keep measuring. Doubling-time estimates depend on the model window chosen, and the pace is faster since 2024 than across 2019-2025. The 4.3-month figure is reported via Wikipedia's summary of METR's TH1.1 update and was not read on METR's own post.

benchmark site

Precautions that actually help

Distilled from the incident record above. Each item maps to at least one real incident in this survey.

Assume every input is an instruction

  • Web pages, emails, calendar invites, documents, repo files, tool descriptions and search results can all carry commands the agent will follow.
  • Never give one agent private data, untrusted content and an outbound channel at the same time without a human checkpoint.

Contain the agent, not the model

  • Run self-hosted agents in a VM or container with no host mounts; OpenClaw's sandbox is off by default.
  • Restrict network egress to what the task needs; many exfiltration findings used an allowlisted domain.
  • Use a separate browser profile with no saved passwords for browser agents.

Least privilege, short-lived credentials

  • Dedicated low-privilege accounts and API keys per agent; rotate them.
  • Scope connectors to read-only unless the task needs writes.
  • Keep secrets out of files the agent can read; infostealers now target agent config files.

Keep a human on irreversible actions

  • Leave approval prompts on for payments, sends, deletes, deploys and permission changes.
  • Never use skip-permission or auto-approve modes on a machine that matters; malware abuses them.
  • Summarization can silently drop your 'ask first' instruction; enforce approvals in configuration, not in the prompt.

Treat skills, plugins and MCP servers as software supply chain

  • Install only what you have read; registries have shipped hundreds of malicious skills.
  • Pin versions; watch for tool-description changes (rug pulls).
  • Prefer vendors who sign and scan, and still verify.

Personal agents: give them their own identity

  • Connect a secondary mailbox or forward only what the agent needs; never the inbox that receives your login codes.
  • Start read-only; allow sending and buying per task, and never approve a security warning mid-task.
  • Opt out of training on your agent's activity where offered, and prune its memory.

Money: single-use cards and limits

  • Pay through single-use or merchant-locked cards (Muse uses Stripe Link cards bound to merchant, amount and time).
  • Set a spending cap and require approval for every new merchant.
  • Check the audit trail after long-running errands, before trusting the next one.

Log, watch and be able to stop it

  • Record every tool call and outbound request; review before trusting a new agent.
  • Have a kill switch: revoke tokens and stop the gateway in one step.
  • Never expose an agent's control UI or gateway to the internet.

Method and caveats

  • Data as of September 27, 2026, from product pages, documentation, changelogs, GitHub, CVE records, vendor security bulletins, security researchers' write-ups and reputable press. Every product and incident lists its sources.
  • User numbers are as reported by the vendor or a named third party, with the date; platform user counts are not agent-mode users. Open-source adoption is shown as GitHub stars or downloads.
  • Autonomy level is our reading: 1 chat only, 2 tools and retrieval, 3 multi-step tasks, 4 long-horizon work with real side effects, 5 runs unattended across systems. The self-hosting score is 0 cloud-only to 4 fully offline on your hardware.
  • Security-issue counts are documented findings we could source, not a vulnerability census. A low count can mean less scrutiny rather than better security.
  • Benchmark scores change monthly and many are vendor-reported or come from aggregators; see each benchmark's caveats.
  • Unlike the Industry map, this survey has not yet had a row-by-row second fact-check. Corrections are welcome by email.
  • The watchlist of upcoming products is labelled by confidence (vendor-confirmed, credible reports, rumor); rumors are included so you can track them, not because they are established.