A new kind of AI helper does not live on your laptop. You type a task, such as “fix the broken sign-up form” or “update these twenty files,” and the work happens on a rented computer somewhere in a data center. You can close the lid and come back to a finished draft waiting for review.
Most of these products let you run several helpers at once, like a small team you hand tickets to. This guide compares ten of them in plain terms: Grok Bot, Devin, Cursor Cloud Agents, OpenAI Codex, GitHub Copilot cloud agent, Google Jules, Claude Code cloud sessions, Kiro, Factory, and OpenHands Cloud. The short version is that they all clone your code into a cloud computer and hand back a proposed change. They differ in where you talk to them, what else they can reach, and how the bill is counted.

What a “cloud agent” actually is
Start with the everyday picture. You have a to-do list for a software project. Instead of doing each item yourself, you describe it in a sentence or two and give it to an assistant. The assistant opens its own copy of the project, makes the edits, tries them out, and sends back the result for you to approve.
The jargon maps onto that picture neatly. The project lives in a repository, or repo, which is the shared folder of code and its history. The agent works on a branch, a side copy where changes can pile up without touching the main version. When it is done, it opens a pull request, which is a proposed change someone else reviews before it goes in. The rented computer is usually a virtual machine, or VM: a software-made computer a cloud company spins up on demand.
The “cloud” part matters for three practical reasons. Work keeps going when your own machine is asleep. You can start five or fifty tasks at once without your laptop fan screaming. And the agent’s mistakes happen inside its own sandbox, a fenced-off space, instead of on the computer you use for banking and email.
The two that work most like a teammate
Grok Bot
Grok Bot is a desktop and phone app from SpaceXAI that gives you named assistants, which it calls Bots, that each have a job and a memory. According to its documentation, all of your Bots share one persistent cloud computer with a browser, a file system, and a terminal (the text window where commands run), so files and sign-ins carry over between tasks. Each Bot gets its own screen on that shared computer, and work continues when your laptop is closed. The apps run on macOS, Windows, Linux, iPhone, iPad, and Android, while the computers themselves run in Cursor’s cloud.
What sets it apart is that coding is only one lane. Bots can use connectors, which are ready-made links to services such as Slack or a ticket tracker, and they can click through ordinary websites when no connector exists. You can save a repeatable process as a skill and put it on a schedule as a routine. For code, a Bot can hand the job to Cursor Cloud Agents on separate machines, then watch the run, read its transcript, and send follow-ups. An engineering guide on the company’s site describes Bots that review the proof attached to each pull request, check failing tests and merge conflicts every half hour, and nudge the coding agent until the result meets the bar. That guide is a first-person account from one of the company’s own engineers, not an independent test.
Grok Bot launched as a beta. Access is included with paid individual Cursor plans (Pro starts at $20 a month) and with Cursor Teams, or you can link an individual SuperGrok, SuperGrok Plus, SuperGrok Heavy, or X Premium+ subscription. Included usage resets weekly, and extra usage can continue as on-demand billing through Cursor if you turn it on. The help pages describe the tiers as relative (“generous,” “highest”) rather than publishing exact task counts, so we could not confirm a number of tasks per week. One security note from the docs is worth knowing: because every Bot on an account shares one computer, separate Bots are not a security wall between each other.
A note on bias: this site uses Grok Bot for some of its own behind-the-scenes work. Everything said about it here comes from its public documentation, the same standard used for every other product.
Devin
Devin, from Cognition, was one of the first products sold as an “AI software engineer.” Its docs describe it as an agent that writes, runs, and tests code, and offer a rule of thumb: if a person could do the task in about three hours, Devin can most likely do it. Each cloud session gets its own VM with a shell, a browser, and copies of your repos, and you can watch along in an embedded editor or take over at any point. You can hand it work from its web app, by tagging it in a Slack or Microsoft Teams thread, from Linear or Jira tickets, or from a command-line tool that can pass a longer task up to the cloud.
Cognition reworked its self-serve plans recently. Its pricing page lists Free ($0), Pro ($20 a month), Max ($200 a month), Teams ($80 a month for the team plus $40 per full developer seat), and custom Enterprise pricing. Cloud agents are part of Pro and up. Paid plans come with a usage allowance that refreshes daily and weekly, and going past it is billed at the underlying model prices. Enterprise customers are billed in Agent Compute Units, or ACUs, at a rate set in each contract, and Cognition does not publish a standard ACU price. Enterprise also offers a deployment inside a company’s own private cloud.
The coding-first cloud agents
Cursor Cloud Agents
Cursor, the AI code editor, renamed its Background Agents to Cloud Agents. Each agent runs in an isolated VM set up to look like a developer’s laptop: your repo cloned, packages installed, secrets loaded, and network access. You can launch them from the desktop editor, the web, the iPhone app, Slack, Linear, a comment on a GitHub pull request or issue, or an API, and the docs say you can run as many in parallel as you want. They work with GitHub, GitLab, Bitbucket Cloud, and Azure DevOps.
The review side is strong. Agents return pull requests along with artifacts, meaning screenshots, short videos, and logs that show what changed and how the agent checked it. You can also take over the agent’s remote desktop to click through the changed app yourself. Billing is at the API price of whichever AI model you pick, drawn first from your plan’s included usage, and you must set a spending limit before the first run. A paid Cursor plan is required.
OpenAI Codex
Codex is OpenAI’s coding agent, and its cloud mode now lives inside ChatGPT. You pick a cloud environment (the saved setup with your repos, tools, and network rules), describe a task, and each task gets its own workspace that keeps working while your computer sleeps. You review changed files and test results, ask for follow-ups, and commit or open a pull request when ready. You can also mention @codex on a GitHub issue or pull request, or start work from Slack or Linear.
Cloud tasks come with ChatGPT Plus ($20 a month), Pro ($100, $200, or $500 a month), Business, and Enterprise or Edu plans. The Free and Go plans and pay-as-you-go API keys do not include cloud features. On Plus and standard Business, cloud and local work share one five-hour usage window, weekly limits may also apply, and OpenAI warns that cloud tasks may use more of the allowance than local ones.
GitHub Copilot cloud agent
GitHub’s agent, recently renamed from “coding agent” to “cloud agent,” lives right where many teams already keep their code. You can assign it an issue, start it from the agents panel or a chat on GitHub.com, or kick it off from Microsoft Teams, Slack, Azure Boards, Jira, or Linear. It researches the repo, plans, edits files, and runs tests in a short-lived cloud environment, then pushes a branch and a pull request.
The limits are spelled out unusually clearly. It works only on repos hosted on GitHub, only in the one repo you pick, on one branch, with one pull request per task, and each session stops after 59 minutes. It is available on every paid Copilot plan. Billing now uses GitHub AI Credits, where one credit equals one cent: Pro ($10 a month) includes 1,500 credits, Pro+ ($39) includes 7,000, and Max ($100) includes 20,000. Each session also uses GitHub Actions minutes, the compute time GitHub meters for automated jobs.

Google Jules
Jules is Google’s asynchronous coding agent, and its own docs still call it experimental. You connect GitHub, pick a repo and branch, and write a prompt; Jules clones the code into a VM and proposes a plan you approve before it changes anything. It then shows a diff, the line-by-line list of changes, and opens a pull request. You can also assign work by adding a “jules” label to a GitHub issue.
The free tier allows 15 tasks in a rolling 24 hours with three running at once. Jules in Pro raises that to 100 tasks and 15 at once, and Jules in Ultra to 300 and 60, through Google AI Pro and Ultra subscriptions. One catch: Google says paid Jules plans currently work only for personal Google accounts ending in @gmail.com, not company Workspace accounts.
Claude Code cloud sessions
Anthropic’s Claude Code can run a session in the cloud instead of on your machine, started from claude.ai/code in a browser, the Claude phone app, the desktop app, or by typing claude --cloud in a terminal. Sessions run on Anthropic-managed VMs, or on a company’s own servers when an organization routes them there. You can also “teleport” a finished session and its branch back into your local terminal. An auto-fix option can watch a pull request and push fixes when tests fail or reviewers leave comments.
Cloud sessions are available on Claude Pro, Max, and Team plans and on certain Enterprise seats. There is no separate charge for the VM, but cloud sessions share the same usage limits as the rest of your Claude use, so running several in parallel drains them faster. Cloning and opening pull requests require GitHub; other hosts can only upload a one-way bundle.
Kiro (from AWS)
Kiro is Amazon Web Services’ agent-centered coding tool, and AWS’s own migration guides now point users of the older Amazon Q Developer extensions toward it. Kiro’s cloud sessions run its agent in a managed sandbox, clone repos from GitHub or GitLab on the server side, and usually deliver results as a pull request. The same session can be opened from the browser, phone, editor, or terminal.
Cloud sessions are included with paid plans at no separate compute charge: Pro ($20 a month, 1,000 credits), Pro+ ($40), Pro Max ($100), and Power ($200), with add-on credits at four cents each. Limits to know: up to 10 cloud sessions at once, sessions run only in AWS’s US East (N. Virginia) region, and there is no step-by-step “supervised” mode in the cloud.
Factory
Factory sells coding agents it calls Droids, available through a desktop app, a command-line tool, a software kit, and Slack. Its Pro plan ($20 a month) includes cloud and local background agents. Plus ($100) adds Droid Computers, which are Factory-managed cloud machines (4 processor cores and 8 GB of memory, per the docs) that keep installed tools and files between sessions and pause when idle. Max ($200) offers about ten times Pro’s usage, and Teams costs $60 a month per team plus $40 per seat. Usage is governed by rolling five-hour, seven-day, and 30-day limits, and you can also register your own server as a Droid Computer.
OpenHands Cloud
OpenHands is an open-source coding agent under the permissive MIT license, so you can run it yourself for free with your own AI model key. Its hosted version, OpenHands Cloud, has a free Individual tier that adds cloud access from desktop and phone, an API, and Jira and Slack hooks, with model usage paid at cost or through your own key. Enterprise pricing is custom and includes a self-hosted option inside a company’s private cloud. OpenHands does not publish daily task caps on its pricing page.
Side by side
| Product | Maker | How you hand it work | Where it works | What comes back | Starting point for cloud use (Oct 2026) |
|---|---|---|---|---|---|
| Grok Bot | SpaceXAI (runs in Cursor’s cloud) | Chat from desktop or phone; scheduled routines | Shared persistent cloud computer, browser, connectors; hands code to Cursor Cloud Agents | Finished work in your tools; pull requests via Cursor agents | Included with paid Cursor plans (Pro from $20/mo) or a linked SuperGrok/X Premium+ account; weekly usage |
| Devin | Cognition | Web app, Slack or Teams, Linear or Jira, CLI | Its own VM with shell, browser, and editor | Pull requests | Pro $20/mo; Max $200/mo; Teams $80/mo + $40 per seat |
| Cursor Cloud Agents | Cursor | Editor, web, iPhone app, Slack, Linear, GitHub comments, API | Isolated VM with your dev setup; MCP tools | Pull requests with screenshots, videos, and logs | Any paid Cursor plan; billed at model API prices |
| OpenAI Codex (cloud) | OpenAI | ChatGPT web, desktop, or phone; CLI; @codex on GitHub; Slack; Linear | A workspace per task with network controls | Diffs, commits, or pull requests | ChatGPT Plus $20/mo and up |
| Copilot cloud agent | GitHub | Issues, agents panel, chat, Teams, Slack, Jira, Linear | Short-lived environment; one GitHub repo per task | One branch and one pull request per task; 59-minute cap | Copilot Pro $10/mo (1,500 AI credits) and up |
| Jules | Web app; “jules” label on a GitHub issue | Cloud VM with your GitHub repo | Plan for approval, diff, pull request | Free (15 tasks a day); more via Google AI Pro or Ultra | |
| Claude Code cloud sessions | Anthropic | claude.ai/code, phone, desktop, claude --cloud | Anthropic-managed or self-hosted VM | Diff and pull request; optional auto-fix | Claude Pro and up; shares your usage limits |
| Kiro cloud sessions | AWS | Web, phone, editor, CLI | Managed sandbox in AWS US East only | Usually a pull request | Pro $20/mo (1,000 credits) and up |
| Factory Droids | Factory | Desktop app, CLI, SDK, Slack | Cloud background agents; persistent Droid Computers on Plus | Session results in your repo | Pro $20/mo; Droid Computers from Plus $100/mo |
| OpenHands Cloud | OpenHands | Web, phone, API, Jira, Slack | Hosted runtime, or self-host the open-source agent | Changes through git integrations | Free Individual tier; pay model costs at cost or bring your own key |
Prices are list prices as of October 7, 2026, from each company’s own pages. “Usage-based” means the bill depends on how much AI model work each task burns, which no vendor can predict precisely in advance.
How to choose, in plain terms
- You want one assistant for code and everything around it. Grok Bot and Devin are built around a persistent helper you talk to like a coworker. Grok Bot reaches furthest into non-coding work and hands heavy coding to Cursor Cloud Agents; Devin stays focused on engineering tickets.
- Your team lives on GitHub and wants the least setup. Copilot cloud agent starts from an issue you already have. Its hard limits keep each task small, which is often a feature.
- You already pay for a chat assistant. Codex comes with ChatGPT Plus and up, Claude Code cloud sessions with Claude Pro and up, and Jules with Google AI plans. Cloud work draws from the same allowance as your everyday chats.
- You want many agents at once and proof they tested their work. Cursor’s artifacts and remote desktop, and Factory’s persistent Droid Computers, are aimed at teams running lots of parallel jobs.
- You need control over where the code goes. OpenHands can run entirely on your own machines, and Devin (Enterprise), Factory, Claude Code, and Cursor each document ways to run agents on company-managed or your own machines.
What none of them fix
Every one of these tools hands you a proposal, not a guarantee. The pull request still needs a human who reads the diff, runs the tests, and asks whether the change belongs at all. That review step is the real bottleneck once agents get fast, a theme we covered in Coding Agents Left Autocomplete Behind.
Three cautions apply across the board. First, an agent can only act within the access you give it, so start with read-only or draft-only permissions and widen them slowly. Second, usage-based billing is hard to predict; set the spending caps each product offers before you fire off fifty tasks. Third, vendor claims about speed and success rates are marketing until someone outside the company measures them, and we found no independent, apples-to-apples test of all ten.
The bottom line
Cloud agents turn “do this for me” into a ticket you can file and forget until the review lands. The coding-first tools are converging on one loop: clone, change, test, and open a pull request. Grok Bot and Devin wrap that loop in a longer-lived teammate that remembers your preferences and can manage the work for you. Pick the one that fits where your code and your team already live, then judge it on how easy its results are to check.
Further reading

AI Engineering — Chip Huyen’s practical guide to building on foundation models, including how to evaluate what an AI system produces before you trust it.

The Pragmatic Programmer — Dave Thomas and Andy Hunt’s classic on the habits behind good code, a handy yardstick for reviewing a pull request an agent wrote.