Coding helpers used to finish the line you were typing. Now many of them open your project, edit several files, run the tests, and hand you a pull request to review. That jump — from autocomplete to an agent — is the biggest everyday change in how software gets written since the IDE itself.
Through 2025 and into 2026, Anthropic’s Claude Code, Cursor’s Composer model, and OpenAI’s Codex tools made that shift concrete. They are not magic junior developers. They are better described as tireless interns with a terminal: useful when you watch the diffs, risky when you rubber-stamp them.

Autocomplete vs. an agent, in plain words
For years, “AI for coding” mostly meant next-token suggestions. You typed a few characters; the model guessed the rest of the line or a short block. You stayed in charge of every file, every test, and every git commit.
An agent is different. You give it a goal in ordinary language — “add dark mode,” “fix the flaky login test,” “open a PR for this bug.” It can search the codebase (look across many files for the right places to change), edit those files, run shell commands, and come back with a plan or a finished patch. The loop is closer to delegating a ticket than to accepting a ghostwriter’s next sentence.
That is why product teams talk about harnesses and tools. A harness is the wrapper around the model: the permissions, the file editors, the terminal, the sandbox (a limited place where commands can run without wrecking your whole machine). The model supplies judgment and text. The harness supplies hands.
Claude Code: from research preview to daily tool
On February 24, 2025, Anthropic launched Claude Code as a limited research preview alongside Claude 3.7 Sonnet. The pitch was blunt: a command-line collaborator that can search and read code, edit files, write and run tests, and work with GitHub — while keeping the human in the loop.
Anthropic said early testers finished, in a single pass, work that often took 45+ minutes by hand. That number is a vendor claim, not an independent study, but it matches how people describe the product: good at the messy middle of a ticket, not only at inventing a function from a comment.
On May 22, 2025, with the Claude 4 announcement, Claude Code became generally available. Anthropic added background tasks through GitHub Actions, native hooks into VS Code and JetBrains (so edits show up inside the editor), and a Claude Code SDK so other apps can reuse the same agent core. Claude Opus 4 was framed as a long-running coding specialist; Anthropic reported 72.5% on SWE-bench Verified, a public test set of real GitHub issues, using a simple bash-and-file-edit scaffold.
Benchmarks are slippery — different labs use different scaffolds and subsets — but the product story is clearer than any single percentage. By mid-2025, a major model lab was shipping a terminal agent as a first-class product, not a demo video.
Cursor’s Composer: speed as a feature
On October 29, 2025, Cursor (Anysphere) shipped Cursor 2.0 with Composer, its first in-house coding model. Composer is a mixture-of-experts model — a design that routes work through specialized subnetworks — trained with reinforcement learning on real software tasks inside large codebases. During training it could call the same kinds of tools developers use in the product: semantic search across a repo, file edits, and terminal commands.
Cursor’s public claim is that Composer is about four times faster than similarly capable models, with most turns finishing in under 30 seconds. The company also rebuilt the interface around agents rather than files: several agents can work in parallel on isolated git worktrees or remote machines, then a human picks the best result. Review and testing showed up as the new bottlenecks — which is why Cursor 2.0 stressed easier diff review and a native browser tool so an agent can click through what it built.

Composer’s research note is honest about limits. On Cursor’s internal “Cursor Bench,” built from real engineer requests, some larger frontier models still outscored Composer. The bet was not “we beat every general model.” It was “we made an agent fast enough that staying in flow feels normal.”
OpenAI’s Codex: local CLI and cloud tasks
OpenAI pushed the same pattern from two directions. In April 2025 it introduced Codex CLI, an open-source coding agent that runs in your terminal and can read and edit local files, run commands, and accept screenshots or rough sketches as input. TechCrunch’s coverage stressed approval controls and sandboxed “full auto” modes — network off, limited to the project directory — so autonomy does not mean a free pass on your whole laptop.
Later in 2025, OpenAI expanded Codex as a cloud software-engineering agent that can run many tasks in parallel, each in its own sandbox preloaded with a repository, then propose pull requests for review. The split matters for teams: some work wants a fast local pair programmer; other work wants overnight agents chewing through a backlog while you sleep.
What actually changes for a working developer
Three practical shifts show up again and again:
- Writing is cheaper; reviewing is the job. Agents draft patches quickly. Humans still own correctness, security, and taste — naming, architecture, and whether the change belongs at all.
- Parallel attempts beat one long monologue. Cursor’s multi-agent setup and cloud Codex both lean on “try a few times, keep the best.” That echoes what reasoning-model research found in other domains: several shorter tries often beat one endless revision.
- Tests and sandboxes are the seatbelts. Agents that can run unit tests and stay inside a limited environment fail safer than agents that only emit text into a chat window.
None of this erases the Quiet AI Job Shift story from elsewhere on this site — judgment still matters. It does change who spends time on boilerplate, migrations, and “make the tests green” chores. Junior work that was mostly typing moves up the stack toward specifying goals and catching mistakes.
What still goes wrong
Agents invent APIs that do not exist. They “fix” a bug by deleting a check. They pass the visible tests and miss a race condition. Vendor SWE-bench scores measure a narrow slice of real work; they do not measure whether your company’s billing quirks survived the refactor.
Permission design matters as much as model IQ. A tool that can push to main without review is a different product from one that opens a draft PR. Teams that treat agent output like untrusted contractor code — read the diff, run the suite, ask why — get the upside. Teams that merge on vibes inherit the downside.
The bottom line
By late 2025, coding agents were no longer a research sideshow. Claude Code moved from preview to GA with IDE hooks. Cursor shipped its own fast agent model and a multi-agent IDE. OpenAI put agents in both the local terminal and the cloud. The everyday idea is simple: software writing is turning into software directing — describe the change, watch the agent work, and spend your attention on the review.
If you are adopting one in 2026, start with a harness that shows every file touch, keeps network and disk limits on by default, and refuses to skip your tests. Speed is the fun part. Trust is the product.
As an Amazon Associate I earn from qualifying purchases. If you buy through links on this page, I may earn a commission at no extra cost to you.
Further reading

Co-Intelligence: Living and Working with AI — Ethan Mollick’s practical guide to working with AI as a collaborator — useful context for teams adopting coding agents without handing them the keys.

Software Engineering at Google — Lessons on code review, testing, and engineering judgment at scale — the human side of the workflow that coding agents still depend on.