Skip to content

Why I could not fairly rank Codex, Copilot, and Claude

Why I could not fairly rank Codex, Copilot, and Claude

By the end of April 2026, I had used OpenAI Codex CLI, GitHub Copilot agent mode, and Claude Code on real development work. I expected three broadly equivalent coding agents, separated mainly by interface or price.

That was not what I found. Codex CLI gave me useful control but asked more of me during the interaction. Copilot agent mode fitted naturally inside the editor but left too much of the context, action, and recovery path outside my control. Claude Code followed my intent more closely and needed fewer explanatory corrections.

I kept all three. The work had been too different, and my records too incomplete, to defend a ranking.

Three agents working against local repositories

The product names matter because each vendor also offered other coding-agent surfaces.

Codex CLI was the terminal agent, not OpenAI’s cloud Codex service. GitHub introduced Copilot agent mode inside VS Code before launching its separate remote coding agent. Anthropic introduced Claude Code as a terminal tool in February 2025.

All three could act against local repositories by the time I used them. Here, local describes where the clients accessed files and ran tools. It does not imply local model inference or a particular privacy boundary.

Immediately before and during my use period, Claude Code 2.1.111 added Opus 4.7 support and new effort controls, Codex CLI 0.123.0 refreshed its bundled model metadata, and VS Code 1.118 documented substantial changes to Copilot context handling and token efficiency.

I cannot recover which exact versions, models, plans, authentication paths, providers, or permission settings I used. Those release notes establish a moving historical backdrop. They do not establish my configuration.

Different work prevented a fair comparison

The three tools handled different bug fixes, feature work, tests, and maintenance. They did not receive the same task against equivalent repository states. I did not define common acceptance criteria or preserve timing, usage, retry, or rework records.

Manual review was the only consistent acceptance boundary. When I found work unacceptable, I usually explained the problem, asked the same agent to correct it, and reviewed the revision again. Each correction loop added rescue effort. Tests and other quality gates varied with the work.

That makes this an account of workflow fit, not a benchmark. I can describe what repeatedly required my attention. I cannot claim that one agent was cheaper, faster, or generally better.

Codex CLI made control visible

Codex CLI’s terminal surface made the interaction easier to inspect. I cannot tie that impression to a recovered permission mode or configuration, but the control was visible in the interaction itself.

That control came with more steering prompts. I spent more of the exchange explaining intent or redirecting the task before I would accept the work. This did not make Codex CLI worse. It made the trade-off clear: visibility and control had a cost in attention and ease.

Copilot’s convenience did not supply enough control

GitHub Copilot agent mode started from the opposite advantage. It lived inside the editor, close to the files and the immediate coding task.

The editor surface did not give me enough control over what context shaped the work, which actions the agent took, or how it recovered when the first direction was wrong. Context selection, agent actions, and recovery all contributed; I cannot isolate one as the decisive problem.

This was a workflow judgement, not a claim about every Copilot configuration. I cannot identify the editor version, model, plan, or permission mode behind those sessions.

Claude Code needed less rescue

Claude Code followed my requested intent more closely. It was less likely to pursue a direction I then had to reverse.

That changed the manual-review loop. I still inspected the work, but I less often had to explain why the chosen direction was wrong before asking for another attempt. This is a remembered workflow impression, not a measured correction rate.

I cannot turn that experience into a measured saving. The tasks differed, and I did not preserve enough evidence to separate the tool, model, task, repository, and my growing familiarity with the workflow.

The failed ranking was the useful result

I had expected price or interface to separate broadly similar agents. Instead, each tool exposed a different trade-off between ease, control, and rescue effort.

A fair ranking would have required the same tasks, equivalent starting states, common acceptance criteria, and records of retries, rework, verification, and human time. I had none of that consistently. Declaring a winner would have converted remembered workflow differences into a benchmark they could not support.

Keeping all three was the defensible decision. It preserved the option to use each where it fitted while I worked out what I actually needed from an agent runtime.

The remaining question was not which interface was most convenient. It was which runtime let me inspect and control context, actions, models, permissions, and recovery without creating more maintenance than that control was worth.

logo

I Create Reach.
I Generate Impact.
I Amplify.