← Latest briefing

Technology

Coding agents disagree on software tools in majority of tests, study finds

An analysis of thousands of automated programming sessions reveals that Claude Code, Codex, and Cursor select matching third-party services only 42% of the time.

The short version

  • Armature published findings from 5,292 validated test sessions evaluating how AI coding assistants choose third-party tools and services across various codebases.
  • The tested assistants—Claude Code, Codex, and Cursor—agreed on the same tool choice in only 42% of identical test cells, displaying differing web-search habits and preferences.
  • Prominent tech brands were frequently discussed during automated evaluations but rarely adopted, with specific vendor documentation details frequently swinging agent decisions.

Key facts

  • Armature analyzed nearly 17,000 runs across 75 test repositories, releasing results from an initial set of 5,292 valid sessions spanning 51 codebases and 18 sectors.[Hacker News]
  • The tested agents agreed on tool selections in 42% of test cells, with Claude Code opting to build custom in-house solutions 19% of the time, compared to 10% for Codex and Cursor.[Hacker News]
  • Web browsing behavior varied widely: Codex utilized web searches in 94% of sessions, Cursor in two-thirds, and Claude Code in roughly 30%, relying primarily on internal priors.[Hacker News]
  • Market dominance diverged by sector, with Stripe winning 90% of payment provider tests while PayPal and Adyen were often cited but rarely selected.[Hacker News]

What remains uncertain

  • Whether the findings reflect real-world developer workflows, given that the test suite relied on synthetic repositories and automated human-in-the-loop orchestrators driven by Gemini 3.7 Flash.[Hacker News]
  • The full results of over 10,000 additional runs that Armature excluded from its initial publication for future release.[Hacker News]

Sources