Evidence, not opinion

What AI coding agents do well, and what they don't

Across product, architecture, engineering, quality and delivery, here is where coding agents stand, judged from published research and engineers' own words.

Every placement links to its sources, all published in the twelve months before . None comes from an AI or developer-tool vendor. Quotes are word for word; “…” marks words left out. Where evidence points the other way, it is shown too.

Product

Does well

Prototypes and throwaway demos

Turning an idea into something people can click and react to is where the evidence is strongest.

  • cloud development environments accelerate prototyping, enabling non-technical users to generate high-fidelity "throw-away" prototypes valuable for experiential exploration

    Kobiella, Breidenstein & SchmidtCHI 2026 · From Throw-Away to Takeaway: How GenAI and Vibe Coding Accelerate Prototyping Across Technical Skill Levels · · Source
  • Evidence is strongest for prototyping and user-interface work and weakest for production, data-intensive, and safety-critical use

    Siddeeq, Waseem, Kemell, Saari, Rasku & AbrahamssonarXiv · Vibe Coding in Software Development: A Multivocal Literature Review · · Source
  • Personal projects, prototypes, internal tools, one-offs etc. I don’t think anybody disputes that this technology is great for those kinds of things.

    Jason GormanFounder, CodemanshipPersonal blog · The Gorman Paradox: An Explanation? · · Source
  • AI-assisted programming significantly reduces the cost of building the wrong thing.

    Simon Willisonsimonwillison.net · Don’t “Trust the Process” · · Source
Mostly

Breaking a spec into tasks

Plans and task lists are useful and traceable, but often oversized for the job.

  • You might be surprised at the quality of the planning, architecture and task breakdown that a simple prompt with some context hints can give you.

    HN user noodletheworldHacker News · · Source
  • A list of tasks that trace back to the requirement numbers

    Birgitta BöckelerDistinguished Engineer, Thoughtworksmartinfowler.com · Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl · · Source

Pointing the other way

  • the workflow was like using a sledgehammer to crack a nut. The requirements document turned this small bug into 4 “user stories” with a total of 16 acceptance criteria

    Birgitta BöckelerDistinguished Engineer, Thoughtworksmartinfowler.com · Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl · · Source
Mixed

Taking a prototype to production

The speed-up fades in the last stretch: isolation, infrastructure, reliability and maintenance still need expert hands.

  • deployment and long-term maintainability remain dependent on technical expertise, with non-technical users consistently encountering barriers when transitioning beyond prototyping

    Kobiella, Breidenstein & SchmidtCHI 2026 · From Throw-Away to Takeaway: How GenAI and Vibe Coding Accelerate Prototyping Across Technical Skill Levels · · Source
  • Across both systems, vibe coding accelerated scaffolding and integration. However, the generated code often under-specified isolation rules and infrastructure constraints when these were not explicitly defined.

    Shuvo, Islam, Hasan, Waseem & AbrahamssonarXiv · Context Before Code: An Experience Report on Vibe Coding in Practice · · Source
  • Toy prototypes proportionally contains a much higher amount of the type of rote greenfield scaffolding that agents are good at writing.

    HN user bccdeeHacker News · · Source
  • When user experience, reliability, security and maintainability matter, we’re forced to drink from the firehose one small mouthful at a time

    Jason GormanFounder, CodemanshipPersonal blog · The Gorman Paradox: An Explanation? · · Source
Mixed

Asking clarifying questions

By default agents assume rather than ask; when they are set up to ask, results improve a lot.

  • current agents are largely optimized for autonomous execution

    Edwards & SchusterarXiv · Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents · · Source
  • while models like GPT-5-Coder excel at coding, they often lack the strategic communication skills required for efficient partnership.

    Li, Wu & ChangarXiv · ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions · · Source
  • better coding models do not always correspond to better dialogue models

    King & FlaniganarXiv · Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents · · Source
  • the clarifying questions were generally not useful because it was pointing out issues it obviously knew the answer to

    HN user samdjstephensHacker News · · Source

Pointing the other way

  • achieves a 69.40% task resolve rate… closing the performance gap with agents operating on fully specified instructions.

    Edwards & SchusterWhen a separate agent is set up to ask the questions.arXiv · Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents · · Source
  • I use Claude Code and it often asks clarifying questions when it is unsure how to implement something

    HN user js8Hacker News · · Source
Mixed

Building the UI it describes

Generated interfaces look structured and polished, but a quarter of the stated design reasoning never makes it into the UI.

  • On average, over 25% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34% for functional requirements

    Imteyaz et al.arXiv · Design Theater: Evaluating the Gap Between User-Facing Design Reasoning and Implementation in Generative UI Tools · · Source
  • effective at creating structured layouts, they face challenges in meeting accessibility standards and providing interactive functionality.

    Sawicki et al.arXiv · Qualitative Evaluation of LLM-Designed GUI · · Source
Weak

Knowing what users need

Models can't stand in for real users, and bigger models don't fix that.

  • would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people.

    Chen, Zhu & ZhengLanguage models used as stand-in survey respondents, compared with real people.arXiv · When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses · · Source
  • Neither failure is remedied by a larger, more capable model.

    Chen, Zhu & ZhengarXiv · When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses · · Source
  • could partially tailor interfaces for different user personas but lacked deeper contextual understanding.

    Sawicki et al.arXiv · Qualitative Evaluation of LLM-Designed GUI · · Source
  • AI atm is not good without proper steering and most importantly AI doesn't decide what to build.

    HN user 6thbitHacker News · · Source
Weak

Pushing back on a bad idea

Agents tend to go along with what they are told, and agent loops make it worse; asked directly for a critique, they can give one.

  • the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior

    Thantham JitthamarXiv · Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models · · Source
  • Decision Flip Rates reaching up to 72% and False Alignment Rates exceeding 90%, indicating substantial instability and agreement with misleading prompts

    Fahad, Asif & TawhidWhen a prompt suggested the wrong answer about code smells.arXiv · Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts · · Source
  • Present even fundamentally flawed suggestions, and you’ll get back “sounds great, boss, I’ll get right on that.”

    PhotoStructurePhotoStructure blog · The LLM sycophancy antidote · · Source

Pointing the other way

  • If you ask it directly it will provide a seemingly objective opinion on your decisions and direction.

    HN user HekkovaHacker News · · Source
  • Claude code will happily tell me my ideas are stupid

    HN user mikkupikkuWhen asked to weigh several alternatives side by side.Hacker News · · Source
Weak

Accessible UI without being asked

Generated interfaces average about two accessibility violations each; asked to fix issues, agents improve most pages but fully resolve few.

  • Across 300 original interfaces, our judges identified 541 semantic accessibility issues before any fault injection, averaging nearly two violations per UI

    Calò, Gurita & De RussisCHI 2026 Extended Abstracts · · Source
  • frequently violate Web Content Accessibility Guidelines (WCAG)

    Yoon, Cho, Kim, Chung, Jeon, Lim & YuarXiv · A11yn: Aligning LLMs for Web Accessibility-Aware UI Generation · · Source

Pointing the other way

  • improve accessibility compliance in 80.2 percent of instances

    Oyelayo, Abushaqra, Asadi, Dey & CostaWhen asked to repair known issues. The same study finds fewer than 26 percent fully resolved.arXiv · LLM Based Web Accessibility Repair: An Empirical Study of Detection, Remediation, and Cost · · Source
Back to the scale

Architecture

Mixed

Explaining existing code

A fast first tour of an unfamiliar codebase, but measured accuracy on real-repository questions is modest, and wrong answers sound confident.

  • It can explore the code base and find and explain patterns and available methods far faster than I can.

    HN user anyonecancodeHacker News · · Source
  • AI-assisted coding helps me a lot to churn through boilerplate, straightforward implementations, test writing, exploration of unfamiliar codebases

    HN user piva00Hacker News · · Source

Pointing the other way

  • Semantic search answered 65.2% of questions correctly against 46.2% for deep agentic search … the single largest share of its failures, 41.8%, occurred at the hand-off between the planner and its sub-agent, and these were usually silent, ending in a fluent and confident answer that was wrong.

    Rafiei Oskooei et al.arXiv · Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study · · Source
  • overall accuracy remains limited for repository-scale comprehension.

    Alebachew, Leary, Vaishampayan & BrownarXiv · Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering · · Source
Mixed

Drafting designs and design docs

Useful for spotting gaps early, but drafts tend to be long, stuck at code level and missing the reasoning behind decisions.

  • The most frequently mentioned benefits are More Effective Search Engine (25 out of 65 participants, 38.46%) and Early Detection (22 out of 65 participants, 33.85%)

    Wang et al.A survey of developers using a chat assistant for software design.arXiv · Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey · · Source

Pointing the other way

  • the generated ADDs are often too long, implementation-focused, and miss the rationale behind the decision.

    Karan, Dhar, Soliman & VaidhyanathanADDs are architecture design decisions.arXiv · Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study · · Source
  • they consistently exhibit granularity mismatches, operating at the code level rather than architectural abstractions.

    Sathvika, Dhar & VaidhyanathanarXiv · LLM-based Automated Architecture View Generation: Where Are We Now? · · Source
  • It does not replace tacit knowledge/domain knowledge required to design a good solution

    HN user piva00Hacker News · · Source
Mixed

Weighing design trade-offs

Agents know architecture theory well, but are least reliable when options must be weighed against each other.

  • Context length is not uniformly beneficial. It helps where the task is recall and hurts where the task is trade-off reasoning.

    Santilli, Daghero & Tourchi MoghaddamarXiv · SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models · · Source
  • Architectural Solutions is consistently the weakest category

    Santilli, Daghero & Tourchi MoghaddamarXiv · SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models · · Source

Pointing the other way

  • codebase design and modularity seems like a concern where computational sensors alone cannot help us much, AI is needed to add semantic interpretation, and consider trade-offs.

    Birgitta BöckelerDistinguished Engineer, Thoughtworksmartinfowler.com · Maintainability sensors for coding agents · · Source
Mixed

Choosing dependencies

Agents rarely add dependencies, but pick known-vulnerable versions more often than people do.

  • rarely add new dependencies (1.3% of PRs)

    Twist & ZhangarXiv · A Study of Library Usage in Agent-Authored Pull Requests · · Source

Pointing the other way

  • select PR-time known-vulnerable versions more frequently (2.46% vs. 1.64%)

    Singla, Çakar, Amusuo & DavisarXiv · Towards a Benchmark for Dependency Decision-Making · · Source
Weak

Preventing design drift

Left alone, agent-assisted code grows more complex with each change; explicit rules help.

  • static analysis warnings increase significantly by 30.3%, and code complexity increases by 41.6%

    He, Miller, Agarwal, Kästner & VasilescuAfter open-source projects adopted Cursor, compared with similar projects that did not.arXiv · Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects · · Source
  • without human review and coupling expertise, AND without these extra AI reviews, the agent was definitely compounding inadvertent technical debt.

    Birgitta BöckelerDistinguished Engineer, Thoughtworksmartinfowler.com · Maintainability sensors for coding agents · · Source

Pointing the other way

  • clean that up, and then continue enforce these layers going forward

    Birgitta BöckelerDistinguished Engineer, ThoughtworksOnce dependency rules between layers were written down and checked.martinfowler.com · Maintainability sensors for coding agents · · Source
Weak

Avoiding over-engineering

Agent code is more verbose and layered than the problem needs.

  • agent code is 2.3x more verbose and 2.0x more eroded

    Orlanski et al.Compared with 473 open-source repositories.arXiv · SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks · · Source
  • It has learned complexity. Then, inside an agent loop, it applies that complexity locally.

    Maxim Saplindev.to · Debloating The AI-Grown Codebase · · Source
  • Often times it's not even that the code is bad but rather that it's overengineered

    HN user _fat_santaHacker News · · Source
Rarely

Keeping a design coherent over many changes

Over long runs, structure erodes in most cases, and telling the agent to write clean code doesn’t slow the decline.

  • Quality degrades across checkpoints, with structural erosion rising in 77% of trajectories and verbosity in 75.5%.

    Orlanski et al.arXiv · SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks · · Source
  • Explicit quality guidance reduces initial verbosity and erosion by up to a third, without affecting degradation rates.

    Orlanski et al.arXiv · SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks · · Source
  • Each addition is defensible in the moment. The bloat comes from accumulation: every agent turn leaves behind a little local compromise, a little explanatory residue, a little defensive abstraction.

    Maxim Saplindev.to · Debloating The AI-Grown Codebase · · Source
  • You need to manually push models to clean up the slop every now and then otherwise it becomes chaotic.

    HN user redox99Hacker News · · Source
Back to the scale

Engineering

Does well

Boilerplate, scaffolding and docs

The repetitive setup and documentation work developers least want to do is where agents are most trusted.

  • writing boilerplate code that we wish we didn’t have to write in the first place

    Huang, Reyna, Lerner, Xia & HempelA surveyed developer. Developers rated boilerplate suitable for agents 25 times and unsuitable 0 times.arXiv · Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 · · Source
  • documentation tasks achieve 82.1% acceptance compared to 66.1% for new features

    Pinna, Gong, Williams & SarroarXiv · Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance · · Source
Mostly

Porting code between languages

Ports work well when a test suite checks the result, but complicated parts tend to get simplified away.

  • Porting entire open source libraries from one language to another via a coding agent works extremely well.

    Simon Willisonsimonwillison.net · I ported JustHTML from Python to JavaScript with Codex CLI and GPT-5.2 in 4.5 hours · · Source
  • We have a complete implementation of Pokemon battle system that produces the same results as the existing JavaScript codebase

    Christopher ChedeauPersonal blog · Porting 100k lines from TypeScript to Rust using Claude Code in a month · · Source

Pointing the other way

  • It ported all the functions very loosely where anything that was remotely complicated would not be ported but instead "simplified".

    Christopher ChedeauPersonal blog · Porting 100k lines from TypeScript to Rust using Claude Code in a month · · Source
  • No matter how much I tried to force it to stick to a mostly line-by-line port, it kept trying to "improve" the code

    HN user thijserHacker News · · Source
Mostly

Implementing well-specified features

With explicit, checkable acceptance criteria, agents usually deliver working features.

  • LLMs are useful for producing code that meets easily and objectively verifiable acceptance criteria which you provide explicitly.

    Jacob O'BryantPersonal blog · 2x, not 10x: coding with LLMs in 2026 · · Source
  • concrete, well-defined, and implementation-focused

    Huang, Reyna, Lerner, Xia & HempelThe tasks one surveyed developer prefers agents for. Developers rated following well-defined plans suitable for agents 28 times and unsuitable twice.arXiv · Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 · · Source
Mixed

Multi-file refactors

Detailed, local refactors go well; large structural ones fail more often or quietly change behaviour.

  • agents perform well at implementing refactorings that are specified in detail, but often fail to discover the human refactoring choices when given a focus area for changes.

    Thillen, Mündler, Raychev & VechevarXiv · CodeTaste: Can LLMs Generate Human-Level Code Refactorings? · · Source
  • agents perform fewer high-level refactorings (43.0%) than humans (54.9%)

    Horikawa, Li, Kashiwa, Adams, Iida & HassanarXiv · Agentic Refactoring: An Empirical Study of AI Coding Agents · · Source
  • It has worked fairly well but it tends to introduce really subtle changes in behaviour (almost always negative) which are very difficult to identify

    HN user mirsadmHacker News · · Source
Mixed

Debugging

Given a failing test or stack trace, agents often find the bug; finding the root cause on their own is much less reliable.

  • it was highly effective at fixing a bug when I could point it to a specific failing test

    Huang, Reyna, Lerner, Xia & HempelA surveyed developer.arXiv · Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 · · Source
  • I just gave Claude a stacktrace and a description and let it go ham. It produced an amazingly accurate thought train and found my issue

    HN user mavamaartenHacker News · · Source

Pointing the other way

  • even frontier agents fail to recover the developer-identified root-cause region in most cases, achieving only 40% overlap with ground-truth buggy lines

    Tu, Gaur, Murtinty, Wang, Shi, Song & WangICML 2026 · FaultLoc: Evaluating Coding Agents For Fault Localization · updated regularly · Source
Mixed

Following codebase conventions

Written-down, consistent conventions get followed; unwritten ones give way to generic defaults.

  • instructions in the context files are well followed by coding agents

    Gloaguen, Mündler-Sasahara, Müller, Raychev & VechevarXiv · Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? · · Source
  • it does a good job following existing conventions in a codebase, as long as they're really consistent

    HN user SwellJoeHacker News · · Source
  • Each line in that file is based on a bad agent behavior, and it almost completely resolved them all.

    Mitchell HashimotoCo-founder of HashiCorpOn the instructions file he keeps for agents.mitchellh.com · My AI Adoption Journey · · Source

Pointing the other way

  • Without this direction LLMs tend to default to stuffing too much code into the models themselves

    HN user JohnBootyHacker News · · Source
Weak

Large and legacy codebases

Fitting changes into big, old systems the agent can’t hold in view at once is a weak spot unless scope is kept tight.

  • poorly on adding new features to existing functionality

    Huang, Reyna, Lerner, Xia & HempelA surveyed developer. Developers rated integrating with existing and legacy code suitable for agents 3 times and unsuitable 17 times.arXiv · Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 · · Source
  • One of the reasons the AI fails constantly is that it has no context of the entire code base. … So it actively adds bloat to the system unless it’s guided by a skilled developer who already knows the system.

    HN user thinkingtoiletHacker News · · Source

Pointing the other way

  • CC can also do well with brownfield development, but the scope _has_ to be constrained

    HN user mwigdahlHacker News · · Source
Rarely

Reusing code instead of writing more

Agents tend to write new code rather than reuse what already exists, and don’t stop to suggest restructuring.

  • leave duplicated logic in 50.8% of task chains by turn 5---all while pass rates barely move.

    Ma et al.arXiv · Do Coding Agents Reuse Existing Code or Reinvent the Wheel? · · Source
  • The agent has no reflex that says "this has become unmanageable, we should stop and refactor." It just keeps adding branches to the pile.

    Rodrigo Rosenfeld RosasPersonal blog · AI Agents and the Refactoring That Never Happens · · Source
  • LLMs seem to have a hard time getting the big picture and reusing code that is already implemented and ALMOST does what you want vs. rewriting everything from scratch

    HN user altern8Hacker News · · Source
Back to the scale

Quality and testing

Does well

Writing unit tests

Agents write many tests readily, with coverage comparable to human-written ones, though they lean on mocks more.

  • AI-generated tests contribute to code coverage comparable to human-written tests

    Yoshimoto et al.arXiv · Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage · · Source
  • HAPRs exhibit a larger extent of testing, nearly doubling the test-to-source line ratio

    Milanese et al.HAPRs are pull requests written by people working with an agent.arXiv · Human-Agent versus Human Pull Requests: A Testing-Focused Characterization and Comparison · · Source
  • My hobby projects have 100x more tests than they used to, because LLMs are great at writing tests.

    HN user brookstHacker News · · Source

Pointing the other way

  • 36% of commits made by coding agents add mocks to tests, compared with 26% by non-agents

    Hora & RobbesarXiv · Are Coding Agents Generating Over-Mocked Tests? An Empirical Study · · Source
  • you're probably churning out repetitive tests, unnecessary tests, tests which aren't great at catching bugs.

    HN user pydryHacker News · · Source
Mixed

Getting failing tests and CI green

Agents can patiently chase failures and sometimes find the real cause, but success rates vary widely.

  • It's not perfect. I have to throw away some of the bad solutions, but shaved 20 minutes off their pipeline and improved pass rate by 35% in a handful of weeks.

    HN user SeanAndersonHacker News · · Source
  • It found the root cause and then asked whether the failure was either a regression or an insufficiently specified test

    HN user DaishimanHacker News · · Source

Pointing the other way

  • the best-performing LLM achieves an 18.9% repair success rate

    Muna, Rafi & ChenA simple repair pipeline on smaller models, not a full agent.arXiv · CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows · · Source
Mixed

Reviewing code

AI reviewers catch some real problems, but most of their suggestions are not adopted, often because they are wrong.

  • GPT 5.4 xhigh thinking was really good at teasing out problems in multi step flows of a process I was refactoring

    HN user rafaelmnHacker News · · Source

Pointing the other way

  • code suggestions made by AI agents are adopted into the codebase at a significantly lower rate than suggestions proposed by human reviewers (16.6% vs. 56.5%). Over half of unadopted suggestions from AI agents are either incorrect or addressed through alternative fixes by developers.

    Zhong, Noei, Zou & AdamsAcross 278,790 review conversations in 300 open-source projects.arXiv · Human-AI Synergy in Agentic Code Review · · Source
  • About 1 in every 10~20 comments is actually useful or something novel that hasn't been caught elsewhere

    HN user emeralddHacker News · · Source
Weak

Writing secure code

Code that works often still has security flaws.

  • Although 57% of the solutions from SWE-Agent with Claude 4 Sonnet are functionally correct, only 11.8% are secure.

    Zhao et al.arXiv · Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks · · Source
Weak

Producing merge-ready changes

About half of agent fixes that pass the tests would still be turned down by the project’s maintainers.

  • roughly half of test-passing SWE-bench Verified PRs written by mid-2024 to mid/late-2025 agents would not be merged into main by repo maintainers

    Whitfill, Wu, Becker & RushThe agents had no chance to revise their work after feedback, as a person would.METR · Many SWE-bench-Passing PRs Would Not Be Merged into Main · · Source
Rarely

Keeping tests honest when stuck

When a task can’t be done, agents often change or game the tests instead, and stronger models do it more.

  • GPT-5, cheats 54.0% of the time on Conflicting-SWEbench when facing these clearly impossible tasks

    Zhong, Raghunathan & CarliniarXiv · ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases · · Source
  • stronger models generally exhibit higher cheating rates

    Zhong, Raghunathan & CarliniarXiv · ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases · · Source
  • It also modified many of the tests to make them pass in mischievous ways.

    HN user debugnikHacker News · · Source
  • CC will claim victory even if test is red, or CC just deletes the failing tests

    HN user kreijstalHacker News · · Source

Pointing the other way

  • a quality overview of the test approaches, strengths, weaknesses, and gaps

    HN user jonathaneuniceWhen asked directly to review a test suite that checked internal state rather than outcomes.Hacker News · · Source
Back to the scale

Delivery and operations

Does well

CI, build and automation

Agent changes to CI, builds and docs are merged more often than any other kind.

  • tasks related to documentation, CI, and build update achieve the highest merge success, whereas performance and bug-fix tasks perform the worst.

    Ehsani, Pathak, Rawal, Al Mujahid, Imran & ChatterjeeAcross 33,000 pull requests written by agents.arXiv · Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub · · Source
  • Agents are good at using gh (GitHub CLI)

    Mitchell HashimotoCo-founder of HashiCorpmitchellh.com · My AI Adoption Journey · · Source

Pointing the other way

  • I got fed up of Claude Code creating GitHub Actions workflows for me that used stale actions

    Simon Willisonsimonwillison.net · simonw/actions-latest · · Source
Mixed

Changes that hold up after merge

Open-source studies disagree: some find agent changes need fixing or reverting more often, others find no penalty. None yet measures production incidents.

  • merged agent PRs attract verified fixes at 1.62 times the odds of merged human PRs in the same repositories over the same period of time

    Takerngsaksiri, Duong & BarnettarXiv · Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests · · Source
  • Agent-authored code shows modestly elevated corrective rates (26.3% vs. 23.0%)

    Rahman & ShihabarXiv · Will It Survive? Deciphering the Fate of AI-Generated Code in Open Source · · Source

Pointing the other way

  • Our analysis does not detect a velocity-quality tradeoff in these coarse proxies.

    Khosravani & MockusarXiv · Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories · · Source
  • Codex-authored PRs were reverted about half as often as human PRs (6.1% vs. 11.5%, odds ratio 0.50), while Devin PRs were reverted more often (14.5%, odds ratio 1.31).

    Obada KraishanarXiv · Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild · · Source
Mixed

Keeping changes small enough to review

In open-source data, agent pull requests are smaller than people’s; engineers at companies report the opposite, and reviews struggle to keep up.

  • I saw a 29,000 line pull request across seventy files recently.

    HN user tombertHacker News · · Source
  • the bottleneck has moved, from writing the code to reviewing it. … the disparity can be jarring when you have multiple thousands of lines of code generated every day and people are used to a review cycle based on tens or hundreds.

    HN user regularfryHacker News · · Source
  • If the AI was good enough to have coded it, please instruct it to make the changes in reviewable chunks.

    HN user greiskulHacker News · · Source

Pointing the other way

  • Agentic PRs tend to introduce smaller and more localized changes than Human PRs

    Ogenrwot & BusingeAcross 24,014 agent and 5,081 human pull requests merged in open-source projects.arXiv · How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests · · Source
Mixed

Migrations and large-scale changes

Repetitive migrations work when each change can be checked automatically; inconsistent systems stall them.

  • we successfully rolled out 240 automated migration PRs

    Devon Edwards JosephSenior Engineer, SpotifySpotify Engineering · Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4) · · Source
  • Today complex migrations are much simpler to perform

    Salvatore SanfilippoCreator of RedisHacker News · · Source

Pointing the other way

  • This lack of standardisation across our data landscape made it hard to write all-in-one prompts for Honk

    Devon Edwards JosephSenior Engineer, SpotifySpotify Engineering · Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4) · · Source
  • I'd let Claude (opus 4.7 max effort) crank away overnight only to immediately find that had added some horrible new bug

    HN user sarchertechHacker News · · Source
Weak

Long tasks without supervision

The best agent succeeds half the time on tasks taking an expert about 17 hours, but reliably only on tasks of about 3 hours.

  • An 8-hour time horizon does not mean that AIs can do 8 hours of work that a (high-context) human professional can do as part of their day-to-day job.

    METRMETR’s data for the best model measured, an early Claude Mythos Preview: 50% success on tasks of about 17 hours (very uncertain), 80% success on tasks of about 3 hours.METR · Task-Completion Time Horizons of Frontier AI Models · updated regularly · Source
  • In our original time horizon paper, we found that AI agents did worse on messier tasks.

    METRMETR · Task-Completion Time Horizons of Frontier AI Models · updated regularly · Source
Weak

Finding why production broke

Working out the cause of a live incident from logs, metrics and traces is still unreliable.

  • Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard

    Gong, Choi, Agarwal, Schechner, Huang, Agrawal, Agarwal & DwivediarXiv · ORCA-bench: How Ready Are Language Model Agents for Oncall? · · Source
Weak

Staying within what was asked

Given vague instructions, agents guess and act, sometimes beyond what was asked, rather than stop.

  • underspecification does not mainly make agents fail; it makes them guess. 55.8-67.8% of runs violate at least one boundary.

    Ji, Zhang, Xu, Li, Gao, Wang & CheungarXiv · Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions · · Source
  • Across 10,000 benign runs, 19.51% trigger overeager behavior

    Qu, Liu, Deng, Zhang, Li, Zhang & ZhangarXiv · SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents · · Source
  • undesirable changes (excluding tests and documentation) in 35 to 65% of cases.

    Gloaguen, Mündler, Müller, Raychev & VechevWhen given issues that had already been fixed.arXiv · Coding Agents Don't Know When to Act · · Source
Rarely

Taking responsibility for what ships

Someone has to stand behind shipped code, and in every source we found, that someone is a person.

  • A computer can never be held accountable. That's your job as the human in the loop.

    Simon Willisonsimonwillison.net · Your job is to deliver code you have proven to work · · Source
  • For high quality system programming tasks you have to still be fully involved

    Salvatore SanfilippoCreator of Redisantirez.com · Redis array type: short story of a long development · · Source
  • Our rule is that if your name is on the PR, you own the code

    HN user regularfryHacker News · · Source
Back to the scale

Find out whether yours is engineered, or slop.

Xpoze measures your code's engineering quality against published research. In your environment, with your LLM.