An NBER working paper circulating this week asks a question every engineering org feels but rarely measures: does writing more code still mean shipping more product?
The productivity paradox in plain English
Researchers circulated an NBER working paper this week that tackles a question every engineering org feels: if developers write—or generate—far more code, does more software actually ship? The study’s headline finding landed in tech Twitter as a sevenfold increase in lines produced alongside flat or modest gains in features merged to production in studied teams.
The distinction between writing code and shipping code is not semantic. Shipping includes review, testing, integration, operational readiness, and product judgment about what should not be built. AI coding assistants excel at the first verb; organizations still bottleneck on the second.
Methodology and what the authors measured
The paper reportedly combines telemetry from version control, issue trackers, and deployment logs across a sample of firms that adopted assistant tools between 2024 and 2026. Outcomes include pull requests merged, incidents tied to recent changes, and product analytics on feature adoption—not just LOC counters.
Authors caution against universal claims: sectors, codebases, and maturity of AI policies differ. Still, the pattern challenges vendor narratives that equate generated lines with economic value.
Why more lines can mean less clarity
Assistants propose verbose implementations, duplicate helpers, and speculative abstractions. Reviewers spend cycles trimming noise. Test suites swell with generated cases that cover happy paths but miss integration edges. The NBER analysis suggests review and rework partially absorb productivity gains—a phenomenon economists call diminishing returns to an input factor.
Teams without strong architectural guardrails may see negative second-order effects: higher maintenance load, unclear ownership of generated modules, and security review backlogs.
Implications for engineering managers
Shift metrics from lines or commits to validated outcomes: deployed features behind flags, customer-visible improvements, defect rates, and lead time from idea to production. Cap the percentage of assistant-generated code per PR if review queues blow out.
Invest in platform tooling—CI speed, test flakiness reduction, and service boundaries—so additional code has a path to ship. Otherwise assistants become typing accelerators for inventory that never leaves the warehouse.
Developer experience and morale
Surveys embedded in the study hint at mixed sentiment. Junior developers feel faster onboarding; seniors report fatigue reviewing low-quality suggestions. The shipping gap may widen if organizations reward visible code volume over maintainability.
Pair assistants with mentorship and design docs. The paper implicitly supports practices long predating LLMs: small batches, trunk-based development, and clear definitions of done.
Connection to AI safety and quality weeks
OpenAI’s misalignment grader story and enterprise insider-threat framing share a theme: automation without verification creates rework or harm. Code is not exempt. Generated patches that pass lint but fail semantics mirror models that pass rubrics while breaking environments.
Organizations should treat assistant output like untrusted contributor code—static analysis, mandatory tests, and staged rollouts.
What leaders should do Monday morning
Run an internal study mirroring NBER metrics for one squad: compare pre- and post-assistant shipping rates, not LOC. Publish results to engineers so incentives align with shipped value.
The seven× lines statistic is catchy; the shipping flatline is the lesson. Writing is cheap now. Judgment, integration, and accountability remain expensive—and decisive.
Additional context for operators
Teams reviewing this story should document which outbound integrations their agents can reach, which identities those integrations use, and whether emergency or government destinations are blocked by default. Run tabletop exercises that assume a model completes a harmful external action before anyone reads the chat transcript. Align communications, legal, and security on escalation paths when automated systems contact the public or authorities. Measure time-to-disable for agent tool access the same way you measure time-to-isolate for compromised workstations. Publish internal guidance that treats near-miss evaluations at major labs as free threat intelligence for your own connector roadmap. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable. Extend tabletop scenarios to include regulators, insurers, and union representatives where applicable.
Further Reading
Discover more articles on similar topics across our network
Shortlisting the 10 best to find your ultimate fleet management software development company
COAX Software ranks first as a fleet management software development company. Its AnchorSpan cut dwell time from 11 to 6.4 hours. Acropolium and Saritasa follow with strong logistics portfolios. See all 10 best options.
Comments
Loading comments…