Fieldstate research · visual edition

The State of AI-Assisted and Agentic Development

2026

AI assistance is mainstream. Agent use is becoming common. High-authority agentic delivery is not. Thirty-nine pages on what that distinction costs you, and how to tell where your delivery system actually stands.

The central thesis

The question is no longer whether you can generate more code.

It is whether your delivery system can convert faster generation into safe, maintainable, valuable change, without overwhelming verification, weakening human capability, or increasing operational risk.

A system with broad permissions and weak controls is not mature. It is unsupervised.

For the developer
A junior who can answer the one question is becoming senior.
For the leader
A CTO who makes room for them to learn it is a leader worth following.

A practical test

Five questions to ask of any AI-assisted change.

Not a maturity scale. A test you can run today, on the next change in front of you.

  1. Can we understand it?

    Not "does it run" — can a human hold its shape in their head?

  2. Can we test it?

    Assurance has to scale with generation, or it becomes a queue.

  3. Can we explain it?

    To a reviewer, to an auditor, to the person on call at 3am.

  4. Can we measure it?

    Flow, reliability, value, capability. Not lines, not accepted suggestions.

  5. Can we safely reverse it?

    Reversibility is the cheapest safeguard you will ever buy.

If the answer is "no," increasing delegated authority is hard to justify. Chapter 1 · Ledger L1–L5

Chapter 3 · the task-shape heuristic

What makes delegation easier or harder to verify?

Six dimensions that shape the verification burden of a delegated task.

AI-favourable

Lower expected verification burden

AI-hostile

Higher expected verification burden

Scope
Bounded, well-specified
Ambiguous, open-ended
Codebase
Greenfield or well-tested
Mature, tacit-knowledge-heavy
Domain
Familiar patterns, common stacks
Novel domains, unusual constraints
Determinism
Mechanical transformation, boilerplate, tests
Judgement-heavy trade-offs
Risk
Low blast radius, easily reversed
Auth, payments, schema, shared infrastructure
Ownership
Individually owned
Cross-team, contested

Heuristic, not a predictive model.

Use it to shape authority, reviews, and safeguards, then validate it against your own delivery outcomes.

See report: ch. 3 · Ledger: L6–L11

Chapter 7 · the pairing

Maturity belongs to the organisation. Authority belongs to a workflow.

One organisation sits at one maturity level. Every workflow inside it sits at its own authority band. That is why these are two axes and not one number.

Both marked cells below belong to the same M4 organisation. High maturity does not mean maximum authority everywhere.

Not a scoring surface. No target cell, no diagonal to travel. Unmarked cells carry no judgement.

Figure 16 · Authority × maturity
Maturity · organisationAuthority · workflow
  • Authentication workflow — restricted to A2
  • Documentation workflow — permitted A5

The one-page version

The whole argument, printable.

Download the PNG
One-page summary: code got cheaper, mistakes didn't — the five-question test and the six task-shape dimensions on a single sheet.
Pin it above the review queue · 1536 × 1024

Executive summary

What this report supports saying. Six claims, each traceable.

Every number traces to an entry in the evidence ledger. If a number’s permitted framing does not fit the chart, the number does not go in the chart.

01 · AdoptionL4 · L5a–b

Using an agent and surrendering authority to one are different decisions.

Stack Overflow 2025 reported coding-agent use by ~31% of item respondents. A late-April 2026 pulse (n=1,100, different instrument) reported 59% at some frequency, while 63% rarely or never permitted fully autonomous operation.

02 · ProductivityL9 · L10

There is no universal productivity number.

METR’s RCT — 16 experienced OSS developers, 246 tasks, familiar mature repositories — found participants took 19% longer with AI while believing they had been sped up. A narrow historical counterexample, not a current estimate.

03 · Displaced workL11–L13

More work can move into the system while verification pressure rises.

In Faros telemetry (~22,000 developers), high-adoption periods showed more completed work per developer and, in the same comparison, median review time ~441% higher. Correlational; not customer value.

04 · MeasurementChapter 4

More code is not more value.

Generated lines, accepted suggestions, and PR counts are activity metrics. Evaluation belongs on a ladder: activity, flow, reliability, value, capability. No activity metric may quietly imply a value outcome.

05 · SecurityL14 · L15a–c

The evidence justifies stronger controls, not panic.

Package hallucination and security-benchmark results demonstrate durable failure modes. None of these are production prevalence rates, and no credible industry-wide prevalence measurement was identified.

06 · PeopleL18a–c

The junior pathway must evolve, not disappear.

Broad verification capability still depends on implementation fluency. The junior role should be redesigned around supervised delivery: not removed, and not converted into cheap AI-review labour.

What's inside

Eight chapters, one evidentiary spine.

  1. Ch 1DefinitionsWhat we are actually talking about, and the A1–A5 authority ladder used throughout.
  2. Ch 2Trust and the gapThree separate measures on three separate dimensions. Never one number.
  3. Ch 3Task shapeSix dimensions that shape the verification burden of a delegated task.
  4. Ch 4The delivery systemThe measurement ladder: activity, flow, reliability, value, capability.
  5. Ch 5Evidence classesFive classes of security evidence, and what each one can and cannot support.
  6. Ch 6The tensionLabour, hiring, and why the apprenticeship pipeline is the real constraint.
  7. Ch 7MaturityTen dimensions, M0–M4, and the authority × maturity pairing.
  8. Ch 8MarketWhere the tooling money is going, and what that does and does not tell you.
  9. A–EAppendicesThe full evidence ledger — thirty-seven entries indexed, twelve given a full card — plus the survey protocol.

Method

What we counted, and what we refused to count.

Inclusion

Primary surveys, peer-reviewed and preprint research, platform telemetry, vendor engineering blogs, named-analyst reports, regulatory filings, documented incidents.

Exclusion

Aggregator and SEO content. We do not infer net industry benefit or harm by counting positive against negative studies.

Positioning

Fieldstate is vendor-independent, not neutral. What we commit to is transparent method, graded evidence, and falsifiable claims.

The governing rule

If a number’s permitted framing does not fit the chart, the number does not go in the chart.

Free · no form · no follow-up sequence

Generating code is getting cheaper. Shipping trustworthy change is not automatic.

Thirty-nine pages, thirty-seven ledger entries, and a framework you can run against your own delivery system this quarter.

Want it walked through with your team? Get in touch.

Fieldstate

Clarity before code.

© 2026 Fieldstate · Aotearoa New ZealandSystems architecture & product engineering