Agentic SDLC · A field guide from Deop

Agentic SDLC, explained with real numbers

The software development lifecycle, redesigned for teams where some members are AI agents — taking issues, opening pull requests, passing the same review bar as humans. This page is the field guide: what changed, the honest numbers, the five maturity levels, and the mechanics high-performing teams install.

Microsoft Solutions Partner — Cloud & AI Platforms — Specialist: Agentic DevOps with Microsoft Azure and GitHub
GitHub Verified Partner

~27%

of new production code industry-wide is AI-generated

+21%

individual task output with AI assistants — org delivery stays flat unmanaged

5–10%

program-level velocity in a managed rollout’s first year — compounding after

46–90%

task-level gains on well-scoped, verifiable work — tests, docs, migrations

The teams that build with Deop

Globys
iMDsoft
Volaris Group
Vela Software Group
Intellicene
Optimal Blue
Trisura

What changed

From autocomplete to accountable teammates

Autocomplete suggests code while a person types — the person stays the author. That era ended when coding agents learned to take an assigned issue, work in their own environment, and open a pull request: GitHub’s Copilot coding agent, Agent HQ as mission control, custom agents with scoped roles, AGENTS.md files as onboarding docs. Authorship shifted — and process has to catch up.

The industry’s paradox in one line: individuals are ~21% faster, organizations mostly aren’t. The gains pool at the task level and die in the delivery system — because the bottleneck moved downstream, into work definition, review capacity, and governance. Agentic SDLC is the discipline of moving it back.

The maturity model

The five levels of agentic SDLC

L0–L1 · Where teams are

L0 Unassisted — no sanctioned AI

L1 Assisted — autocomplete

Shadow AI, zero review discipline

Individual gains, flat delivery

The plateau most rollouts hit

No baseline, no evidence

L2–L3 · The working middle

L2 Delegated — agents take scoped issues and open PRs behind one review bar

L3 Orchestrated — custom agents with defined roles; specs drive the work; gates enforce the standards

Each level scored on evidence — six dimensions per team, weakest first

L4 Team-integrated

✓ Agents in planning, not just execution

✓ Instructions files maintained like code

✓ Scorecards steer delegation share

✓ Humans own intent, verification, risk

Scored across six dimensions —

tooling & access · repo readiness · work definition · review & gates · governance · people

— and a team’s level is its weakest one

How teams climb these levels: DeopShift, Deop’s adoption framework →

The data

Four numbers that explain the whole market

Individual ≠ organizational

+21% individual task output is real. Organizational delivery on unmanaged rollouts is roughly flat. Both are true — the gap is the work.

Autocomplete accelerates typing, not delivery

Gains pool in individuals, not in flow

Org-level lift requires process redesign

This gap is the entire adoption problem

46–90% at the task level

Where agents are dramatic today: well-scoped, verifiable work with clear success criteria and fast feedback loops.

Test backfills and coverage lifts

Documentation and doc-drift repair

Dependency bumps and mechanical migrations

Small features against a written spec

5–10% program velocity, year one

The honest program-level number for a managed first year — compounding as instructions files, gates, and delegation share mature.

Grows with delegation share, not seat count

Compounds through failure-clinic discipline

Leading indicators move first: review latency, PR size

Scale decisions belong on your data, not benchmarks

~27% of new code is AI-generated

The industry-wide share of new production code — which makes review capacity and governance the binding constraint.

Review latency is the new bottleneck

One review bar for human and agent code

Provenance and audit trails in regulated work

Change failure rate is the safety gate

The mechanics

What high-maturity teams install

Delivered with

Microsoft

+

GitHub

1

Instructions files

The agent’s onboarding docs: repo-wide AGENTS.md and copilot-instructions.md, path-scoped .instructions.md — versioned, reviewed, and improved every time an agent gets something wrong.

2

Spec-driven development

Specs become the source of truth for delegation — constitution → specify → plan → tasks with GitHub Spec Kit — so agents build from written intent, not guesswork.

3

One review bar

Branch protection, required reviews, and deterministic CI gates — the same standard for human and agent PRs. Rules live in gates, not in prose.

4

Team scorecards

DORA four keys plus DX Core 4, baselined before rollout and reviewed monthly — team-level only, never individual surveillance.

Harness engineering

Performance lives in the harness, not the model

agent = model + harness

A documented team took their coding agent from 30th to 5th on Terminal Bench 2.0 by changing only the harness — the environment, constraints, and feedback loops around the model. The principle: “follow our standards” in a prompt is probabilistic; a linter that blocks the PR is deterministic.

Tool orchestration

scoped toolsets per agent — a test-writer that literally cannot deploy

Verification loops

fast tier-1 CI · meaningful tests · rules moved from prose into gates

Context & memory

instructions files · specs · ADRs — productive from the repo alone

↻ Plus guardrails & observability — five layers, and the weakest one gets fixed first

Why delivery stays flat

The bottleneck moved downstream

Work definition

Agents amplify clarity and punish ambiguity: a vague issue produces a confident, wrong PR. High-maturity teams write agent-ready issues — acceptance criteria, scoped context, explicit boundaries.

Review capacity

More code, produced faster, lands on the same reviewers. Without smaller PRs, faster tier-1 CI, and review working agreements, the throughput gain queues up and dies in review.

Governance

Security asks “who wrote this and why” — and needs an answer. Provenance, scoped permissions, enterprise AI controls, and freeze rules turn agent output into deployable code.

Governance & measurement

Measured like a team, governed like production

4 keys

DORA: lead time, deploy frequency, change failure rate, MTTR

Core 4

DX: speed, effectiveness, quality, impact — baselined before rollout

0

individual-level metrics — team scorecards only

“Expansion freezes if change failure rate rises. Delegation expands only on evidence. One review bar for human and agent code — no exceptions, no shadow lanes.”

Plain answers

The questions everyone searches

What is agentic SDLC?

A software development lifecycle where AI agents hold real team responsibilities — taking issues, opening pull requests, passing the same review bar as humans — inside governed workflows. The shift is from AI as a typing assistant to AI as an accountable teammate.

What’s the difference between Copilot autocomplete and a coding agent?

Autocomplete suggests code while a person types — the person stays the author. A coding agent takes an assigned issue, works in its own environment, and opens a pull request for review. Authorship shifts, and team process has to catch up.

What is an AGENTS.md file?

A repo-level instructions file that teaches agents how your codebase works — build commands, conventions, boundaries. Treat it like onboarding documentation: versioned, reviewed, and improved every time an agent gets something wrong.

Are AI agents replacing developers?

No — the evidence says roles shift rather than disappear. Work moves toward specification, review, and system design. Teams that thrive treat agents as capacity for well-defined work, and developers as the ones who define, verify, and own it.

How do you measure AI coding agents?

At the team level: DORA four keys plus flow metrics like review latency and agent-authored share, baselined before rollout so change is attributable. Never individual surveillance — it corrodes the trust adoption depends on.

What is spec-driven development?

Writing an executable specification before code — with GitHub Spec Kit: constitution, specify, plan, tasks — so humans and agents build from the same source of intent. Specs turn delegation from a gamble into an engineering practice.

Where does your team sit on the five levels?

A two-week, fixed-fee assessment places every team on the model, names the weakest dimension, and hands you a 90-day roadmap — yours to keep either way.

Book the 2-week assessment →