
The AI-native SDLC is a playbook from Anthropic's Applied AI team that rebuilds the whole software development lifecycle around coding agents, not only the build step. It turns the six-stage line (plan, design, build, test, deploy, maintain) into a loop. Each stage commits one artifact to Git — intent.md, spec.md, plan.md, the diff and its tests, the reviewed PR, the incident record — and the next stage starts by reading it. Governance moves out of committees and into hooks, skills, and evals that run while the agent works.
That's the answer. Here's the part I care more about.
This morning I wrote the first approval gate the playbook recommends: a Claude Code hook that blocks edits to database migrations unless the branch carries a change ticket. It took about fifteen minutes. Then I broke it on purpose, and it waved the edit straight through. No error on screen. No warning. The gate just quietly stopped being a gate.
I'll show you exactly how that happens, and a second trap in CI that does the same thing at scale. Both are one line of config away from the "governance enforced as the AI acts" story the playbook tells. Before that, though, you need the frame — because the playbook's best idea isn't the loop diagram everyone is screenshotting.
What is the AI-native SDLC, and where does it come from?
The AI-native SDLC is Anthropic's framework for redesigning planning, review, testing, deployment, and maintenance so they keep pace with agent-written code. Anthropic published it on August 21, 2026 as "The AI-native SDLC playbook" (lead author Louis Claxton, with Jim Blackhurst, Will Steuk, and Jamal Arif), and it's also a free 14-lesson course on Claude Academy.
The fourteen lessons are an introduction, twelve plays, and closing thoughts. The twelve plays are:
- Capture as intent.md
- Requirements and design
- Plan mode as the default starting point
- The CLAUDE.md
- Skills as institutional knowledge
- Parallel sessions and subagents
- Give Claude a feedback loop
- Continuous evals in CI
- AI in the PR review loop
- Hooks as approval gates
- CI/CD integration and deployment
- Closing the loop on metrics
To be clear about credit: the framework is Anthropic's. Nothing in the structure below is my invention. What I'm adding is the practitioner layer — the actual files, the configs I ran against Claude Code 2.1.286 on October 1, 2026, the places where I think the playbook is right, and the places where I'd push back.
If you've read my earlier piece on the Agentic Development Life Cycle, note the difference. ADLC is about building agents as products. The AI-native SDLC is about using agents to build any software. Different problem, different playbook.
So why does the old lifecycle need replacing at all? The playbook's diagnosis is sharper than most "AI changes everything" takes.
Why the traditional SDLC's controls assume a human at every step

The classic SDLC isn't heavy by accident. Every stage is a discrete phase owned by a different role. PMs write requirements. Architects design. Engineers build. QA — at regulated companies, a whole department — verifies. Release teams ship, and ops watches production. Work crosses each boundary as a document, a ticket, or a sign-off.
That weight bought accountability. And it made sense when writing code was the most expensive stage by a wide margin. PRDs, estimation rituals, and security review boards all existed to force alignment before weeks or quarters of expensive engineering time got spent.
Here's the assumption buried in all of it: a human performs every step. Every control was sized for human output.
Once agents write most of the diff, the playbook argues three things become true:
- The bottleneck moves. Build gets fast. Plan, review, test, and deploy still run at human speed, so they become the constraint.
- The controls stop matching reality. Line-by-line human review can't keep up when an agent produces most of the code.
- Governance gets more expensive, not less. Exceptions still route through weekly or monthly committees, and there are now far more changes generating exceptions.
The security example in the guide is the one that landed for me. Security teams are staffed for the volume of code humans write. Multiply that volume and you get one of two outcomes: the review queue backs up, or code ships under-reviewed. A regulated organization can accept neither. So the security and policy checks themselves have to run at agent speed.
The diagram above recreates the playbook's framing. Before agents, Build is the long block. After agents, Build shrinks to a sliver, and the width it gives back is "cycle time reclaimed" — but only if the surrounding stages change shape too. Plan and design compress into requirements. Test becomes review. Deploy becomes release.
That compression is what the loop is for.
The line becomes a loop, and the commit chain becomes the audit trail
The playbook's headline shift is visual. The traditional lifecycle is a line, and a single slow loop back means a new release cycle. The AI-native version is a loop that turns in hours, with Claude in the middle and humans "above the loop" — instigating, directing, and governing rather than performing each step.
Pretty diagram. But my honest read is that the loop is the packaging, and the artifact chain is the product.
Every stage ends by writing one file (or one set of files) to version control, and the next stage begins by reading it:
| Stage | Traditional SDLC | AI-native SDLC | Committed artifact |
|---|---|---|---|
| Plan | Requirements by committee, workshops, sign-offs, written by hand | Claude synthesizes pain points straight from the sources | intent.md |
| Design | Analyst writes a spec, designers parse it | Requirements and design compressed into one agent session, guided by skills | spec.md |
| Build | Handwritten code and tests, docs written after | Agent-generated code and tests; knowledge lives in CLAUDE.md and skills | plan.md, then the diff and its tests |
| Test | QA gates at stage boundaries | Continuous evals woven through implementation | Eval results in CI |
| Deploy | Humans review every line; governance varies by reviewer | Layered agentic review; humans review regulated or critical code; hooks gate actions | The PR with review findings |
| Maintain | Humans watch production for bugs | Agents monitor; a breached control band becomes a new intent.md |
The incident record |
Why do I think this is the real idea? Because git log becomes your audit trail without anyone writing an audit report. Who asked for what lives in intent.md and its author. What the agent produced lives in the diff. Who approved it lives in the PR review and the merge. An auditor can reconstruct the decision chain from commits alone.
The early artifacts are Markdown on purpose: a product owner and an agent can both read them. From Build onward, the artifact is code plus its records.
Pushback number one, right here: the playbook makes this sound tidy. In practice, the artifact chain only works if people stop making decisions in Slack DMs. The moment the real scope change happens in a thread and never lands in spec.md, your audit trail has a hole shaped exactly like the decision an auditor will ask about. The tooling can't fix that. Team habit has to.
With that caveat parked, here's how the plays look as actual files.
Plan and design plays: intent.md, requirements, and Claude Design
Capturing intent as intent.md
Intent enters three ways: someone has an idea, someone files a ticket, or an alert fires. The playbook's move is to skip the backlog-grooming relay — user stories, story points, refinement, ownership changing hands at each step — and have the originator brainstorm with Claude until a Markdown proto-spec exists. What's wanted, why, and under which constraints, in the originator's own words.
The product owner reviews and corrects that agent-written file before it's committed. That's the first human gate.
The playbook leaves the template to you ("an agreed intent.md template"). Here's the one I'd start with. Short is the point — if this turns into a 4-page PRD, you've rebuilt the old process in a new file extension.
# intent: Faster CSV export for large order histories
- **Status:** draft <!-- draft | accepted | rejected -->
- **Originator:** Support lead (ticket SUP-2291)
- **Source:** ticket <!-- idea | ticket | alert -->
- **Accepted by:** — <!-- PO name + date when accepted -->
## What we want
Customers with 50k+ orders can export their order history without the
request timing out.
## Why now
Three enterprise accounts raised it this month; support is exporting by hand.
## Constraints
- No change to the CSV column format (customers have import scripts).
- Must respect existing tenant isolation.
- No new paid infrastructure without approval.
## Out of scope
- PDF export. New filters.
## How we'll know it worked
- Export for the largest current tenant completes without a timeout.
- No support tickets about export timeouts for 30 days.
## Open questions
- Is an emailed download link acceptable, or must it stay synchronous?
Where does it live? The guide's simplest answer: an intent/ folder in the product repo. In a monorepo it's just a directory. A separate intent repo only pays off when one piece of intent spans many repositories. A platform team sets it up once and decides who can write to it.
Non-engineers don't need Git skills to take part. The playbook's route is claude.ai or Cowork with a GitHub connector, so the PO commits through the same chat they brainstormed in. I like that a lot. It's the first time I've seen a credible answer to "how do we get product people into version control without a training program."
Requirements and design in one session
An accepted intent.md triggers the next pass. Claude produces a combined requirements-and-design spec, guided by organizational skills for brand, security, compliance, and UX. The PO reviews spec.md but doesn't write it. The output is something engineering can plan against, with concerns flagged rather than buried.
For front-end work, the guide's example routes through Claude Design (beta). The PO mocks the screen up from intent.md, iterates, and exports it to Claude Code — the Export menu's "Handoff to Claude Code" option — to build. Claude Design is an Anthropic Labs product released April 17, 2026, and it's still a beta on Pro, Max, Team, and Enterprise plans. My own write-up on the Claude Design to Claude Code website workflow covers the handoff mechanics in more depth.
My opinion: this is the stage where the playbook's speed gain is largest and least proven. Collapsing analyst, designer, and spec-writer into one session is fast. Whether the resulting spec catches what a skeptical architect would have caught depends almost entirely on how good your skills are — which is a Build-stage play. So the dependency graph (we'll get to it) is honest about this: requirements and design needs both intent capture and skills in place first.
The spec is approved. Now an engineer opens a terminal.
Build plays: plan mode, CLAUDE.md, skills, and parallel sessions
Plan mode as the default starting point
Traditionally, the plan lived in the engineer's head, and the first reviewable artifact was the diff. That's too late. By the time you're reviewing a diff, the design decisions are sunk cost.
The playbook says: start every task in plan mode. Hand Claude the approved spec.md, let it interview you, and iterate until the plan is right. Plan mode reads the codebase without changing anything. The approved plan gets committed as plan.md, so later stages — the reviewer, the verifier — can check the work against what was agreed.
Two ways in, both verified against the current docs:
# Start a session already in plan mode
claude --permission-mode plan
# Or, mid-session: press Shift+Tab until the status bar shows "⏸ plan mode on"
Then a prompt like: "Read intent/sup-2291/spec.md. Interview me about anything ambiguous before proposing a plan. When we agree, write the plan to intent/sup-2291/plan.md." I've covered heavier planning setups in my Ultra Plan test, but plain plan mode plus a committed plan.md covers most of the value.
The CLAUDE.md as onboarding for the agent
CLAUDE.md gives Claude the context a new hire needs: conventions, commands, architecture, and common mistakes. It loads at the start of every session, the whole team maintains it, and — this is the bit teams skip — you update it every time Claude makes a mistake worth not repeating.
For the AI-native SDLC specifically, I'd add a section that teaches the agent the artifact chain itself:
## Commands
- Test: `php artisan test --parallel`
- Lint: `./vendor/bin/pint --test`
- Static analysis: `./vendor/bin/phpstan analyse --memory-limit=1G`
## Artifact chain (read before you change anything)
- Work items live in `intent/<id>/`: intent.md -> spec.md -> plan.md.
- Before editing code, read the plan.md for the current work item.
If none exists, stop and ask for one. Do not invent scope.
- If you must deviate from plan.md, write the deviation and the reason
under "## Deviations" in plan.md before continuing.
## Protected paths
- `database/migrations/` and `infra/` need a change ticket (CHG-####)
in the branch name. A hook enforces this; don't try to route around it.
## Known mistakes
- Never use `DB::raw()` with interpolated input. Use bindings.
- Queue jobs must be idempotent; exports can retry.
The "Deviations" rule is mine, not the playbook's. It closes the gap I complained about earlier: when scope changes mid-build, the change lands in the artifact chain instead of a chat thread. For more on keeping CLAUDE.md lean, see my CLAUDE.md and skills install guide.
Skills as institutional knowledge
The playbook's rule of thumb for skills is the clearest one I've seen: write a skill for institutional knowledge that must be applied the same way everywhere. Don't write one for things that belong in CLAUDE.md or a prompt.
My translation:
- CLAUDE.md: facts about this repo. Commands, layout, gotchas.
- Skill: policy that applies across many repos and changes centrally. Your security review checklist. Your accessibility standard. Your logging and PII rules.
- Prompt: this task, this time.
A skill is a folder with a SKILL.md at .claude/skills/<name>/SKILL.md, frontmatter with at least a description, and the instructions below it. When the security team updates the PII rule, they change one skill in one place, and every session that loads it picks up the new policy. That's the "updated centrally when policy changes" property the playbook wants, and it's the property a wiki page never had. I go deeper on the format in the Claude Code agent skills guide.
Parallel sessions versus subagents
These two get conflated constantly, and the playbook separates them cleanly:
- A parallel session is another full Claude Code instance working a separate task in its own Git worktree. Sessions share nothing except you.
- A subagent runs inside one session as a scoped helper with its own context window and tool limits.
Parallel sessions raise how many tasks are in flight. Subagents keep each session focused. Your job becomes steering and reviewing.
Starting an isolated parallel session is one flag now:
claude --worktree sup-2291-export
# second terminal, second task
claude --worktree chg-4821-orders-index
I wrote up the worktree mechanics and the cleanup gotchas in git worktrees for parallel Claude Code agents. The subagent side matters more for the next stage, because the playbook's best testing idea is a subagent.
Test plays: feedback loops, verifier subagents, and continuous evals

Give Claude a feedback loop
The single highest-payoff habit in the whole playbook: always give Claude a way to check its own work — tests, a build, a screenshot diff — so it fixes its mistakes before you ever see them. Without that, you're the feedback loop, and you're the slowest component in the system.
Concretely, that means the test, lint, and analysis commands sit in CLAUDE.md (they're in the snippet above) and your prompts end with "run the tests and fix failures before reporting back."
The verifier subagent is a different thing
The playbook draws a distinction I hadn't articulated before reading it. The feedback loop runs throughout the task. A verifier subagent runs once, at the end, in a fresh context window — so its verdict isn't colored by the assumptions that produced the code.
That second point is subtle and correct. The session that wrote the code believes its own plan. A fresh context reading plan.md and the diff cold doesn't.
Here's a verifier definition, checked against the current subagent docs (file: .claude/agents/verifier.md):
---
name: verifier
description: Independent final check before a task is reported done. Use after implementation claims completion. Runs build and tests, compares the diff to plan.md, returns PASS or FAIL with evidence.
tools: Read, Grep, Glob, Bash
model: sonnet
maxTurns: 15
---
You are a verifier. You did not write this code and you do not trust
the summary you were given.
1. Find the work item's plan.md under intent/. Read it fully.
2. Run `git diff main...HEAD --stat`, then read every changed file.
3. Run the test, lint, and static analysis commands from CLAUDE.md.
4. Check each step in plan.md: implemented, partially implemented,
or missing. Check "## Deviations" for anything undocumented.
5. Flag any change outside the scope of plan.md.
Respond with exactly:
VERDICT: PASS or FAIL
EVIDENCE: command output excerpts and file:line references
SCOPE DRIFT: list, or "none"
Never edit files. If something is broken, report it.
Note the tools line: no Edit, no Write. A verifier that can "helpfully" fix what it finds has stopped being independent. Invoke it explicitly with @agent-verifier, or let Claude delegate based on the description.
Continuous evals in CI
This is the AI-native replacement for stage-gate QA, and it's aimed at a different thing than unit tests. An eval suite runs whenever the agent's configuration changes — a new model, a rewritten prompt, an edited skill — and answers one question: does the agent still work to the same standard?
The playbook's warning here is the one most teams will ignore. Treat the suite as alive. As models improve, old cases stop discriminating — everything passes, and a suite where everything passes tells you nothing. Add new cases from what monitoring catches. Some teams run evals offline on a schedule instead of on every change, which is fine as long as someone owns the cadence.
If you've made it this far, you have the plan and build halves of the loop. The next section is where the "governance enforced as the agent acts" promise either holds or quietly doesn't.
How do Claude Code hooks work as approval gates?
A Claude Code hook is a command that runs at a fixed point in the agent's lifecycle — for approval gates, PreToolUse, which fires before every tool call — and can allow, deny, or "ask" for a human's approval. Because hooks run wherever Claude acts, the rule holds in a terminal session, a worktree, or CI.
The playbook frames two modes. During Build, hooks are guardrails: allow or block, no human involved. The "ask" decision pauses until someone approves, which is what release gating needs. Its two examples: block edits to migrations and infrastructure without a change ticket, and stop the agent from editing test files during a fix task (so it can't make the test pass by changing the test).
Here's the migration gate as I built it. The script reads the hook's JSON from stdin, checks the file path, and looks for a CHG-#### ticket in the branch name. No ticket: deny. Ticket present: "ask," so a named human still approves the edit.
File: .claude/hooks/change-gate.sh
#!/usr/bin/env bash
# PreToolUse gate: protected paths need a change ticket in the branch name.
set -euo pipefail
input=$(cat)
file=$(jq -r '.tool_input.file_path // empty' <<<"$input")
[[ -z "$file" ]] && exit 0
# Only gate migrations and infrastructure code.
if [[ ! "$file" =~ /(database/migrations|infra|terraform)/ ]]; then
exit 0
fi
branch=$(git -C "${CLAUDE_PROJECT_DIR:-.}" rev-parse --abbrev-ref HEAD 2>/dev/null || echo "")
ticket=$(grep -oE 'CHG-[0-9]+' <<<"$branch" || true) # <- this "|| true" matters
if [[ -z "$ticket" ]]; then
decision="deny"
reason="Protected path ($file). Create a branch named with a change ticket, e.g. feat/CHG-1234-add-index, then retry."
else
decision="ask"
reason="Protected path under $ticket. A human must approve this edit."
fi
jq -n --arg d "$decision" --arg r "$reason" '{
hookSpecificOutput: {
hookEventName: "PreToolUse",
permissionDecision: $d,
permissionDecisionReason: $r
}
}'
Registered in .claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/change-gate.sh",
"timeout": 10
}
]
}
]
}
}
I tested the script on October 1, 2026 (Claude Code 2.1.286, jq 1.7.1, git 2.50.1) by piping it the same JSON shape a PreToolUse event sends — a scratch repo, an Edit on database/migrations/2026_10_01_add_orders_index.php. On a branch with no ticket:
--- branch: feat/add-orders-index
{
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason": "Protected path (.../database/migrations/2026_10_01_add_orders_index.php). Create a branch named with a change ticket, e.g. feat/CHG-1234-add-index, then retry."
}
}
exit=0
On feat/CHG-4821-add-orders-index, the same edit returned "permissionDecision": "ask" with the reason "Protected path under CHG-4821. A human must approve this edit." An edit to app/Models/Order.php produced no output and exit 0 — no opinion, normal permissions apply. Exactly the behavior I wanted.
The fail-open trap I promised you
Now delete || true from the ticket= line and run it again on a branch with no ticket:
exit=1
No JSON. Just exit code 1. Here's why that's a disaster: grep returns 1 when it finds no match, set -e kills the script, and per the hooks documentation, exit code 1 is a non-blocking error. The tool call proceeds. The one situation the gate exists for — no ticket — is precisely the situation where it silently waves the edit through.
The docs are explicit about the exit codes: exit 2 blocks the tool call no matter what; exit 0 with valid JSON applies your decision; anything else, including 1, is a non-blocking error. So the rule I now follow for every gate script:
- Never let a gate script's failure path be exit 1. Either emit a JSON decision, or
exit 2with a reason on stderr. - Test the deny path, not the happy path. The happy path working proves nothing about the gate.
- Treat
set -ein hook scripts as a footgun unless every command that can legitimately "fail" (grep, test, diff) is guarded.
Two more gaps worth naming honestly. The matcher above covers Edit|Write. An agent with Bash access could still write a migration with sed -i or a heredoc, so pair the hook with permission rules that restrict Bash in protected paths, or add a Bash-matched check. And the test-file lock from the playbook's second example is the same script shape: match tests/ paths, deny when a marker such as .claude/fix-mode exists in the worktree.
If you'd rather have someone build this governance layer for your repos — hooks, skills, the verifier, and the CI wiring — I take on exactly these Claude Code setup engagements. You can see what I've built at fiverr.com/s/EgxYmWD.
Hooks hold the line inside a session. The second trap shows up when you move the agent into your pipeline.
Deploy plays: AI in PR review and CI/CD with claude -p
Claude on both sides of the review
The playbook has Claude giving and receiving reviews. It reviews incoming PRs against organizational policy, and it addresses review comments on its own PRs. Engineers move up a level: judge intent and risk, not formatting.
The lowest-effort route is the official GitHub Action. Running /install-github-app inside Claude Code installs the Claude GitHub App, stores the secret, and opens a PR with the workflow. Manually, the core is a few lines — this is the current shape from the docs, using anthropics/claude-code-action@v1:
name: Claude Code
on:
issue_comment:
types: [created]
pull_request_review_comment:
types: [created]
jobs:
claude:
if: contains(github.event.comment.body, '@claude')
runs-on: ubuntu-latest
permissions:
contents: write
pull-requests: write
issues: write
id-token: write
actions: read
steps:
- uses: actions/checkout@v6
with:
fetch-depth: 1
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
claude_args: "--max-turns 10"
Then @claude address the review comments on this PR does what it says. A --max-turns cap plus a workflow timeout is the cheapest insurance against a runaway job; the docs recommend both.
A policy gate with claude -p
For a deterministic pass/fail check in any CI system, non-interactive mode is the tool. Here's a policy review that fails the build when the verdict isn't "pass":
git diff origin/main...HEAD | claude --bare -p \
"Review this diff against the policy in the system prompt. Judge only what is in the diff." \
--append-system-prompt-file .claude/review-policy.md \
--settings .claude/ci-settings.json \
--allowedTools "Read,Grep,Glob" \
--output-format json \
--json-schema '{"type":"object","properties":{"verdict":{"type":"string","enum":["pass","fail"]},"findings":{"type":"array","items":{"type":"string"}}},"required":["verdict","findings"]}' \
| jq -e '.structured_output.verdict == "pass"'
--json-schema puts schema-conforming output in the structured_output field, and jq -e exits non-zero when the expression is false, so the job fails cleanly. Piping the diff in means Claude doesn't need Bash to read it.
The CI trap: --bare deletes your gates
Look at the --bare flag in that command. The headless docs recommend it for CI — it starts faster and gives the same result on every machine, and the docs say it'll become the default for -p in a future release.
Read what it skips: hooks, skills, custom commands, subagents, plugins, MCP servers, auto memory, and CLAUDE.md. All of it.
So the playbook's promise that "hooks run wherever Claude acts" has an asterisk. If your pipeline follows the docs' recommendation and runs claude --bare -p against your repo, every approval gate in .claude/settings.json is gone. Nothing errors. The agent just runs ungated. That's the same failure shape as the set -e bug, at pipeline scale.
The fix is the reason --settings .claude/ci-settings.json sits in the command above: bare mode only loads what you pass explicitly, so pass your hooks back in through --settings, and your policy through --append-system-prompt-file. Then prove it — run a canary job that attempts a protected edit and asserts it was denied. I'd do that before trusting any agent pipeline with write access. A mirror-image risk exists too: without --bare, a -p run executes the project's hooks and connects its .mcp.json servers with no workspace trust dialog, so running Claude on an untrusted fork's code is its own problem.
The rest of the CI/CD play is operational discipline, and I agree with all of it: sandbox long-running agents, expose deployment through MCP integrations rather than raw credentials, and rehearse the rollback path before the agent ever needs it.
One stage left, and it's the one that makes the loop a loop.
Maintain: closing the loop on metrics
Every stage so far needs a person to start it. Stage six changes that. A continuously running monitoring agent sees a bug ticket or a breached control band, writes a new intent.md, and the work flows through requirements, plan, build, test, and review — headless — with an independent confidence gate between stages. That gate is either a deterministic check or an adversarial reviewing agent, and it decides whether output continues or escalates to a human.
This is the end state the playbook describes: you start by prompting each step by hand, and eventually each accepted artifact fires the next gate on its own. Human attention concentrates at the gates, reviewing what the agent flagged.
I believe in the direction. I'm skeptical of the timeline for most teams. Autonomous intent creation is only as good as the control bands you define, and most teams I've seen can't say what "breached" means for their service beyond "it's down." Define the bands first. Let an agent draft intent.md from alerts, but keep a human accepting it for a long while. A false positive here costs more than a bad commit. It costs an agent a full day building a fix for a problem that never existed.
Which raises the practical question: where do you actually start?
The dependency graph is a rollout order
The most underrated page in the playbook is the dependency graph. It tells you which plays need which others, and read top to bottom, it's a rollout plan:
| Row | Play | Needs |
|---|---|---|
| 1 | Capture intent, CLAUDE.md, Skills, Feedback loop, Hooks, Plan mode | Nothing |
| 2 | Subagents | CLAUDE.md (feedback loop helps) |
| 2 | Evals | CLAUDE.md + feedback loop |
| 3 | Requirements & design | Capture intent + Skills |
| 3 | PR review | Evals + Subagents (skills help) |
| 4 | CI/CD | PR review + Hooks |
| 5 | Closing the loop | CI/CD + Capture intent |
CLAUDE.md helps skills but isn't required for them.
Six plays have zero prerequisites. You don't have to pick all six. I'd start with four: CLAUDE.md, the feedback loop, plan mode, and hooks. My reasoning:
- CLAUDE.md and the feedback loop are prerequisites for evals and subagents, which open the door to PR review. They also pay off on day one in a single engineer's terminal.
- Plan mode costs nothing to adopt and produces the first reviewable artifact before any code exists.
- Hooks are the one Row 1 play that works as a control rather than a speed boost. Get your gates right — including the deny-path tests — before you let anything run headless.
What I'd hold back: intent capture and requirements-and-design. Both involve changing how product people work, and that's a people project. Engineers adopting plan mode doesn't need anyone's permission. A PO committing intent.md through a GitHub connector does.
What changes for regulated teams
If you work in finance, healthcare, or anywhere with auditors, the playbook's pitch boils down to this: keep the old control objectives, swap the enforcement mechanism.
Separation of duties survives — the PO accepts intent.md, an engineer approves plan.md, a reviewer approves the PR, and those are different commits by different people. Change control survives — the hook ties protected edits to a ticket, and "ask" puts a named human on the approval. Traceability gets better than most manual processes, because the chain is in Git rather than in a ticketing tool someone can edit after the fact.
Human line-by-line review doesn't disappear either. The playbook reserves it for regulated and critical code, with layered agentic review handling the rest. That's the honest version of the trade: human review stays, and gets aimed at the code where it matters most.
For the security side of agent pipelines specifically — sandboxing, secrets, and what a hook can and can't stop — xCyberSecurity covers that ground in AI cloud security skills for 2026.
Where I think the AI-native SDLC goes wrong
The playbook is good. Here's where I'd expect real teams to hurt themselves with it.
Over-automation before the gates are proven. The loop diagram is seductive. Teams will wire claude -p into CI before they've tested a single deny path. Both traps above are invisible in a demo, because demos only exercise the happy path.
Rubber-stamping agent review. Once an agent reviews every PR and the findings are usually right, humans stop reading them. That's how "layered agentic review" degrades into zero review with extra steps. The countermeasure is boring: sample agent-approved PRs and re-review some by hand, and track how often humans overturn the agent.
Stale evals. The playbook warns about this, and I'll repeat it louder. A suite where every case passes after a model upgrade isn't a green light. It's a sign your cases stopped discriminating. Add a new case every time monitoring catches something.
intent.md becomes the new PRD. Give it six months and someone will add twelve required sections and a sign-off matrix. Guard the template's length like you'd guard an API contract.
Skills sprawl. "Write a skill for institutional knowledge" turns into eighty skills nobody owns. Each skill needs an owner and a reason it isn't a CLAUDE.md line.
Honestly, I'm not sure the full closed loop is a good target for small teams at all. For a five-person product team, Rows 1 through 3 probably capture most of the value, and the autonomous Stage 6 may never earn its operational overhead. That's a judgment call the playbook leaves open, and I think it should be stated more plainly.
A 30-day adoption path
Here's how I'd sequence the first month, using the playbook's dependency order and my four-play starting set.
Days 1–7: context and verification.
- Write CLAUDE.md with commands, layout, and a "Known mistakes" section. Have the engineer who knows the codebase best own it.
- Make sure test, lint, and static analysis commands run headless and fast. That's your feedback loop.
- Adopt the habit: every Claude prompt ends with "verify, then report."
Days 8–14: plans and gates.
4. Default every non-trivial task to plan mode. Commit plan.md next to the work.
5. Write one hook — the protected-path gate above. Test the deny path with a piped payload before enabling it.
6. Add the verifier subagent and use it on every task that touches more than a couple of files.
Days 15–21: shared knowledge and review.
7. Turn your top one or two cross-repo policies (security checklist, PII rules) into skills with named owners.
8. Install the GitHub Action and let Claude respond to @claude on PRs. Keep humans as the approvers.
9. Start a small eval set: ten cases from real past bugs.
Days 22–30: the first pipeline step.
10. Add the claude -p policy gate to CI, passing hooks back in through --settings.
11. Run a canary job that attempts a protected edit and asserts the denial.
12. Pilot intent.md with one PO and one product area. Keep the template short.
Notice what's missing: autonomous monitoring and self-generated intent. That's month three at the earliest, and only after the gates have survived real traffic.
The hook I wrote this morning works. It works because I broke it before trusting it. That's the whole AI-native SDLC in miniature: speed is easy now, and the artifacts and gates are what make speed safe to keep. Pick one gate this week, pipe it a payload that should fail, and watch what actually happens.
FAQ
Frequently Asked Questions
Everything you need to know about this topic
The AI-native SDLC is Anthropic's playbook for redesigning the software development lifecycle around coding agents like Claude Code. It replaces the linear six-stage process with a loop where each stage commits an artifact to Git and governance runs as hooks, skills, and evals. See "The line becomes a loop" above.
Yes. Anthropic published it on August 21, 2026 as a downloadable playbook and as a free 14-lesson course on Claude Academy covering twelve plays, from capturing intent.md to closing the loop on production metrics.
intent.md is a short Markdown proto-spec that starts every piece of work. The originator brainstorms it with Claude, stating what is wanted, why, and under which constraints, and a product owner corrects and accepts it before commit. A template is in the plan and design section above.
Yes. A PreToolUse hook can return a permissionDecision of "ask", which pauses the tool call until a person approves it. Returning "deny" blocks the call outright. Exit code 2 also blocks, but exit code 1 is non-blocking, so a crashing gate script fails open.
They do by default, but not with --bare. Bare mode skips project hooks, skills, subagents, MCP servers, and CLAUDE.md, so you must pass hooks back in with --settings. Test it with a canary job, as described in the deploy section above.
Let's Work Together
Looking to build AI systems, automate workflows, or scale your tech infrastructure? I'd love to help.
- Fiverr (custom builds & integrations): fiverr.com/s/EgxYmWD
- Portfolio: mejba.me
- Ramlit Limited (enterprise solutions): ramlit.com
- ColorPark (design & branding): colorpark.io
- xCyberSecurity (security services): xcybersecurity.io