Skip to main content
Claude Code

Agentic OS on Claude Code: The Three-Layer Build

Architecture, memory, observability: the three-layer agentic OS I actually run on Claude Code, with the real configs and memory files behind it.

7 min
Read time
1,258
Words
Published
Last revised
Engr Mejba Ahmed

Written by

Engr Mejba Ahmed

Share Article

Agentic OS on Claude Code: The Three-Layer Build

An agentic OS is not a product you download, and I have grown suspicious of anyone selling one. It is three disciplines layered on top of Claude Code: architecture (what the agent is allowed to do and how work is organized), memory (what survives between sessions), and observability (what you can verify actually happened). I say this with some confidence because I am not describing a concept; I am describing my literal setup, the one that runs my site, my content operations, and a chunk of my client work. Every file I mention below exists on my machine. Here is the build, layer by layer, including what each layer looks like when it is missing.

Agentic OS on Claude Code: The Three-Layer Build - overview of layer one: architecture — constitution, org chart, borders, layer two: memory — what survives the session

Layer One: Architecture — Constitution, Org Chart, Borders

Architecture answers: when a session starts, what does the agent already know it may do, must do, and must never do?

The constitution is CLAUDE.md, and it is hierarchical. My global file (~/.claude/CLAUDE.md) carries standards that apply everywhere: security non-negotiables (no hardcoded secrets, parameterized queries only, flag exposed credentials immediately), git rules (conventional commits, never push unasked), and a workflow spine of plan, implement, verify. Each project's CLAUDE.md carries that repo's reality: the domain map, the verify commands, and the accumulated sharp edges. The most valuable lines are corrections written down once so they never recur; my project file knows that composer test runs static analysis rather than the test suite, and that certain innocuous-looking method calls power a six-locale setup and must not be "cleaned up." A convention written in CLAUDE.md is executable culture. A convention in your head is a future incident.

The org chart is skills. I run about 35 in ~/.claude/skills, each with one narrow job: a family of SEO auditors, a design critic, a test-first enforcer, a token-diet compressor, a session-handoff writer. Narrow matters, because skills are routed by their descriptions, and a skill that claims everything matches nothing reliably. The full argument for skills as the unit of organization is in Claude Skills: the automation feature nobody talks about.

The borders are permission allowlists. .claude/settings.local.json pre-approves classes of safe commands (test runners, linters, read-only git and CI checks) while every destructive operation still requires me. This asymmetry is what makes autonomy safe enough to be useful: the agent moves freely inside the fence and knocks at the gate.

Missing this layer looks like: the slot machine. Every session starts from zero, you re-explain your standards daily, and output quality depends on how well you prompted that particular morning.

Layer Two: Memory — What Survives the Session

Memory answers: what does the system know next week that it learned this week? This is where most "agentic OS" writing gets hand-wavy, so let me be concrete about what actually persists on my setup and in what form.

An index plus topic files. My persistent memory is a MEMORY.md index pointing at per-topic files, one per long-running concern. Real examples from my project's memory directory: a file tracking my site's SEO recovery (what crashed, what shipped, what dates matter), one tracking translation status across six locales, one per completed remediation project with its gotchas. The index stays short; depth lives in the topic files. When a session touches a concern, it reads the relevant file and updates it before ending. That is the entire mechanism. It is unglamorous and it works.

Durable manifests for resumable jobs. For batch work spanning many sessions, chat history is not state; files are. When I rewrote 84 blog posts across seven batches, the system of record was a JSON status manifest (done/pending per post) plus a row-backup file per post written before any change. Any session, on any day, could resume by reading the manifest. This pattern (manifest plus backups, no memory in the conversation itself) is the single most reusable idea in this post.

The lesson I paid for: append-only, human-readable notes beat clever memory infrastructure for a solo operator or small team. I have watched elaborate vector-memory setups decay into junk drawers, because what agents need across sessions is not semantic recall of everything ever said; it is a small set of curated, current facts. Write those down in markdown. The six maturity levels of this practice are mapped in Claude Code memory systems, and if your notes already live in a vault, Obsidian as persistent Claude Code memory covers that integration.

Missing this layer looks like: competence without continuity. Sessions are individually excellent and collectively amnesiac; every multi-day project pays a context-reconstruction tax that compounds with size.

Layer Three: Observability — Trust Through Verification

Observability answers: how do you know what the system actually did? With agents doing real work (editing production content, touching client code), "it said it did it" is not an answer.

My practice, in descending order of how often it has saved me:

  1. Verification is mandatory, and mechanical. Every change runs the project's gates: formatter, static analysis at a strict level, relevant tests, and for anything user-facing, a real browser check. A pre-commit hook blocks unformatted code from ever landing, which binds the agent and me equally.
  2. Backup before mutation. Any batch job that changes data writes a restorable backup of each row before touching it. Not because the agent is usually wrong; because "reversible" converts a possible disaster into an annoyance. Invariant, not preference.
  3. Git as audit log. Small, conventional commits with real messages mean the history is the observability. When my blog redesign shipped, the commit message contained the measured contrast ratios; six months later that is documentation no dashboard would have kept.
  4. Run reports and usage audits. Long jobs end with a written summary of what changed, what was verified, and what is pending, and /insights periodically audits how I use the tool itself. If you want this layer as a literal surface rather than a practice, dashboards over agent activity are a real option, and the visual-layer end of this idea is its own rabbit hole.

Missing this layer looks like: slow-motion trust erosion. The agent is right 95% of the time, the unverified 5% accumulates silently, and one day you stop delegating, which quietly kills the whole system.

Build Order, and What to Skip on Day One

Do not build all of this in a weekend; mine accreted over months of real work, each piece added when its absence hurt. The order that works:

  1. Day one: project CLAUDE.md with your verify commands and known traps, plus a permission allowlist. Thirty minutes, immediate payoff.
  2. First week: a memory file for your one longest-running concern, updated at session end. Skip the vault, skip embeddings.
  3. First month: encode your two most-repeated procedures as skills. Then, only when a batch job appears, adopt the manifest-plus-backups pattern.
  4. Skip until it hurts: dashboards, multi-agent orchestration, anything with "swarm" in the name. Several buried Claude Code features carry you a long way first, and I cataloged those in features most developers never find.

The point of an agentic OS is not sophistication. It is that Tuesday's work builds on Monday's, provably, without you being the message bus. Architecture makes behavior consistent, memory makes progress cumulative, observability makes both trustworthy. Three layers; everything else is decoration.

A good number of the components in my three layers (the SEO auditors, the handoff writer, the token compressor, the skill that builds skills) are published on my agent skills marketplace. If you are assembling your own OS, start by installing one layer-two or layer-one piece from there and grow it the same way I did: only when the absence hurts.

Coffee cup

Enjoyed this article?

Your support helps me create more in-depth technical content, open-source tools, and free resources for the developer community.

Related Topics

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 8+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Related Articles

Browse All

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support