YouMind
تسجيل الدخول

The Missing Layer: Structuring AI Agents for Organizational Hierarchies

@joseemv88
الإنجليزية02 أكتوبر 2026
200K
82
7
26
18

ليرة تركية؛ د

This article proposes a reference model for adding hierarchical layers to AI agent platforms like Claude, addressing gaps in instruction inheritance, permission scoping, and conflict resolution to improve enterprise governance and reduce token costs.

A reference model for how organizations structure their work with AI

Jose Martinez · 1 October 2026 · v1.8.1 (measured in Claude Code; re-checked against Anthropic’s documentation on 1 October 2026)

In five lines.

  1. Organizations are trees: firm, line of work, project, task. Claude’s projects are flat: chat has one 3,000-character organization block, projects in chat and Cowork don’t nest or inherit, and in a project a connector account belongs to a person or to the whole organization, never to a branch.
  1. So every project gets a hand-made copy of its line’s rules, the copies drift, and nobody can see which rule produced an answer.
  1. Anthropic has already built the tree twice. Claude Code nests instructions by folder, and since 1 October 2026 it runs an organization’s mods ahead of a person’s. Claude Tag, in Slack, inherits instructions and credentials from organization to workspace to channel. Neither reaches a project. The same model can: a line-of-work layer, projects born from their folders, typed tasks, a compiler that rejects conflicts before the model sees them, and permissions per node that work on Team.
  1. I measured it in Claude Code, twice. An organization rule and a project rule that disagreed, in separate CLAUDE.md levels and neither marked binding: the project rule won 20 of 20 times. The same organization rule marked as enforced, in words: it won 20 of 20. Precedence was set by wording that anyone editing any layer can change, not by the structure.
  1. It runs today. A desktop app I built runs Claude Code inside the tree. It mounts each layer as a CLAUDE.md, limits Claude to what the person may do on that node, rejects the conflict when it is written, and logs every answer with its rules, its tokens and its reviewer. Measured, the tree loaded 20–34% fewer instruction tokens than flat copies.

Organizations are trees. A firm sets policy, a line of work sets its standards, a project applies them to one job, and a task delivers one thing. Every quality system I have worked in is built this way, and so is the folder tree on almost every company file server: firm, line of work, year, then one folder per job, named by a job number that encodes all of it.

AI projects are flat. I use Claude as the case because it is the product I work in every day, and because it already ships the answer in two of its own surfaces. In Claude chat, projects cannot be nested, and organization instructions are a single block of up to 3,000 characters that applies to everyone. On Enterprise, admins can scope permissions by group. On Team, roles apply to the whole organization. Nothing documented in chat or Cowork carries instructions down to a department or line of work and on into its projects. On Enterprise, skills and projects can be shared with a group, but that is distribution, not inheritance.

That is the missing layer: the one between the organization and the project. This article sets out what is missing, why it matters, what it really costs, and a reference model any platform could implement, including its permissions and its data structure.

1. What exists today

Checked against Anthropic documentation on 29 September–1 October 2026. Claude has four surfaces where a team’s work runs, and each has its own instruction model:

Jose Martinez - inline image

Permissions depend on the plan:

Jose Martinez - inline image

Anthropic has built half of the tree for permissions.On Enterprise, custom roles are assigned to groups; “custom roles also control which connectors, and which tools on those connectors, a role can use”; across the platform, organization, role and user levels “the most restrictive level wins”, while a member’s several roles add up; admins can “View effective role”, with a “Granted by” label; and groups can have their own spend limits. But these are flat groups, not a tree, and none of it reaches instructions. The nearest thing is plugins: on Enterprise, an owner can make a plugin and its skills required or installed by default for one group, with a stated order (“group setting, then org-wide setting, then marketplace default”). That targets a group of people, not a project, nothing inherits between projects, and a skill still loads only when Claude judges it relevant. Team has roles for the whole organization and sharing person by person; its plugin settings have, in the documentation’s words, “no group setting”. And Team is the plan built for small and mid-sized firms.

Jose Martinez - inline image

Anthropic has built the whole tree twice, outside projects.

In Claude Code, by folder. Claude Code loads CLAUDE.md files from four levels; “all discovered files are concatenated into context rather than overriding each other”, ordered “from the filesystem root down to your working directory”, and files in subdirectories “load on demand”. An organization’s block can arrive as a managed file on each machine or as text from the admin console (the claudeMd key). The documentation is candid about the limit: Claude treats these files “as context, not enforced configuration”, and “if two instructions contradict each other, Claude may pick one arbitrarily.”

In Claude Tag, by Slack channel. Claude Tag is Claude in a team’s Slack, in public beta on Team and Enterprise. Its settings attach to a scope, and “a scope is where a bundle applies: Default Slack access (the organization-wide root), a workspace, or a single channel.” Three things the rest of this article asks for are already there:

• Instructions inherit. “Per-scope custom instructions are concatenated, Default Slack access first, then workspace, then channel. A channel’s instructions add to, rather than replace, what’s set above it.”

• Credentials belong to the branch. In channels, Claude acts with service accounts an admin attaches to a scope, and “the credential from the narrowest scope is used: channel beats workspace, which beats Default Slack access.”

• You can see where access came from. Each connector, repository and plugin row says “Inherited from” a wider scope or “Attached from” a bundle. Spend is capped for the organization and per channel, and reported per channel.

The limits are documented too. The tree has Slack’s shape: three fixed levels, and Slack channels don’t nest. Instructions are “guidance, not an enforced guardrail”, and the documentation describes no check for conflicts between scopes. There is “no per-action log of every task and who asked”. And it stops at the edge of a project: “Projects in claude.ai don’t apply here; Claude doesn’t read a Project’s instructions or knowledge in Slack, and a channel can’t be pointed at a Project.”

What changed on 1 October 2026. Claude Code 2.1.287 turned on mods, which had been in early access: functions inside a plugin that run inside Claude Code and can rewrite a prompt or a section of the system prompt, block or rewrite a tool call, approve or deny a permission request, and draw panes in the interface. Three things about them matter here:

• They carry a declared order. A built-in guard and the organization’s own mods run first, then the mods a person installs. “The first mod is outermost: it sees the event before the others and the result after them, and it decides whether the others run at all.” Where the guard loads (a machine with managed settings, or a Team or Enterprise login), a person’s mod can’t change “the system prompt, your managed CLAUDE.md and other managed instructions”, and it can’t approve a tool call that a deny rule refuses. That is precedence set by structure, which is what this article asks for. It is a default, not a lock: a person who starts Claude Code with --safe-mode runs without installed mods, the organization’s included, while managed hooks and deny rules still apply.

• The order has two owners. The organization’s mods run before the person’s, or after them if the organization chooses. There is no level for a line of work. Settings delivered from the admin console “apply uniformly to all users in the organization. Per-group configurations are not yet supported.” An organization that wants a different policy per group has two routes, both through IT: deploy a different settings file to each group’s machines, or run a self-hosted gateway that “delivers managed settings per IdP group”.

• They don’t reach chat, and reach Cowork unevenly. “Mods work in the Claude Code CLI and in the Code tab of the Claude Desktop app.” A mod ships in a plugin’s hooks/hooks.json, and Anthropic’s plugin support table marks that file “Ignored” in chat. The same table marks it “Loads” in Cowork, because “Cowork in the Claude Desktop app runs its sessions on Claude Code”; the mods pages don’t list Cowork, and I haven’t tested it. What the organization controls there is thinner: in a Cowork session, Claude Code “never fetches server-managed settings from the claude.ai admin console, even when the user signs in with a Team or Enterprise account”, and remote Cowork sessions have no device policy to read.

What is rolling out. Cowork is merging into Claude: the help center now says “Claude Cowork is now just Claude”, on Pro and Max first, while Team and Enterprise “keep chat and Claude Cowork as they are today”. On 6 October 2026, new Cowork tasks on Pro and Max move to the cloud. And a new version of projects is in public beta on Pro and Max, starting in Claude Code, with chat, Cowork, Team and Enterprise to follow; in it, “a project is one conversation” that Claude splits into parallel threads. It is still a single level: “A project belongs to one user”, and “there are no organization-level controls for projects during the beta.”

So the tree is not a new idea to Anthropic. It exists by folder for code and by channel for Slack, and in both the organization comes first. The place where a firm files its work, the project, has neither a parent nor a line above it. Two other Anthropic offerings point the same way and are outside this article’s scope: Claude for Government resolves settings through a tenant, group and organization chain, and Claude Desktop on third-party providers has per-group policies in beta.

2. Why it matters

Projects are not containers. A real job is a container. It has a key (the job number), a parent (its line of work), a lifecycle (open, closed, archived by year), a folder, a client, people, rules and deliverables. A Claude project has a free-text name, no key, no parent and no children, and its documented lifecycle is archiving and deleting. Cowork can bind a project to a local folder, but only by hand, one project at a time, with no key pattern and no parent. A line of work may open hundreds of jobs a year. That leaves two bad options: one hand-made Claude project per job, each with its own copy of the rules, or one project per line of work, where different clients’ context sits side by side.

The load path breaks. In structural engineering, every load needs a continuous path to the foundation. Take out one member and nothing above it transfers down. Rules behave the same way. When there is no layer between the organization and the project, a line of work’s standards, templates and sign-off rules have nowhere to live. So every project gets a hand-made copy.

The copies drift. Fix a rule in one project and the others keep the old version. Each project still passes its own local check. The mismatch only shows up when someone compares projects side by side, and in regulated work that is usually an auditor. Infrastructure teams know this failure well. In Firefly’s 2026 research, about a third of respondents tied configuration drift to costly production incidents, and about one in five had no process to detect or fix it.

Identity is flat too. Many people work across more than one organization, and each needs its own sealed context. Inside any one of them, in chat and Cowork, a connector account belongs to a person, not to a branch. Enterprise roles can decide which connectors a group may use, and an admin can authorize a connector once for the whole organization; a custom connector can even carry one shared credential for everyone (in beta). Either way the account is the person’s or the organization’s, never a branch’s: a consultant serving two clients can’t bind each client’s Drive to that client’s projects. Shared projects make it sharper: “Connectors are only available in private projects.” Claude Tag shows the other design is possible, with a service account attached to a Slack channel, and it shows where it stops, at the project. Anthropic’s Google connector page describes one connected Google account, and the documented way to change it is to disconnect and reconnect; three open issues (below) ask for more than one account. The boundary lives in the person’s head, which is exactly the kind of manual boundary that fails quietly.

Nobody can see which rule applied. Claude Enterprise shows admins a “View effective role” for permissions, and Claude Code’s /context lists which memory files were loaded. A Claude Code mod can now draw its own pane, so an effective-instructions view is something anyone can build there. Claude Tag labels each connector and repository with the scope it was inherited from, but for instructions its documentation says to “ask Claude to repeat its admin instructions”. No surface shows, for a given answer, which instruction came from which layer. Without provenance there is no audit trail, and without an audit trail there is no quality system.

People are asking for pieces of this. Open issues in Anthropic’s public issue tracker (github.com/anthropics/claude-code), checked 2026-10-01:

Jose Martinez - inline image

A seventh, #47741, asked for an organization-managed CLAUDE.md and was closed because Claude Code already has one. That is the point: the layers exist in Code and in Slack, and the issues ask for them where the projects are.

3. What it costs, and what the token hides

3.1 Tokens are not the barrier

More layers could mean more context with every message, and AI is billed by the token. That explanation is only partly right. On usage-based Enterprise plans, usage is billed at API rates, so more context means more revenue, not less. On Team, seats are flat unless extra usage is enabled, and extra tokens show up as members reaching their weekly limit sooner. And Claude Code, billed by the same tokens, already ships a four-level cascade. If tokens were the barrier, it wouldn’t exist. Measured in section 3.2, Claude Code’s own system prompt and tools came to about 30,200 tokens before any instruction of ours; a project’s full instructions added 3.5–5.2% to that.

3.2 A worked example

Tags on every image: REAL = measured, or checked against Anthropic’s documentation, between 29 September and 1 October 2026; EST = simulated; IND = illustrative.

Jose Martinez - inline image

Measured first. I ran the same question in Claude Code (claude -p, Claude Sonnet 5.5) against one demo project, five times per condition, and took the input tokens Claude Code itself reported. I ran it on Claude Code 2.1.286 on 30 September and again on 2.1.287 on 1 October; the instruction counts were identical. Subtracting a run with no project instructions at all leaves what each layout costs:

Jose Martinez - inline image

Two things the simulation below could not show. The cascade costs 224 tokens more than the compiled file for the same rules: every extra file carries overhead, here the app’s own marker and heading on each level’s file plus Claude Code’s framing around each file it loads. More levels means more overhead. And Claude Code actually loaded 21–26% more than the public tokenizer estimated, even × 1.30; part of that gap is the same per-file overhead. Ratios survive it; absolute dollar figures don’t, so read the simulation’s dollars as low.

Then simulated at firm scale. I simulated one month of instruction tokens for an illustrative firm: 40 people in three lines of work, 250 active projects, six report types per line, 35 messages per person per workday in sessions of five, 29,400 messages in total. Token sizes were counted on sample instruction texts with Anthropic’s public legacy tokenizer (which Anthropic itself calls “a very rough approximation” for Claude 3 and later) and scaled by 1.30 for the Claude 4.7+ tokenizer. The organization block is extrapolated from a 597-character sample to the 3,000-character cap, and the line manual is three times a 953-character sample. Prices are Claude Sonnet 5.5 list prices (input $2, 5-minute cache write $2.50, cache read $0.20 per million tokens). The prompt cache lasts five minutes and is refreshed on each hit.

• Flat (today’s workaround): the organization block, then each project’s own copy of the line manual, all six report templates and the project’s specifics.

• Tree: organization, line, only the report template in use, then the project’s specifics, compiled most-shared first.

Jose Martinez - inline image

Sizes behind the table: organization block 830 tokens (3,000 characters), line manual 729, one report template 147, project specifics 98 (rounded; totals were computed before rounding). The simulation ignores Claude’s own system prompt, which sits before the organization block.

Two honest caveats. First, the tree does not save tokens by itself. The −29% comes from typed tasks: only the template for the report being written is loaded, not all six. The −62% comes mostly from compiling the most-shared layers first, so hundreds of projects share one byte-identical prefix: ordering alone, with all six templates still loaded, takes the shared-cache cost from $36 to $18 (−51%), and typed tasks add the rest. That second saving only exists if the cache is shared across users. On the Claude API, caches are isolated between organizations and, within one, per workspace, so identical prefixes are reused across requests within a workspace; for claude.ai this is not documented. It does not happen in Claude Code as shipped: there “the cache is effectively scoped to one machine and directory”, so two people in two project folders miss each other’s cache. Read the last column as what a native layer in chat could earn, not as something available today. A fair objection: Skills already load on demand, so a flat workspace that moves its templates into Skills gets part of the −29% today. What Skills lack in chat and Cowork is scope and inheritance: a skill can’t belong to one line and flow down to that line’s projects. (In Claude Code, a skill in a subfolder does load for sessions started in or below it.) Second, these are instruction tokens only, and at this scale they range from $14 to $149 a month depending on caching. Conversation history and output dominate real bills. The strong argument for the tree is correctness and, on Team, capacity. It is not the invoice.

The measurement above reproduces the typed-task effect on a real tree, in Claude Code, instead of assumed sizes: −20% as a cascade and −34% compiled, against the simulation’s −29%. A line with a single task type would save nothing from typed tasks.

3.3 The same answer, from real Claude

I asked Claude Code the same question forty times: ten independent answers under each of four conditions, in two batches of five a day apart. The question was whether a field density test on subgrade passes, at 112.3 pcf against a 115.8 pcf maximum dry density with 98% required. The instructions were either the firm’s six rules or a one-line “helpful assistant” prompt, and the language was English or Spanish. All forty answers reached the same verdict: 97.0%, fail.

Claude Code reports two output numbers: the tokens billed, and how many of those were thinking that the reader never sees.

Jose Martinez - inline image

Four findings:

• The firm’s rules made the answers 1.4 times longer on screen, and 1.8–1.9 times longer on the bill. The visible extra was the flags, standards and limitations sections the rules asked for. In a quality system, those are the valuable part.

• Under the firm’s rules, over a third of the billed output was invisible. 37% of billed output tokens were thinking, against 16% under the plain prompt. In English, the reader sees 447 tokens and pays for 711.

• Spanish cost 1.2 times the visible tokens of English for answers within 4% of the same length in words. Measured the same way, the firm’s rules took 1.53 times the input tokens when written in Spanish.

• The same verdict was billed from 317 to 953 tokens, three times as much for the longest answer as for the shortest. Per-token billing can’t tell rigor from padding. Acceptance criteria can.

These ratios move between batches of five. The on-screen ratio for the firm’s rules was 1.43–1.52 in the first batch and 1.27–1.31 in the second; the billed ratio was 1.78–1.82 and then 1.70–2.05; the Spanish ratio was 1.29–1.38 and then 1.14–1.18. The direction never changed. The size is good to one significant figure.

Method: Claude Code 2.1.286 on 2026-09-30 and 2.1.287 on 2026-10-01, claude -p --output-format json, Claude Sonnet 5.5. Every tool was disallowed and the personal ~/.claude files were excluded, so only the stated instructions differed between conditions. Token counts are Claude Code’s own usage report, including thinking_tokens; monthly cost applies Sonnet 5.5’s $10 per million output tokens to the billed mean. The measurement script, both batches, the pooled figures and every answer are kept by the author and available on request. The first version of this article estimated these numbers with subagents and the public tokenizer; those estimates are superseded.

3.4 When rules conflict, the wording decides

Anthropic’s documentation concedes that contradictory instructions may be resolved “arbitrarily”. I tested one conflict of the kind a missing line layer produces. The organization rule said to use U.S. customary units; a project rule said to report densities in SI. I placed them the way a tree would: the organization rule in a CLAUDE.md at the storage root, the project rule in a CLAUDE.md in the project folder, both loaded by Claude Code’s own cascade. Then I ran the same setup with the organization rule marked as enforced, in words only: the tag ENFORCED in its label and one added sentence, “This rule is enforced: no line or project rule may override it.” Each setup ran ten times on 30 September and ten more on 1 October.

Jose Martinez - inline image

It wasn’t arbitrary. With nothing declared, Claude took the nearer, more specific rule every time. Fourteen of the twenty answers said why (“that rule is more specific than the firm-wide U.S. customary rule”, or that it overrides the firm’s rule); five cited only the project rule and never mentioned that the firm’s rule disagreed. With the binding declared in words, the organization rule won every time, and every answer said the enforced firm rule took precedence. The second batch reproduced the units, 10 of 10 each time, and the difference between the two rows is far beyond chance (Fisher’s exact test, p < 0.0001). The explanations held less well: nine of ten answers said why the project rule won in the first batch, five of ten in the second.

That is good news for the model and bad news for the workspace. Precedence exists, but it lives in the wording of the rules, which anyone editing any layer can change, and which nobody reviews as a precedence decision. And when the lower rule won, a quarter of the answers didn’t tell the reader that a higher rule had been set aside. A pilot run in the first version of this article, with both rules in one prompt instead of the cascade, also shifted when only their order changed (org rule first: SI 5 of 5; org rule last: SI 2 of 5, and 3 of 5 gave both units or asked which applies).

No model should be asked to arbitrate a conflict that the organization could have caught when the rule was written. Claude Code has two partial answers. /doctor prompt-audit asks Claude to look for instruction files that contradict each other, when a person runs it. And since 1 October a mod can enforce an order in code: the organization’s mods run before a person’s, and where the built-in guard loads, deny rules beat a person’s mod. Neither covers instruction text. Instruction files are still concatenated, and the documentation describes the outcome three ways: Claude “may pick one arbitrarily”; when a user rule and a project rule conflict, “Claude may follow either one”; and “when instructions conflict, Claude uses judgment to reconcile them.” The twenty-of-twenty result is what that judgment looked like here. Claude Tag declares an order for its three scopes, and calls the result “guidance, not an enforced guardrail”. The reference implementation’s check rejects this exact change, “R-22 sets units.density=SI; R-01 (org:firm) enforces US”, before anything reaches Claude, and enforced is a field on the rule, not a sentence in it.

Instruction density makes it worse. In the IFScale benchmark (2025), Claude Sonnet 4’s accuracy fell from 100% at 10 simultaneous instructions to 42.9% at 500.

3.5 Would compute or energy be a fairer unit?

The token is a reasonable proxy for compute within one model: more tokens really are more work for the hardware. That is also why a compute- or energy-based price would not fix the language penalty. Spanish costs more because the tokenizer compresses it less, and those extra tokens are real compute. The fix for that is a better tokenizer, or metering normalized by content.

A normalized compute unit would still help in three ways. It makes different models and vendors comparable. It is physical and reportable, for example for sustainability reporting. And if the coefficient is fixed against reference hardware, the vendor keeps its own efficiency gains, which is the right incentive. There is precedent: cloud providers once sold normalized units such as the EC2 Compute Unit.

It has real problems. The customer cannot verify it without a standard and an auditor. Actual energy depends on hardware, data-center efficiency, batching and the grid. And Anthropic does not publish energy per request; I found only third-party estimates. Most importantly, compute is still an input. It doesn’t say whether the answer was right.

My conclusion is three separate layers:

  1. Bill in tokens or a normalized compute unit.
  1. Disclose energy per task and per node.
  1. Manage by cost per verified outcome.

For Team, the minimum step is to publish the weekly limit in a stated unit. The per-session allowance is given as “1.25x the Pro plan’s per-session usage allowance”; the weekly limit has no published number at all. Neither can be budgeted.

As the FinOps Foundation puts it, “the token is the billing unit, not the value unit.” A hierarchy is what makes value definable, because it is where acceptance criteria can live.

3.6 More likely reasons it hasn’t been built

  1. Implicit precedence. Anthropic’s own documentation concedes that directly contradictory instructions can produce varying behavior, and section 3.4 shows precedence following whatever the rules’ wording says. Stacking layers multiplies conflicts, and instruction-following degrades with density: in the IFScale benchmark (2025), even the best models tested reached only 68% accuracy at 500 simultaneous keyword instructions (the benchmark cited in section 3.4).
  1. Permission inheritance. If knowledge inherits down a tree, access has to inherit too. That means rebuilding the permission model under every layer.
  1. A preference for memory and retrievalover static layers.
  1. Consumer-first simplicity. Code tools inherit a tree for free from the filesystem. Chat products have to invent one.

Anthropic hasn’t said publicly why Claude chat and Cowork lack hierarchy. Everything in this section is inference from what has shipped.

4. The reference model

The design borrows from systems that already solved this problem: cloud resource hierarchies (AWS Organizations, Google Cloud Org Policy, Azure management groups), directory policy (Active Directory Group Policy), and Claude Code’s own CLAUDE.md cascade. Relations between nodes use five words: contains, inherits, uses, sealed and shared.

Jose Martinez - inline image

4.1 Mount the tree the organization already has

Don’t make people rebuild their organization inside the AI workspace. The file server or document system is already the source of truth. Take one path:

Jose Martinez - inline image

The job number 26GT301 already encodes the tree: year, line of work, sequence. The workspace should mount this structure, not copy it.

Jose Martinez - inline image

4.2 Rule-bearing nodes, grouping nodes, projects and tasks

• Rule-bearing nodes: organization, line of work, project, task. Each carries the same three things: context(instructions and knowledge), policy (which tools, data and connectors are allowed), and identities (the connector accounts bound to it).

• Grouping nodes: series, year, region. They carry no rules. They exist for navigation, retention and lifecycle. Separating them keeps the rule tree shallow, three to four levels, as Microsoft’s own guidance for management groups recommends (“no more than three to four levels”).

• The project is a keyed container. It is created automatically: when a folder matching the line’s key pattern appears (for example {YY}GT{NNN}_{Name} under the line’s root), a project node is created, inherits from its line, and gets connector access scoped to that folder only. It holds only what differs from its line: members, client, specifications. It moves from open to closed to archived (Diagram 2).

• The task is typed. Its type comes from the line’s catalog (a density report, a boring log). The type carries a template and acceptance criteria. The output is filed back into the project folder under the firm’s naming convention, and a reviewer accepts it. Skills are the closest thing Claude has to task types today. On Enterprise they can be shared with a group, but that is distribution, not inheritance: nothing flows down a branch.

Jose Martinez - inline image
Jose Martinez - inline image

The recursion is deliberate. Stafford Beer’s Viable System Model puts it directly: “In a recursive organizational structure any viable system contains, and is contained in, a viable system.”

4.3 One primary parent, plus overlays

Christopher Alexander argued in 1965 that “a city is not a tree.” Real structures overlap. A client, an agency’s specification or a task type can span several lines of work. So every node has one primary parent, and cross-cutting rule sets attach as overlays (uses). Conflicts resolve the same way every time: deny wins, otherwise the nearest node wins.

4.4 Two channels, two semantics

This is the core of the design, and it is where most hierarchies go wrong.

• Context concatenates. Instructions and knowledge are merged from the root down, as CLAUDE.md does.

• Policy is deny-by-default. A tool or connector is allowed only if an allow exists on the whole path from the root, and an explicit deny anywhere above wins, as AWS Service Control Policies do. A parent can mark a rule as enforced, and no child can block it, as in Group Policy.

Mixing the two is the classic mistake. Advisory context should blend. Enforcement should not.

4.5 Compile before the model reads it

Today, conflicting instructions are resolved by the model at answer time. The fix is an effective-instructions compiler that runs before the model sees anything:

  1. Merge context from root to leaf.
  1. Apply policy: deny wins, and an allow must hold on the whole path.
  1. Honor enforced rules from parents.
  1. Stamp every rule with an ID and its layer.
  1. Order the block by how widely each part is shared, and apply a token budget per layer.

Conflicts never reach this step: they are rejected earlier, when a rule is written, and approval comes from someone other than the author (Diagram 3, lower lane).

Because conflicts are resolved at compile time, the output can be ordered by how widely each part is shared, not strictly by depth: organization, line, task-type template, then project specifics. That order maximizes cache hits (section 3.2). It maps onto Anthropic’s four cache breakpoints, with a token budget per layer, though in practice one breakpoint may be needed for the conversation itself.

Jose Martinez - inline image

4.6 Identities bound to branches, not people

A connector identity (account, tenant, scope) attaches to a node, not to a person. Claude Tag already works this way for Slack channels: an admin attaches a service account to a scope, and the narrowest scope’s credential wins. The model here asks for the same thing one level down, on a line of work and its projects. A person working in two organizations has two sealed trees; they switch trees, not accounts. Nothing crosses between them unless the owners of both explicitly share it. The technical primitive already exists: the MCP authorization specification uses audience-bound OAuth tokens (RFC 8707 resource indicators, RFC 9728 protected resource metadata) and requires that servers “MUST NOT accept or transit any other tokens” (specification version 2026-07-28).

4.7 Permissions follow the tree

Principals: people, groups, service accounts, external guests (for example a client), and the agent itself.

The agent never exceeds the person or the node. Claude acts with the invoking user’s permissions, intersected with the node’s policy. It cannot change rules or permissions; it can only propose changes. (Today, Claude can update Cowork folder instructions on its own. In this model, that becomes a proposal someone approves.)

Jose Martinez - inline image

Evaluation. Grants flow down only, never up or sideways. The effective permission at a node is what the roles on its path grant, within what policy allows on the whole path, minus any deny above. An overlay grants access to its own content only: reading a specification does not open the projects that use it.

Lifecycle.

• Open: roles apply as granted.

• Closed: no new tasks, but pending reviews can finish.

• Archived: read-only for everyone; only the owner can restore, and the restore is logged.

• Sealed tree: nothing crosses without an explicit share.

Exceptions and delegation. Exceptions are time-bound and justified, and someone other than the requester approves them. They expire on their own and are counted, because every override is a permanent maintenance island; SharePoint’s broken-inheritance limits are the cautionary tale. Delegation can never grant more than the delegator holds. Owner break-glass access exists, is always logged, and is reviewed afterward.

On Team, this works without groups: the grant lives on the node, so a four-role organization still gets per-branch permissions.

Jose Martinez - inline image
Jose Martinez - inline image

4.8 How the layers interact

A tree is only worth building if changes travel through it. Three interactions carry most of the work (Diagram 5):

• Push down. A line lead publishes a new version of a rule. Every project in the line reads it at its next compile. An approved exception keeps the old version until it expires, and artifacts already filed keep the version they were made with.

• Pull up. A member improves a template inside one project and proposes it. The line lead, not the proposer, approves it, and the sibling projects inherit it. Today that improvement stays in the project where it happened.

• Across. An agency specification changes once. Projects in three lines recompile with it, and the lines themselves don’t change. A conflict with a line rule is rejected when the update is written.

One request shows all the layers at once (Diagram 6): the tree checks the member’s grant and the project’s state, the compiler builds the block, Claude reads field data through an identity scoped to that project’s folder, files a typed artifact back into the folder, and a reviewer who didn’t write it accepts it. Every step lands in the log, and the cost is charged to the project key.

Jose Martinez - inline image
Jose Martinez - inline image

4.9 The data structure

Jose Martinez - inline image

Provisioning is event-driven: a new folder matching key_pattern under a line’s storage_root creates the project node. The answer_log gives provenance for every answer and cost per node.

4.10 Measure outcomes, not tokens

Each task type carries acceptance criteria: the definition of done. With those in place, a better unit of AI work becomes measurable:

cost per verified outcome = (token cost + review time) ÷ accepted deliverables

Because every answer is logged against a node, AI cost can be charged to a project key the same way labor and materials are. Pieces of this exist: Claude Tag reports and caps spend per channel, and Claude Code’s telemetry can be tagged by department, cost center or repository, by hand. None is keyed to a project, and none divides by accepted deliverables. For a firm that bills by job number, AI becomes a direct job cost instead of overhead. In engineering, you pay for the checked, sealed deliverable, not for pencil lead. AI work should be measured the same way.

4.11 Objections, answered

“Skills and plugins already do this.” A skill loads when Claude judges it relevant, which is relevance, not a guarantee. Provisioning gives a skill to everyone; on Enterprise, a plugin that carries it can be required for one group. That is the closest thing to a line of work in chat and Cowork today, and it falls short in three ways: it is Enterprise only, it targets people and not projects, and nothing flows down from a line to its projects. A rule that must always apply in one line can’t depend on relevance detection.

“Mods already do this.” In Claude Code, partly, since 1 October 2026. A mod can rewrite the system prompt, refuse a tool call and draw a panel, and an organization’s mods run before a person’s. So the compiler, the write-time check and the effective-instructions view in this article could be built as a mod today, and section 6 says so. Three limits remain. Mods don’t run in Claude chat, and in Cowork the organization’s console settings don’t apply. Their order has two owners, organization and person, with no line of work between them; on Enterprise, a plugin required for one group can carry a mod to that group, but it runs as one of the person’s own mods, with no precedence. And a mod is unsandboxed code: to run ahead of people’s mods, an organization’s mod must sit in a directory on each machine, and settings delivered from the admin console “can’t put the directory on a machine”. A firm with no device management can push a mod to everyone, but it runs among people’s own mods, not ahead of them. A firm should not need to write TypeScript to say that one department reports in different units.

“Memory will learn the rules.” Memory is written mostly by Claude, for one person or one project, and an owner can’t read or edit a member’s memories. An auditor needs rules that were written by a person, versioned, approved and traceable to each answer. Anthropic has built that three times already: for permissions, with “View effective role” and its “Granted by” label; for skills and plugins, with version history and a review step where “you can’t approve your own”; and for Claude Tag’s access, with its “Inherited from” labels. Instructions in a project have none of the three.

“Claude Tag already does this.” For Slack channels, largely yes, and section 1 says so. Three things are still missing. A channel is not a project: it has no folder, no key, no lifecycle, and “a channel can’t be pointed at a Project”. The tree is three fixed levels, so a firm with lines of work and hundreds of jobs has to flatten two of its levels into channel names. And the documentation describes no check for a conflict when an instruction is written; the scopes are concatenated and the model is left to reconcile them. If anything, Claude Tag is the strongest evidence for this article’s design: the same company chose inheritance, scope-bound credentials and an origin label when it built for teams.

“Hierarchies add complexity.” Only if their depth is unlimited. Microsoft’s own guidance for management groups is “no more than three to four levels.” This model fixes four rule-bearing levels, and grouping folders such as series and year carry no rules at all.

“Inheritance is a security risk.” It is, if access is inherited carelessly. The cloud answer applies: an allow must exist at every level, a deny anywhere wins, Claude acts as the user intersected with the node, and Claude’s own rule edits become proposals.

“More layers cost more tokens.” Section 3.2 measured the opposite in Claude Code: 20% fewer instruction tokens as a cascade, 34% fewer compiled, against flat copies. Each extra file does add a little overhead, so compiling beats cascading. Typed tasks drop the templates that aren’t in use, and compiled prefixes are byte-identical across projects, so prompt caching can reuse them. On the Claude API, caches are isolated per workspace, so a workspace per organization maps onto the root of the tree. Claude Code doesn’t get this today: its cache is scoped to one machine and directory.

“Teams can just maintain their own projects.” That is today’s workaround, and the prototype in section 5 measured what it produces: two of six pasted copies out of date in a small demo.

5. It runs today: a reference implementation

Jose Martinez - inline image

To show that the model is buildable, and not only arguable, I built Worktree, a small desktop app (Node and Electron, 24 passing tests on Windows and Linux) laid out like Claude Desktop. The video above is a real run of it, shortened only where Claude was working. It is a separate app that drives Claude Code from outside, not a mod. It does not call a model of its own. Every chat runs the Claude Code already installed on the computer (claude -p), with whatever login Claude Code has: a Claude subscription or an API key. I ran it on an invented firm with three lines of work and six projects. When someone sends a message:

• Permit. The person acting needs a grant on the project or above it, and the project must be open. An admin with no grant on the line was stopped before Claude started.

• Check. The write-time check runs first. The SI rule from section 3.4 is rejected as a conflict with an enforced firm rule, so it never reaches a CLAUDE.md.

• Mount. The tree is written into the real project folders as one CLAUDE.md per level (organization at the storage root, then line, then project), and Claude Code’s own cascade loads them. A compiled mode writes one file per project instead. Files without the app’s marker are never overwritten.

• Run. The task template goes in --append-system-prompt-file. What Claude may do is enforced by Claude Code, not by CLAUDE.md: --allowedTools is the person’s role intersected with every level’s policy, writes are limited to the project folder, and --permission-mode dontAsk refuses everything else. In real runs, a web search was refused because the line’s policy doesn’t allow web, and a write outside the project folder was denied and logged.

• Log and review. Every answer is logged with the person, the node, every rule tag, the rules Claude cited, and the input, cache and output tokens Claude Code reported. A reviewer who did not write the answer accepts or returns it. A real density-report run on the demo took about 30 seconds; across three runs, Claude Code reported $0.08–0.22 per answer at list prices. Claude cited the rules it applied, flagged results near the acceptance limit, and left the engineer’s fields blank.

• Drift. Every mounted CLAUDE.md is compared with the tree and hand edits are flagged. An earlier command-line prototype ran the same comparison on copies pasted into flat projects and found two of six out of date: one still on R-07 v3, and one where a rule had been deleted by hand.

Three findings from building it, all relevant to anyone layering instructions on Claude Code:

  1. Your personal instructions leak into the organization’s runs. By default, every run also loaded my personal ~/.claude/CLAUDE.md, rules, agents and MCP servers. The app now excludes them with the claudeMdExcludes setting, plus --strict-mcp-config. On my machine, that took a run’s context from 29.6k to 21.4k tokens. The obvious alternative, --setting-sources project,local, did the opposite of what was needed in Claude Code 2.1.284 on Windows: it kept the personal file and dropped the parent-folder CLAUDE.md files that carry the organization and the line.
  1. A person’s mods leak in too. Mods arrived the day after the app was built, so I tested it. I installed a one-hook mod in my own user scope that adds a line to every prompt. It reached the app’s runs: the organization’s three rules loaded, and so did my personal line, and the answer obeyed it. Adding disableAllHooks to the run’s settings kept it out and left the three levels of CLAUDE.md intact; the app now does this. --safe-mode is not a substitute: it removed the mod and the whole CLAUDE.md cascade with it. Per the documentation, disableAllHooks in a person’s own settings leaves what the organization manages running.
  1. CLAUDE.md is context, enforcement is configuration. Anthropic’s documentation says so: “Settings rules are enforced by the client regardless of what Claude decides to do. CLAUDE.md instructions shape Claude’s behavior but are not a hard enforcement layer.” The app depends on it. Everything a rule must guarantee is mapped to tool permissions; everything in CLAUDE.md is guidance with a tag.

It works on today’s Claude Code, and the compiled text can be pasted into a chat or Cowork project’s instructions on any plan. One dependency is dated: the app relies on claude -p loading CLAUDE.md files, and Anthropic’s documentation says --bare, which skips them, “will become the default for -p in a future release”. When that lands, the app will have to pass the tree another way; the compiled mode and --append-system-prompt-file already can. It is a working specification for the native version, not a security boundary: “acting as” is a demo switch, not a sign-in.

6. A path from what already ships

In Claude Code, now. Since 1 October the tree can ship as a mod: compile the node’s rules into a section of the system prompt, refuse tool calls the node’s policy denies, and draw the effective instructions in a pane. An organization can run that mod ahead of anything a person installs. I have not built this; it is the first item on the reference implementation’s roadmap. It would cover Claude Code, and possibly Cowork sessions on the user’s machine, which run on the same engine. It would not cover chat.

The 30-day version, for chat and Cowork. Anthropic has the parts. Let a project name a parent project whose instructions it inherits, the way a Slack channel inherits from its workspace in Claude Tag. Compile the two in order, with a tag on every rule, and add a “View effective instructions” panel next to the existing “View effective role”. That alone gives every line of work one place to keep its rules.

After that, each step is useful on its own, starting with what helps Team plans:

  1. Line-of-work nodes and key-pattern provisioning, reusing the CLAUDE.md semantics that already work in code. In Claude Code itself, the matching step is a tier for a group between the organization’s mods and a person’s, and per-group managed settings.
  1. Task types as skills scoped to a line, with acceptance criteria.
  1. A write-time conflict check, so contradictions are rejected by the tree, not arbitrated by the model.
  1. Grants on nodes, which work on Team without groups and extend Enterprise’s custom roles.
  1. Branch-bound connector identities for projects, the way Claude Tag already binds a service account to a Slack channel.
  1. Per-node accounting keyed to the project, as Claude Tag already reports per channel; a published weekly limit in a stated unit; and energy disclosure per task.

7. Limitations

• Coverage. On 1 October 2026 the full page indexes of code.claude.com/docs and claude.com/docs were triaged by title (466 pages) and about 150 pages were read, with the help-center articles cited here. Earlier versions of this article missed Claude Tag entirely; this one may miss something else. Claude for Government and Claude Desktop on third-party providers are mentioned and not analyzed.

• Platform features change monthly, and one changed while this was being written. Every product claim here is dated 29 September–1 October 2026 and should be re-checked before relying on it. Mods are one day old; I have read their documentation and tested one case, not used them in production.

• Measured numbers in sections 3.1–3.4 come from Claude Code’s own usage reports, in two batches: Claude Code 2.1.286 on 2026-09-30 and 2.1.287 on 2026-10-01, both on Claude Sonnet 5.5. The script, both batches and every answer are kept by the author and available on request. They include Claude Code’s own prompt (about 30,200 tokens), which claude.ai and Cowork don’t share, and they cover one model, one demo project and one question. Samples are small: ten answers per condition for answer length, twenty per condition for the conflict test. Answer-length ratios moved between the two batches (section 3.3); two batches a day apart also differ by Claude Code version, and I can’t separate that from chance.

• The run time, cost and context sizes in section 5 come from the app’s own answer log of three demo runs and one manual isolation test, which are the author’s own records.

• The conflict answers were coded by the author, unblinded, by reading every answer (units reported; whether the answer said which rule won and why; whether it asked). All forty are available on request for recoding. The test used one wording of “enforced”; other wordings, models and rule pairs may behave differently.

• The mod test in section 5 is one mod with one hook, on one Linux machine, with Claude Haiku, signed in with a subscription and no managed settings. I did not test an organization’s policy mod, the built-in guard on a Team or Enterprise login, or the Desktop app.

• The cost model in section 3.2 is a simulation of an illustrative firm, not measured billing. It covers instruction tokens only; conversation history, output and thinking usually dominate real bills. Its token sizes use Anthropic’s public legacy tokenizer × 1.30; in the measurement, Claude Code loaded 21–26% more than that estimate, so its dollar figures read low. Its percentages are ratios and hold. Session patterns and cache sharing are assumptions.

• Whether chat usage on Enterprise receives API cache pricing and whether caches are shared across users on claude.ai are not documented. Cowork runs its sessions on Claude Code, and plugin hooks load there, but the mods pages don’t list Cowork and I didn’t test it; on 1 October the desktop app still bundled Claude Code 2.1.286, one version before mods were turned on. A device’s managed policy reaches Cowork sessions on the user’s machine unless the organization runs them in a full VM sandbox. Two pages disagree on whether a person’s own ~/.claude files reach Cowork, so this article makes no claim either way. I also didn’t test whether CLAUDE.md files in parent folders load in a Cowork session; if they do, the reference implementation’s mounted tree would reach Cowork on the user’s machine today. On Pro and Max that window narrows on 6 October 2026, when new Cowork tasks move to the cloud.

• Motives are inferred. Anthropic hasn’t explained publicly why Claude chat and Cowork are flat.

• The code and the raw data are not published with this article. A reader can’t reproduce the measurements from the article alone; the video shows the app working, not how it is built.

• The reference implementation runs on an invented firm. It is not integrated with claude.ai or Cowork, and “acting as” is a demo switch, not a sign-in. Permissions are enforced by Claude Code’s tool lists, not by the app. Isolation turns off a person’s own hooks and mods for the run; mods built into Claude Code keep running, and the names of personal agents can still appear in context. The app depends on claude -p loading CLAUDE.md, which Anthropic says will stop being the default. The personal-file leak and the --setting-sources result were tested on Windows; the measurements and the mod test ran on Linux.

Sources

• Anthropic, Set organization instructions

• Anthropic, Roles and permissions

• Anthropic, What is the Team plan? · Plans and pricing

• Anthropic, Manage custom roles on Enterprise plans

• Anthropic, Organize your tasks with projects in Claude Cowork

• Anthropic, Use Google Workspace connectors

• Anthropic, Manage groups and group spend limits on Enterprise plans

• Anthropic, What are projects? (new version of projects, beta)

• Anthropic, Get started with Claude Cowork(global and folder instructions)

• Anthropic, How Claude remembers your project (CLAUDE.md) · All settings (claudeMdExcludes, disableAllHooks)

• Anthropic, Customize Claude Code with mods (1 October 2026) · Mods overview · Manage mods for your organization · React to events with a mod · Mods reference

• Anthropic, Claude Tag: What is Claude Tag? · Configure per-channel access · Customize Claude Tag · How agent identity works · Audit · Set a spend limit

• Anthropic, Projects in Claude Code (new projects beta) · Projects in Cowork · How Claude Code uses prompt caching · Run Claude Code programmatically (--bare) · Extend Claude Code · Manage project visibility and sharing · Use connectors · Authorize MCP connectors for your entire organization · Provision and manage skills

• Anthropic, Configure server-managed settings (no per-group configuration) · Manage plugins for your organization (plugin availability by group on Enterprise) · Plugin feature support across platforms (hooks ignored in chat, loaded in Cowork) · Deploy managed settings (Cowork runs its sessions on Claude Code)

• Anthropic, Pricing (Sonnet 5.5 rates, tokenizer note) · Prompt caching (caches isolated per organization and per workspace on the API)

• Anthropic, @anthropic-ai/tokenizer(public tokenizer used for the counts)

• Microsoft, Management groups · Landing-zone management group design

• AWS, SCP evaluation · Google Cloud, Hierarchy evaluation

• Microsoft, Group Policy processing · SharePoint fine-grained permissions

• FinOps Foundation, Token economics

• Jaroslawicz et al., How Many Instructions Can LLMs Follow at Once? (IFScale)

• Firefly, 2026 State of IaC research

• MCP, Authorization specification, version 2026-07-28

• GitHub, anthropics/claude-code issues #68262, #14467, #30554, #27567, #30250, #27302, #47741

• Beer, S. (1979), The Heart of Enterprise; Alexander, C. (1965), “A City is Not a Tree,” Architectural Forum; Simon, H. (1962), “The Architecture of Complexity,” Proc. Am. Phil. Soc.106(6)

• Reference implementation and data: the Worktree app, its tests, the measurement script, both result batches and every answer are held by the author and available on request.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية