Skip to content

Mental Mode Enterprise: 119 Lines of Anti-Bureaucracy Rules

Adityo Guni Waluyo

A commit added 119 lines of execution rules to AGENTS.md. I expected fresh bureaucracy; the opposite: capped interviews, scaled cards, capped iterations.

TL;DR

A 119-line update to the AGENTS file initially looked like heavy process but actually aims to cut friction. It uses lightweight work cards scaled to task size, caps fix attempts at three, and requires verifiable checks before marking done. Clear priorities and scope fences prevent guesswork, silent creep, and wasted context while letting you step away confidently.

There is a commit with zero lines of code in it: the AGENTS.md file in the my-apps repo grew by 119 lines, all of them execution rules for coding agents. It is called Mental Mode: Enterprise Full-Stack Engineer. When I first heard the word "enterprise", my gut said one thing: new bureaucracy. A trivial fix now needs a project card, an interview first, then a wait for review?

Once I read the whole thing, the direction turned out to be the opposite. This mode cuts friction instead of adding it: interviews are used only for genuinely ambiguous or high-risk decisions, work cards scale with task size down to a one-line plan for a trivial fix, fix iterations are capped at three, and the agent is told to stop the moment the Definition of Done is met. Every rule answers one real failure mode: fake completion without evidence, silent scope creep, guess-based fixes, or interviews that eat time over trivia.

Context Is Finite, Rules Must Be Sharp

Why not just write rules that are as complete as possible? Because agent context is expensive. On its best practices page, Anthropic explains that [2] nearly all of their best practices rest on one constraint: the context window fills up fast and performance degrades as it fills. A pile of instructions without priority drains the exact budget the task itself needs.

That is why the priority order in this mode is rigid: security and user data first, explicit user instructions above global rules, and when two rules collide, the conflict gets rewritten explicitly, not resolved by feel. This reset-how-you-work pattern is not new. [1] The Agile principles state that the best architectures emerge from self-organizing teams, and the last principle calls for regular reflection to tune the team's behavior. A retrospective that produces new rules, not new meetings.

Work Cards and Delegation Contracts

The most practical part is how work gets handed over. The form is a work card whose contents can be verified:

# kartu kerja minimal, tugas sepele cukup satu baris rencana
task: perbaiki validasi email di form kontak
goal: form menolak format invalid dengan pesan yang jelas
priority: security_first          # keamanan selalu urutan pertama
scope: vertical_slice             # backend + form + verifikasi sekaligus
constraints:
  - max_fix_iterations: 3         # berhenti saat DoD terpenuhi
  - scope_creep: kartu backlog baru, bukan dikerjakan senyap
verification:
  - pytest -k contact_form        # exit 0 = lulus
  - isi form manual di preview, cek pesan error
definition_of_done: verification lulus dan diff bisa direview

The verification section is the core of evidence-based verification. Anthropic puts it sharply: [2] give the agent a check it can run; it is the difference between a session you watch and one you walk away from. Without it, the human becomes the verification loop, and agent autonomy stops at the word "looks done".

Handing work to an executor agent uses a contract too, not hope:

# kontrak delegasi: dibaca agen pelaksana sebelum mulai
goal: tambah paginasi offset di endpoint daftar aplikasi
context:
  - apps/api/routes/apps.py (handler list_apps)
  - ikuti pola router yang sudah ada, jangan ubah skema database
acceptance_criteria:
  - tes paginasi lulus
  - halaman 2 tidak menduplikasi baris halaman 1
verification: pytest -k apps_pagination
expected_output: satu commit di branch staging, diff plus hasil verifikasi

The ecosystem already supports this pattern. opencode ships [3] specialized agents configured with their own prompts, models, and tool access, from the wide-open Build to the read-only Plan. Claude Code separates out subagents [4] that work in their own context and return only a summary, so a long research detour never floods the main conversation. The card is what decides which slice of work is worth throwing over the fence.

A Fence for Autonomy

My favorite failure mode is silent scope creep: an agent is asked to fix one handler, quietly restyles buttons in a neighboring file, and two hours later the diff cannot be reviewed because everything is mixed together. A card that closes the scope, the three-iteration cap, and the rule that scope creep becomes a new backlog card close that hole.

The fence works both ways. In this blog's pipeline, every article draft goes through a DETECT-mode style audit before it gets pushed: style findings get edited manually, while numbers, claims, and sources must not be touched. Evidence-based verification applies not just to agents, but to anyone claiming work is finished.

My experience runs one way: five minutes writing a sharp card and verification criteria saves two hours of undoing an agent's wrong guesses. Those 119 lines are not extra bureaucracy. They are the price of the ticket that lets me walk away from the keyboard without flinching, knowing exactly which card is running on which track.

Sources

Related articles