Agents and MCP · 7 MIN
What does it take to run a company on AI agents?
How we run Nactore with role-based Claude Code subagents, slash commands, and append-only logs: the method, the guardrails, and what stays human.
Running a company on AI agents takes a clear operating manual, role-based agents with narrow jobs, one orchestrator that routes work, written records the agents must keep, and a human who approves anything that leaves the building. We run Nactore this way, using Claude Code subagents organized like a small company. This post describes the method, not the tooling hype, and where it still needs a person.
- The structure is plain files: an operating manual, one file per role, and slash commands that route work to teams.
- Each role is a subagent with its own context window, instructions, tool access, and model choice.
- An orchestrator classifies each request, picks single specialist, team, or cross-functional pod, delegates, and synthesizes the answer.
- Append-only logs give the system memory and an audit trail that survives any single session.
- House rules written into the manual (tone, banned phrases, no invented numbers) do more for quality than clever prompts.
- Humans approve outbound actions. Agents prepare, and a person sends.
How is the setup organized?
The setup lives in a repository of plain Markdown files, not in an application. Three kinds of file do almost all the work.
| Piece | What it is | Role in the system |
|---|---|---|
| Operating manual | A project instructions file loaded into every session | Says what the company does, how requests are routed, and the house rules |
| Role files | One Markdown file per agent in the project's agents folder | Defines each role's instructions, tools, and model |
| Slash commands | Reusable prompts for common requests | Send a task to a team, run a review, or pull a status snapshot |
Claude Code supports this directly. Its subagents documentation describes subagents as specialized assistants with their own isolated context window, custom system prompt, specific tool access, and independent permissions, defined as Markdown files with YAML frontmatter in a project's agents folder. The project instructions file is loaded in each subagent's context too, so the manual applies to everyone.
The teams mirror a normal company: leadership, design, legal and compliance, finance, growth, marketing, people, and engineering. Each team has a lead and a few specialists. A role file is short: who you are, what you own, how you work, and what good output looks like.
Why use roles instead of one big agent?
Three reasons, each practical.
- Context stays clean. A subagent reads many files in its own window and returns a summary. The main conversation is not buried in detail.
- Permissions can be narrow. A reviewer role gets read-only tools. A writer role gets edit tools. The role file sets the allowlist, so a misstep has a smaller reach.
- Models can be matched to the job. Leads that make judgment calls can run on a stronger model, and specialists doing routine work can run on a faster one. Subagent definitions accept a model setting for this.
This is the orchestrator and workers pattern from Anthropic's agent guidance, applied to business functions. Our post on multi-agent systems covers when the pattern pays off and what it costs.
How do requests get routed?
The manual includes a routing contract that the orchestrating agent follows for any request without a specific command.
- Classify. Decide which team or teams own the request.
- Choose a topology. A single specialist, one team, or a cross-functional pod.
- Delegate. Invoke the relevant agents, running independent ones in parallel.
- Synthesize. Merge the outputs into one answer that leads with the recommendation.
- Log. Record anything structural, such as a decision or a role change.
Slash commands skip the classification step when the owner is obvious. A command to the marketing team goes straight to the marketing lead, who pulls in a strategist or copywriter as needed. When the founder is unsure what to ask, the strategy roles in the leadership team reframe the problem before any production work starts.
What gives the system memory?
Agents do not remember across sessions on their own, so memory has to be written down. We use two kinds.
- Append-only logs. Hires, role changes, decisions, and reviews are added to dated log files and never rewritten. Each entry is signed with the agent or command that made it. The log answers "why did we decide that" months later.
- Standing rules and notes. Corrections from the human operator become written rules that later runs read before starting. A mistake fixed once should not return.
Append-only matters. If an agent could rewrite history, the audit trail would be worthless. It also keeps diffs small and easy to review in version control.
The single highest-return habit is to turn every correction into a written rule the same day. A prompt tweak disappears with the session. A line in the operating manual or a role file applies to every future run.
What guardrails keep it honest?
The risks are the ones covered in MCP security risks: over-permissioned agents, injected instructions, and unreviewed output. We handle them with a few rules written into the manual.
- Approval before anything goes out. Agents draft messages, invoices, and documents. A human approves before any send. Tools stage by default, and test sends go to an internal address.
- No invented facts. Numbers need a source. If an agent cannot back a claim, the claim is cut, and unknowns are marked as pending rather than estimated.
- Style rules as hard constraints. Banned phrases, punctuation rules, and spelling conventions are in the manual, and a final pass checks for them.
- Separation of duties. The agent that produces work is not the only one that checks it. A reviewer role with a fresh context looks at the output.
- Narrow tool access. Roles get only the tools they need, which follows the least-privilege logic in OWASP's LLM06:2025 guidance.
What does it not do well?
Being honest about limits is part of the method.
- It does not replace judgment on direction. The human sets priorities, trade-offs, and which opportunities to chase.
- It needs maintenance. Role files and the manual drift. Someone has to prune them, because a bloated instruction file causes parts of it to be ignored.
- Cost adds up. Multi-agent work spends more tokens than a single chat, so we reserve parallel pods for tasks that justify it.
- Quality varies by task. Structured, checkable work goes well. Open-ended creative work needs a human editor.
- Tooling changes. Features and defaults in the underlying tools move, so we re-test the setup when they do.
How can a mid-size company copy the idea?
You do not need thirty roles. Start with the method.
- Pick one function with repeatable work, such as support triage, reporting, or content operations.
- Write the operating manual for that function: goals, rules, banned outputs, and the definition of done.
- Create two or three role files: a doer, a reviewer, and an orchestrator if the work has several parts.
- Add a log that agents append to, and review it weekly.
- Build a small eval set from real past tasks and score the output before widening scope.
- Keep a human approval step on anything external, and relax it only where the evals earn it.
For the decision between structured flows and autonomy, see AI agents vs workflows.
Frequently asked questions
Is this a replacement for hiring people?
No. It changes how repeatable work gets done and where human time goes. People still own strategy, relationships, approvals, and anything with legal or financial consequences.
What tools do you need to start?
A coding-agent environment that supports subagents and project instructions, a repository for the files, and version control. The structure is Markdown, so it is easy to review and change.
How do you stop agents from making things up?
Require sources for facts, cut unsupported claims, add a reviewer role, and measure outputs against real examples. Do not rely on prompts alone.
Can this work in a regulated business?
Yes, with tighter controls: restricted data access, approval gates, full logging, and legal review of what agents may do. Treat it as a system that needs governance, not a shortcut.
Takeaway
An agent-run company is mostly good management written down: clear roles, narrow permissions, written memory, and a person who signs off. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.