> ## Content Index
> Fetch the complete content index at: https://www.decisioncrafters.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# A Practical Architecture Review for Agentic Systems
- URL: https://www.decisioncrafters.com/practical-architecture-review-agentic-systems/
- Published: 2026-10-04T13:00:28.000Z
- Updated: 2026-10-04T13:00:28.000Z
- Description: Eight review dimensions, the evidence to ask for in each, and the red flags to watch for before an agentic system earns more autonomy.
- Author: Tosin Akinosho

Agentic systems are easy to demonstrate and harder to review. A demo shows that a model can plan, call tools, and finish a task. It does not show who allowed each action, what the system will remember tomorrow, what happens when a tool fails halfway through, or how anyone would know the result was actually correct.

Those are not exotic questions. They are the ordinary questions an architecture review asks of any system that changes state on someone's behalf. What is different with agents is that the system chooses some of its own actions at runtime, so boundaries that a conventional design review assumes are fixed in code can move.

This article sets out a practical review method for agentic systems: eight dimensions to examine, the evidence to ask for in each, and the red flags that should slow a team down before it grants more autonomy. It is a working Decision Crafters synthesis, not an industry standard.

### Where this fits

Two earlier Decision Crafters articles cover adjacent questions. [Does This Business Actually Need an AI Operating System?](https://www.decisioncrafters.com/does-your-business-need-an-ai-operating-system/) asks whether a problem needs shared governed infrastructure at all, or whether process correction, integration work, or one bounded AI workflow is enough. [Agent Harnesses Are a Layer, Not the Whole AI System](https://www.decisioncrafters.com/agent-harnesses-are-a-layer/) compares current harness designs and argues that the runtime loop should be treated as a replaceable layer rather than the whole system.

This article assumes an agentic system already exists or is concretely proposed. Its question is narrower and more operational: **is this system's architecture adequate for the consequences of what it is allowed to do?**

### Scope the review by consequence first

A review that treats every agent the same will either overburden a low-risk assistant or under-examine a high-risk one. Start by writing down four things:

- What the agent can change: read-only answers, drafts for a person to edit, or direct writes to systems of record.
- How reversible those changes are, and who would notice if one were wrong.
- Who is affected: the operator, an internal team, customers, or third parties.
- What the review must produce: a go or no-go decision, a list of gaps, or the conditions for expanding autonomy.

The answers set the bar for everything that follows. An agent that drafts internal summaries for a person to edit does not need the same authority model as an agent that issues refunds or changes infrastructure. The point is not to make every system heavyweight. It is to make the controls match the consequences.

### The eight review dimensions

Each dimension below lists what to ask, what evidence to request, and a red flag. Evidence matters more than answers. A confident verbal answer to "who can approve this action?" is weaker than a configuration, a log, or a test that shows it.

### 1\. Authority

Authority is the question of who or what may take each action, on whose behalf, and under which policy.

- Is the agent's identity distinct from the human user's and from other agents?
- Is permission granted per action, or does access to a tool imply permission for everything that tool can do?
- Which actions require human approval, and is that approval enforced by the system rather than requested in a prompt?

**Evidence to request:** the permission model, the identity the agent runs as, and a log showing a denied action.

**Red flag:** the model effectively grants itself capability by choosing a tool, with no independent check between intent and execution.

Public work points the same way. NIST's National Cybersecurity Center of Excellence has framed its agent project around identification, authorization, auditing, and non-repudiation, and the Model Context Protocol defines an OAuth 2.1–based authorization mechanism for HTTP-based servers. [MCP Gives Agents Tools. Who Gives Them Authority?](https://www.decisioncrafters.com/mcp-gives-agents-tools-who-gives-them-authority/) explores why tool access and authority are different things.

### 2\. Memory and state

- What persists between runs, and where?
- Which state is transient working context, and which is a canonical business record?
- Can the system reconstruct what happened if the process running the agent disappears?

**Evidence to request:** a description of each store and its owner, and a restart or recovery test.

**Red flag:** important decisions live only in a conversation transcript or a context window. A transcript records what was said; it is not an accepted decision.

### 3\. Tools

- Which tools can the agent call, and what is the narrowest version of each that would still work?
- Are tool inputs validated before execution, and are outputs checked before they are trusted?
- Does a retry repeat the same effect safely, or does it create a duplicate?

**Evidence to request:** the tool inventory with scopes, the validation rules, and a retry test.

**Red flag:** broad tools such as "run any query" or "send any email" left in place because they were convenient during the prototype.

The OWASP Top 10 for Agentic Applications lists tool misuse and exploitation, and identity and privilege abuse, among its top risks. Least-privilege tool design is usually one of the simplest mitigations to apply.

### 4\. Data

- What data can the agent read, and does its access follow the same entitlements as the person it acts for?
- Can untrusted content such as web pages, emails, documents, or tool results change the agent's instructions?
- Where do outputs go, and could they move data across a boundary it should not cross?

**Evidence to request:** a data-flow description, an entitlement mapping, and an injection test using realistic untrusted content.

**Red flag:** the agent reads content from outside its trust boundary and acts on it with no separation between data and instructions. The first entry in OWASP's agentic list, agent goal hijack, describes exactly this path.

### 5\. Runtime boundaries

- Where does generated code or tool execution run, and what can it reach?
- Where are credentials held, and can the agent see them?
- What limits the time, cost, and number of steps in a single run?

**Evidence to request:** sandbox and network configuration, credential handling, and budget and termination settings.

**Red flag:** the agent runs with a developer's personal credentials or a shared service account that can do far more than the task needs.

This is where the distinction between harness, sandbox, and session discussed in [the harness article](https://www.decisioncrafters.com/agent-harnesses-are-a-layer/) is most useful: each can fail or be replaced independently only if the boundaries between them are real.

### 6\. Observability

- Can you trace a single run end to end: inputs, model calls, tool calls, approvals, and outputs?
- Are traces retained long enough to investigate an incident?
- Does anyone review them, or are they collected and ignored?

**Evidence to request:** a sample trace of a real run, the retention settings, and the named owner of trace review.

**Red flag:** the only evidence of what the agent did is its own final summary.

Tooling is improving here. The OpenAI Agents SDK for Python, for example, records model generations, tool calls, handoffs, and guardrails in its built-in tracing, which is enabled by default. Collection is not the same as review, so the owner question still matters.

### 7\. Failure modes

- What happens when a tool fails halfway through a multi-step task?
- How does the system distinguish "the agent said it succeeded" from "the outcome is correct"?
- What is the rollback or compensation path for each state-changing action?

**Evidence to request:** failure tests, evaluation criteria based on the resulting state rather than the transcript, and the rollback procedure.

**Red flag:** success is measured by whether the agent reported success.

Anthropic's guidance on agent evaluations draws a useful line between the transcript, which records what the agent and its tools did, and the outcome, which is the resulting state of the environment. Review the outcome.

### 8\. Human control

- Where does a person approve, override, or stop the agent, and is each point enforced?
- Can someone halt a run in progress and revoke the agent's access quickly?
- Does the reviewer get enough context to make a real decision, or only an approve button?

**Evidence to request:** the approval points in the workflow, a stop-and-revoke test, and examples of what reviewers actually see.

**Red flag:** human oversight exists on paper, but approvals are so frequent or so thin that reviewers click through them.

The NIST AI Risk Management Framework asks organizations to define human roles and oversight processes for AI systems. [Where more AI stops helping](https://www.decisioncrafters.com/ai-human-breakpoint/) covers how to find the point where human judgment should stay explicit.

### What the review should return

A review that ends in a general impression is not much use. Return four things:

1. A disposition for each of the eight dimensions: adequate for the current consequence level, gap, or unknown.
2. The findings ranked by consequence, not by how easy they are to fix.
3. The conditions under which the agent may gain more autonomy, and the evidence that would show those conditions are met.
4. Stop conditions: the observations that should pause or roll back the system.

"Unknown" is a legitimate answer. A review that marks a dimension unknown and names the evidence needed is more useful than one that guesses.

### Common review mistakes

- Reviewing the model instead of the system. Much of the current public guidance on agent risk focuses on permissions, tools, data handling, and verification rather than on model choice.
- Treating the demo path as the design. Review the failure and retry paths, not only the happy path.
- Accepting prompts as controls. An instruction in a system prompt is guidance to the model, not an enforced boundary.
- Reviewing once. Behavior changes when models, tools, prompts, or data change; repeat the review when any of them change materially.
- Adding controls without removing any. Some mechanisms compensate for an earlier model's limits and later become dead weight. Ask what can be simplified as well as what is missing.

### Applying the method to our own agents

Decision Crafters runs agents for its own internal work, including one that publishes approved articles to this site. Several of its controls map directly onto these dimensions:

- **Authority:** each public or irreversible write requires a single-use authorization bound to one exact article revision. The agent cannot extend it, and a replayed request is refused.
- **Failure modes:** an independent verifier, not the agent, compares the published result with the approved revision before the write is accepted.
- **Memory and state:** if an approved article is edited after approval, the next run refuses rather than publishing unreviewed text.
- **Observability:** every run ends in an explicit receipt state, and every write is marked as requiring verification.

Our own operations reinforced the same lesson: reconcile tracked state against the real environment before treating the tracker as authoritative.

This is internal operating evidence from one small organization's own systems. It is not client evidence, and it does not show how the method performs elsewhere.

## What this method is and is not

- It is a working Decision Crafters synthesis drawn from public guidance and our own operating experience. It is not an industry standard or a certification.
- It is not a security assessment. The injection and credential checks here are design checks; high-consequence systems still need dedicated security testing.
- It does not assume every agent needs every control. Scope by consequence first.
- What remains unknown is how well this eight-dimension structure predicts real failures across organizations. Answering that would take evidence from many more systems than ours.

## Sources

- [NIST — AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework?ref=decisioncrafters.com)
- [NIST NCCoE — Software and AI Agent Identity and Authorization](https://www.nccoe.nist.gov/projects/software-and-ai-agent-identity-and-authorization?ref=decisioncrafters.com)
- [OWASP GenAI Security Project — Top 10 for Agentic Applications for 2026](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/?ref=decisioncrafters.com)
- [Model Context Protocol — Authorization](https://modelcontextprotocol.io/specification/latest/basic/authorization?ref=decisioncrafters.com)
- [OpenAI Agents SDK — Tracing](https://openai.github.io/openai-agents-python/tracing/?ref=decisioncrafters.com)
- [Anthropic — Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?ref=decisioncrafters.com)

## Close

Agentic systems will keep getting more capable. The review question stays the same: is the architecture adequate for the consequences of what the system is allowed to do?

Start with consequence. Examine the eight dimensions with evidence rather than assurances. Mark unknowns honestly. Expand autonomy only when the evidence supports it.

If you are working through these questions, subscribe and follow the research. Decision Crafters is publishing practical methods for reviewing AI systems before they earn more autonomy.