Agentic systems are easy to demonstrate and harder to review. A demo shows that a model can plan, call tools, and finish a task. It does not show who allowed each action, what the system will remember tomorrow, what happens when a tool fails halfway through, or how anyone would know the result was actually correct.

Those are not exotic questions. They are the ordinary questions an architecture review asks of any system that changes state on someone's behalf. What is different with agents is that the system chooses some of its own actions at runtime, so boundaries that a conventional design review assumes are fixed in code can move.

This article sets out a practical review method for agentic systems: eight dimensions to examine, the evidence to ask for in each, and the red flags that should slow a team down before it grants more autonomy. It is a working Decision Crafters synthesis, not an industry standard.

Where this fits

Two earlier Decision Crafters articles cover adjacent questions. Does This Business Actually Need an AI Operating System? asks whether a problem needs shared governed infrastructure at all, or whether process correction, integration work, or one bounded AI workflow is enough. Agent Harnesses Are a Layer, Not the Whole AI System compares current harness designs and argues that the runtime loop should be treated as a replaceable layer rather than the whole system.

This article assumes an agentic system already exists or is concretely proposed. Its question is narrower and more operational: is this system's architecture adequate for the consequences of what it is allowed to do?

Scope the review by consequence first

A review that treats every agent the same will either overburden a low-risk assistant or under-examine a high-risk one. Start by writing down four things:

The answers set the bar for everything that follows. An agent that drafts internal summaries for a person to edit does not need the same authority model as an agent that issues refunds or changes infrastructure. The point is not to make every system heavyweight. It is to make the controls match the consequences.

The eight review dimensions

Each dimension below lists what to ask, what evidence to request, and a red flag. Evidence matters more than answers. A confident verbal answer to "who can approve this action?" is weaker than a configuration, a log, or a test that shows it.

1. Authority

Authority is the question of who or what may take each action, on whose behalf, and under which policy.

Evidence to request: the permission model, the identity the agent runs as, and a log showing a denied action.

Red flag: the model effectively grants itself capability by choosing a tool, with no independent check between intent and execution.

Public work points the same way. NIST's National Cybersecurity Center of Excellence has framed its agent project around identification, authorization, auditing, and non-repudiation, and the Model Context Protocol defines an OAuth 2.1–based authorization mechanism for HTTP-based servers. MCP Gives Agents Tools. Who Gives Them Authority? explores why tool access and authority are different things.

2. Memory and state

Evidence to request: a description of each store and its owner, and a restart or recovery test.

Red flag: important decisions live only in a conversation transcript or a context window. A transcript records what was said; it is not an accepted decision.

3. Tools

Evidence to request: the tool inventory with scopes, the validation rules, and a retry test.

Red flag: broad tools such as "run any query" or "send any email" left in place because they were convenient during the prototype.

The OWASP Top 10 for Agentic Applications lists tool misuse and exploitation, and identity and privilege abuse, among its top risks. Least-privilege tool design is usually one of the simplest mitigations to apply.

4. Data

Evidence to request: a data-flow description, an entitlement mapping, and an injection test using realistic untrusted content.

Red flag: the agent reads content from outside its trust boundary and acts on it with no separation between data and instructions. The first entry in OWASP's agentic list, agent goal hijack, describes exactly this path.

5. Runtime boundaries

Evidence to request: sandbox and network configuration, credential handling, and budget and termination settings.

Red flag: the agent runs with a developer's personal credentials or a shared service account that can do far more than the task needs.

This is where the distinction between harness, sandbox, and session discussed in the harness article is most useful: each can fail or be replaced independently only if the boundaries between them are real.

6. Observability

Evidence to request: a sample trace of a real run, the retention settings, and the named owner of trace review.

Red flag: the only evidence of what the agent did is its own final summary.

Tooling is improving here. The OpenAI Agents SDK for Python, for example, records model generations, tool calls, handoffs, and guardrails in its built-in tracing, which is enabled by default. Collection is not the same as review, so the owner question still matters.

7. Failure modes

Evidence to request: failure tests, evaluation criteria based on the resulting state rather than the transcript, and the rollback procedure.

Red flag: success is measured by whether the agent reported success.

Anthropic's guidance on agent evaluations draws a useful line between the transcript, which records what the agent and its tools did, and the outcome, which is the resulting state of the environment. Review the outcome.

8. Human control

Evidence to request: the approval points in the workflow, a stop-and-revoke test, and examples of what reviewers actually see.

Red flag: human oversight exists on paper, but approvals are so frequent or so thin that reviewers click through them.

The NIST AI Risk Management Framework asks organizations to define human roles and oversight processes for AI systems. Where more AI stops helping covers how to find the point where human judgment should stay explicit.

What the review should return

A review that ends in a general impression is not much use. Return four things:

  1. A disposition for each of the eight dimensions: adequate for the current consequence level, gap, or unknown.
  2. The findings ranked by consequence, not by how easy they are to fix.
  3. The conditions under which the agent may gain more autonomy, and the evidence that would show those conditions are met.
  4. Stop conditions: the observations that should pause or roll back the system.

"Unknown" is a legitimate answer. A review that marks a dimension unknown and names the evidence needed is more useful than one that guesses.

Common review mistakes

Applying the method to our own agents

Decision Crafters runs agents for its own internal work, including one that publishes approved articles to this site. Several of its controls map directly onto these dimensions:

Our own operations reinforced the same lesson: reconcile tracked state against the real environment before treating the tracker as authoritative.

This is internal operating evidence from one small organization's own systems. It is not client evidence, and it does not show how the method performs elsewhere.

What this method is and is not

Sources

Close

Agentic systems will keep getting more capable. The review question stays the same: is the architecture adequate for the consequences of what the system is allowed to do?

Start with consequence. Examine the eight dimensions with evidence rather than assurances. Mark unknowns honestly. Expand autonomy only when the evidence supports it.

If you are working through these questions, subscribe and follow the research. Decision Crafters is publishing practical methods for reviewing AI systems before they earn more autonomy.