How this guide started
This guide grew out of a practical review of permissions, approval mechanisms, and risks in a coding-agent workflow. I wanted to understand which restrictions were enforced by the environment, which depended on an instruction, and what evidence would support a claim that a control worked.
With AI assistance, I turned the material into a general guide, removing company-specific details and internal operational information. The downloadable document is in Hebrew. This post explains the approach behind it and the questions I would bring to another team's review.
Start with what the agent can access
A coding agent may work through several interfaces: a terminal, a repository host, a database client, and connected services. Reviewing the permissions of one tool does not establish the boundaries of the whole session. The useful question is what the agent can actually reach through every interface available to it.
I start by defining the workspace, the required tools, and the data needed for the task. A code-editing task should not inherit production credentials simply because they are convenient to keep in the same environment. When read access is enough, the identity used for that task should not also be able to update or delete data.
Three layers worth checking separately
The first layer is the execution environment: file access, network access, available tools, and command permissions. The second is the change process: tests, review, and controls on publishing or merging code. The third is identity: which account each tool uses and the permissions that account holds in each environment.
These layers answer different questions. A code reviewer can find a risky change without preventing a direct database command. A read-only database identity can protect records without preventing an unauthorized deployment. A useful review tests each boundary independently and then checks the paths between them.
| Layer | Question to test |
|---|---|
| Execution environment | Can the agent access files, tools, or network destinations outside the task? |
| Change process | Can a change be published when a required check fails or is unavailable? |
| Identity and data | Can the same operation be attempted through a more privileged identity or another tool? |
Make approval specific enough to mean something
Human approval is most useful when the person can see the action, target environment, affected resources, and scope of the change. An approval to prepare a migration is different from permission to run it. An approval for one environment should not silently carry over to another.
If the content or target changes materially, the approval needs another look. After execution, verify what actually happened. A request being accepted is not the same as the intended result being present, and a timeout does not establish that nothing happened. Read back the state before repeating an uncertain operation.
Test the boundary, including the inconvenient paths
A configuration file is evidence of intent. A test shows what happened under specific conditions. Keep both, and describe the difference. Use a test environment with synthetic data to exercise permitted operations and operations that should be refused.
Include alternative paths: a direct command instead of the usual wrapper, a different identity, an unavailable review service, an expired approval, or a partially completed operation. Where a check is mandatory, its failure or absence should stop the action that depends on it.
Record what was tested, which environment and source state were involved, what the result demonstrated, and what remains unverified. One successful test should not become a claim that the entire system is secure. Controls also need another look when tools, permissions, or deployment settings change.
Use the guide as a team conversation
The guide brings together permission boundaries, human approval, common risk scenarios, and a verification checklist. It is a starting point for deciding which actions an agent may perform independently, which need review, and which credentials it should never receive for the task at hand.
The most useful conversation is concrete: who can approve a deployment, what happens when the reviewer is unavailable, how temporary access expires, and who owns a gap until it is resolved. Agree on those decisions before a long-running agent has to encounter them.
If your team uses coding agents, I would be interested to hear how you test these boundaries and what you would add to the checklist.