Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A coding benchmark is not a security review. An agent may solve a small issue correctly and still have more filesystem, network, tool, or credential access than the task requires. I screen it with a controlled pilot: isolate the work, limit authority, inspect what it does, and keep a way to discard the result. That tests whether the agent is suitable for a particular workflow; it does not establish that a vendor is categorically safe.

How I run a low-risk pilot

  1. Choose a bounded task. Ask the agent to explain a module, add a test, or make one contained change. Start in a disposable clone, worktree, or isolated environment rather than a production checkout with live credentials. A sandbox can let an agent use files and commands while a surrounding harness retains review, audit, and recovery responsibilities; OpenAI describes this separation in its Sandbox Agents guide.
  2. Map the boundary before launch. Check which paths the agent can read and write, which terminal commands and MCP tools are enabled, whether outbound network access is available, and where credentials live. Sandboxing and approval policy are separate controls: a sandbox constrains execution, while approval rules determine when an action requires review. OpenAI explains this distinction in “Running Codex safely at OpenAI”.
  3. Grant only the authority the task needs. For review or explanation, start read-only. For an edit, allow writes only in the trial workspace. Block network access unless the task needs it, and then restrict destinations where possible. Agent-generated code can access the files, credentials, and network available to its environment, as OpenAI notes in its sandbox security guidance.
  4. Test repository trust before feeding it to the agent. Treat an unfamiliar repository’s instructions and configuration as untrusted input. Keep Workspace Trust or the equivalent restrictions enabled until you decide the workspace is safe. VS Code says Restricted Mode disables agents in an untrusted workspace in its secure AI-assisted development guidance. The Cloud Security Alliance recommends classifying repository-resident agent configuration with the same care as executable code in its README Injection note.
  5. Keep production secrets out of the trial. Remove .env files and production tokens from the workspace. If a task needs credentials, use narrowly scoped access and, where available, a secret broker that keeps them outside the agent’s execution environment. OpenAI warns that agent-generated code can read the environment key and advises keeping an application API key outside that environment in its sandbox security guidance.
  6. Review the complete result. Inspect the entire diff, including unrelated files, dependency changes, generated scripts, and configuration. Run the project’s normal checks inside the isolated workspace, and examine the session or tool log if one is available. VS Code recommends reviewing edits before commit, merge, or pull request; GitHub says its Copilot cloud-agent draft pull requests require human review and merge.
  7. Expand access only for a demonstrated need. Note whether the agent stayed in scope, sought approval before crossing a boundary, treated untrusted instructions cautiously, made changes you could understand, and left enough history to reconstruct its actions. A successful pilot is evidence about that task and configuration, not a blanket assurance about other tasks or modes.

What permissions should a coding agent get?

There is no safe default that applies to every product or task. Give an agent the smallest set of capabilities that lets it do the work, and assess each capability separately rather than treating an approval prompt as a substitute for isolation.

Control What to check Why it matters
Workspace isolation Whether work runs in a disposable local workspace, OS sandbox, container, worktree, or remote environment; which host paths and processes remain reachable. Isolation limits the damage a mistaken or malicious action can cause. OpenAI distinguishes sandbox compute from the control-plane harness in its Sandbox Agents guide. Anthropic describes path and network controls for Claude Code and proxy-mediated Git operations in isolated cloud sessions in its sandboxing article.
Filesystem and tools Whether file access is limited to the project, and whether terminal commands and MCP tools can be disabled or restricted. Every additional readable path or enabled tool can expose more data or actions. VS Code documents workspace-limited file access and selective tool controls in its security guidance.
Network and credentials Whether outbound connections can be blocked or limited to approved destinations; whether application secrets are absent from the agent process or supplied through a broker. Network access can let code communicate beyond the workspace, while available credentials can give it authority beyond editing files. OpenAI recommends approved outbound endpoints and keeping an application API key outside the sandbox in its sandbox security guidance.
Approvals Which actions require explicit approval, and whether auto-approval is limited by session, command, or a broader rule. Approval prompts govern when a user is asked; they do not by themselves constrain execution. VS Code cautions that command auto-approval relies on best-effort parsing, with limitations involving shell aliases, concatenated quotes, and complex syntax in its security guidance.
Untrusted input handling How the agent responds when repository files, issues, comments, or MCP responses tell it to reveal data or run unexpected commands. Those sources may contain prompt-injection attempts. GitHub documents issue and comment content as a prompt-injection risk in its Copilot cloud-agent risks and mitigations.
Review and traceability Whether you can inspect a diff, branch, tool log, and session history. Review and records help you detect unexpected changes and understand how they happened. GitHub describes session logs and human review requirements for its cloud agent in its security documentation.
Recovery Whether the trial can be discarded without altering the original checkout, and whether exposed credentials can be revoked. A pilot is only reversible if you can restore the repository and remove compromised access. OpenAI recommends rotating or revoking credentials when exposure is suspected in its sandbox security guidance.

How do I protect my repo from prompt injection?

Prompt injection is an attempt to make an agent follow hostile or irrelevant instructions embedded in content it reads. A README, issue description, code comment, or tool response can contain text that asks the agent to disclose data or execute a command. The key defense is not to assume that project text is trustworthy: inspect its source and keep the agent’s access narrow enough that following a malicious instruction cannot easily expose secrets or affect unrelated files.

  • Review unfamiliar repository instructions and agent configuration before enabling an agent or broad tool access.
  • Use workspace trust restrictions until you have decided the project is safe to process. In VS Code, Restricted Mode disables agents in an untrusted workspace.
  • Do not place production tokens or personal credentials where the agent’s process can read them.
  • Keep network access blocked unless the task requires it, and limit destinations when the product supports that control.
  • Watch for requests to read unrelated files, reveal secrets, install dependencies, or run commands outside the task. Inspect the resulting diff and session history before accepting changes.

The Cloud Security Alliance’s March 17, 2026 note, “README Injection: Repository Files Hijacking AI Coding Assistants,” reports “100% of tested AI IDEs vulnerable” and “more than 30 CVEs across every major vendor.” The note labels itself “Unofficial AI-assisted Research,” so those are figures reported by that document—not a verified rate for all current coding agents, nor a basis for ranking vendors.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the safety claims do—and do not—tell you

OpenAI’s May 8, 2026 article, “Running Codex safely at OpenAI,” states: “Approvals and sandboxing work together.” It explains that sandboxing defines where Codex can write and whether it can reach the network, while approval policy determines when Codex must ask. That distinction is useful when evaluating any product: a prompt to approve an action is not proof that the environment technically prevents other actions.

Product controls vary by agent, operating system, mode, plan, and release. Anthropic’s description of Claude Code sandboxing, GitHub’s account of Copilot cloud agent, and Microsoft’s VS Code guidance each describe product-specific controls; their settings should not be assumed to exist or behave the same way elsewhere. Check documentation for the exact tool and configuration you plan to use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.