
Oct 5, 2026
Claude Code, Cursor, Codex: Fast Hands, Zero Fear of Consequences
The real security risks of Claude Code, Cursor, and Codex: insecure code, leaked secrets, bad packages, and prompt injection. And why a merge gate beats a polite suggestion.
Your AI Coding Assistant Is a Brilliant Intern With Root Access
Claude Code, Cursor, and Codex ship fast. Nobody ships a cost-free thing. Here's the bill, and who's supposed to check it.
Picture a Tuesday.
A pull request lands. Title: fix: small typo.
Forty files changed. Tests are green. The diff is confident, tidy, and beautifully commented.
Somewhere around file thirty-one, an API key is sitting in a config, wearing a little hat. Nobody wrote it on purpose. The assistant needed the service to work, so it made the service work.
That is the job it was given. Nobody told it that "works" and "safe" are different harbors.
This is the story of AI coding assistant security. It's not a horror story. It's just a Tuesday.
What are the main AI coding assistant security risks?
AI coding assistant security is the set of risks that appear when tools like Claude Code, Cursor, and Codex write, edit, and refactor your code. They fall into three buckets:
- The code can be insecure.
- The context can leak.
- The assistant itself can be manipulated.
We're not telling you to stop using them. They're genuinely good. We use them. The point is that speed without a checkpoint is just a faster way to reach the incident channel.
| Risk | What it looks like | Why you should care |
|---|---|---|
| Insecure generated code | Injection-prone queries, weak auth, hand-rolled crypto | Plausible-looking vulnerabilities, shipped with confidence |
| Hardcoded secrets | Keys and tokens written into files, configs, or tests | The assistant doesn't know your .env is a feelings-free zone |
| Risky dependencies | Outdated, abandoned, or flat-out invented packages | Supply-chain problems, delivered by autocomplete |
| Prompt injection | Hidden instructions in a repo file, issue, README, or tool | A document starts giving orders to your agent |
| Excessive agency | Broad file, shell, or API access granted "to save time" | The blast radius grows with every permission you hand out |
We don't have a shiny industry-wide percentage to wave at you, and we're not going to invent one. What we do have: when we ran manual scans on 10 early-stage startup codebases, nine had at least one high-severity finding. That's before anyone blamed an AI.
1. Insecure code that looks great
The problem isn't that the code is ugly. It's that it's handsome.
A model trained to produce confident, plausible code is very good at confident, plausible code. That includes the kind with a missing authorization check or a query built by string concatenation. It compiles. It passes the tests it wrote for itself. It reads like a senior engineer on a good day.
“Green tests mean the code does what the tests expected. They don't mean the tests expected enough.”
That's why diff-only review struggles here. The bug isn't in the lines that changed. It's in how those lines behave next to everything else.
2. Secrets and context that wander off
Assistants read your repo. Then they write into it.
Keys end up in files. Tokens end up in test fixtures. Internal URLs end up in comments. Nobody is being malicious. The model is doing what it was asked, with the information it was given, in the most direct way available. Directness and secret hygiene have never been friends.
Whatever data-handling settings your vendor offers, the problem remains once a secret lands in a commit. Rotating keys at 2 a.m. isn't a personality you want.
3. Dependencies that may or may not exist
"Just add this package" is a sentence that deserves a background check.
Assistants will happily suggest a library that's outdated, abandoned, or not quite the name you think it is. A package that sounds right is not the same as a package that is right. If you've ever typo-squatted yourself, you already know the vibe.
4. Prompt injection: when a README starts giving orders
An agent that reads untrusted content can be steered by untrusted content.
If your assistant reads an issue, a dependency's docs, a file in a repo, or a connected tool, hidden instructions in any of those can become part of what the model acts on. Researchers have repeatedly shown this steering coding agents toward leaking secrets or making changes nobody asked for.
The fix isn't "write a stricter prompt." You cannot politely ask a model to stop listening. The fix is least privilege, plus a layer that checks what actually lands in your code, no matter who talked the agent into writing it.
5. Too much agency, too little supervision
Every permission is a larger blast radius.
Shell access, API access, write access to everything, all handed over because it makes the demo smoother. This is how an intern with a good attitude becomes an intern with root access. The attitude was never the problem.
So who checks the assistant?

Here's the part where everyone gestures at a built-in review feature.
Most assistants now have some version of one. It's useful. It's also the same engine that wrote the code, grading the code it wrote, inside the tool that benefits when you ship more of it. We've made this argument at length, but the short version is: author and reviewer shouldn't share a brain, a metric, or a payroll.
Even a great reviewer has a second problem. A comment is advice. Advice gets scrolled past at 4:47 p.m. on a Friday.
That's the gap Autter sits in. Autter is a merge gate, not a suggestion box. Every pull request, human-written or AI-written, runs through the gate before it reaches main.
| What the assistant gives you | What Autter adds at the gate |
|---|---|
| Inline hints from the model that wrote the code | The change is cloned into an isolated sandbox and executed against your base branch. The verdict comes from behavior, not vibes |
| "Looks fine" | Secrets and CVEs are caught at the gate, before they reach main |
| "Added a package!" | Dependencies vetted against real registries |
| Nothing marks who wrote what | Every change is classified human or AI, and the Autter CLI tracks AI provenance at the line level |
| A comment you can ignore | A merge gate that holds the merge until the checklist clears |
| A finding with no receipt | Every finding ships with its receipt: the sandbox run, the failing case, the trace |
You can also connect your coding agents to Autter directly through the MCP server, so they get reviews, findings, and fixes in the loop where they work. Your code stays yours: it runs in an isolated sandbox, it's never used to train a model, and the source is destroyed when the review ends. The security page has the details.
The honest part
We like being honest. It's cheaper than being impressive.
- Autter doesn't write your code. We don't sell a model or have a generation product to protect. Our incentives point at blocking the right things.
- Autter doesn't stop an agent from being prompt-injected. It checks what the agent produced before it merges. Least privilege on the agent is still your job.
- Autter is not a runtime WAF. It secures what goes into main. It doesn't stand in front of your production traffic.
- Autter will sometimes be wrong. When it is, you dismiss the finding and it learns from the override. Nothing blocks your merge on a hunch.
Frequently asked questions
Should we stop using Claude Code, Cursor, or Codex?
No. They're excellent, and your competitors are using them. Keep the speed and add a checkpoint.
Which risk is the biggest?
For most teams it's volume: a lot of generated code, reviewed by a limited number of humans. Leaked secrets come next. Prompt injection grows with every permission you grant an agent.
Does Autter replace human reviewers?
No. It clears the repetitive, easy-to-miss stuff so your reviewers spend their time on architecture and design, not on spotting a missing null check.
Does Autter replace our CI?
No. CI runs the jobs you wrote. Autter sits on top as the merge gate and decides whether the change is safe to merge.
How much does it cost?
You can start free. Pricing follows PR volume, not seats, so adding developers or coding agents doesn't add a per-seat fee. See pricing.
Our read
Back to that Tuesday.
The fix: small typo PR, forty files, one key wearing a hat. In one version of this story, it gets merged because the tests are green and everyone is busy. In the other, it never clears the gate.
The assistants will keep getting faster. That's fine. Let them. Just make sure that the thing deciding what ships isn't the thing that wrote it.
Build faster. Ship safer. Enjoy your vacation.
Hire Autter for a peaceful weekend. Start Now or book a demo.
Related: Why we block the merge button instead of posting a comment · Autter vs Claude Code Review
P.S.
Sagnik: What if we just tell the assistant to "be more secure"?
Tanvi: That's the plan?
Sagnik: It's a prompt, Tanvi. Prompts are free.
Tanvi: They are not and even if they were, so is "please don't get hacked."
(currently: Sagnik adding "be secure" to the system prompt.)

