The Agent Wrote the Code—Now What?
If an AI agent writes most of your code and you feel like you’re learning less, you’re looking in the wrong place.
Agents now absorb the syntax errors, library quirks, and API docs. What’s left, and what an agent can’t judge for you, is the system around the code. Your job is knowing whether what the agent produced actually belongs there.
This is a quick reference. Each section gives you the why in one line, then the questions to ask and the things to do. Work through them on your own codebase, one at a time.
Table of Contents
- Part 1: Understand the problem
- Part 2: Understand the system
- Part 3: Build
- Part 4: Ship
- Part 5: Run
- Part 6: Protect
- Part 7: Sustain
- Before you merge agent code: the 60-second checklist
- None of this works without the fundamentals
- The bottom line
Part 1: Understand the problem
1. Know what you’re building and why
Why: Agents build exactly what you ask for. If you asked for the wrong thing, you get the wrong thing, fast.
- What problem does this solve, and for whom?
- How will we know it worked? Name a metric.
- What’s explicitly out of scope?
- What are the edge cases: empty input, huge input, bad input, no network, first-time user?
- Who else is affected: support, ops, other teams, users with disabilities, users in other languages or time zones?
Part 2: Understand the system
2. Trace one request end to end
Browser → DNS → load balancer → service → cache → database → response
DNS turns a name into an address. A load balancer spreads traffic. A cache keeps recent answers close. The database is the source of truth.
Why: Every slowdown, bug, and outage lives somewhere on this path.
- Open the browser Network tab, then follow the request through logs or traces.
- At each hop: Where does latency show up? Where can it break? Where does the data change?
- Which hops are synchronous (user waits) vs. asynchronous (queued for later)?
3. Find out where state lives
Why: The nastiest bugs are state bugs. The code looks right, but the data is stale, duplicated, or overwritten.
- Which service owns each piece of data? One owner only.
- Where is it cached, and for how long?
- What happens when the cache is stale? (Password changed, old one still works for five minutes.)
- What happens when two requests update the same record at once? (Two buyers, one ticket.)
- What’s the source of truth when two copies disagree?
4. Define the contracts
Why: APIs are promises. Break one and you break every caller, including ones you didn’t know existed.
- Is the request/response shape documented (e.g. OpenAPI, protobuf)?
- Is every change backward compatible? Add fields; don’t rename or remove them.
- Is there a versioning plan for breaking changes?
- Are errors consistent and useful: status code, machine-readable code, human message?
- Do list endpoints paginate?
- Is every write idempotent (safe to repeat)? Use idempotency keys for payments and orders.
5. Change data safely
Why: Code rolls back in seconds. Data doesn’t.
- Use expand → migrate → contract: add the new column, dual-write, backfill, switch reads, then drop the old one. Never rename in one step.
- Will the migration lock a big table? Test on production-sized data.
- Does every frequent query have an index?
- Is there a tested backup, and have you ever actually restored it?
Part 3: Build
6. Direct the agent well
Why: Output quality tracks input quality.
- Give context: the goal, the relevant files, the constraints, the conventions.
- Break work into small, reviewable steps.
- Ask for tests alongside the code.
- Ask the agent to list assumptions, trade-offs, and alternatives before it writes.
- When it’s wrong, say why it’s wrong, not just “try again.”
7. Review every line the agent writes
Why: You own the code. “The AI wrote it” isn’t an excuse in a postmortem.
- Do you understand every line? If not, ask or rewrite.
- Does it call APIs or libraries that actually exist, at the version you use?
- Does it swallow errors silently (empty
catch, ignored return values)? - Does it match the existing patterns in the codebase, or invent new ones?
- Did it change anything you didn’t ask it to?
- Is there dead code, debug output, or a hardcoded value?
8. Test the right things
Why: Tests are how you trust code you didn’t write.
- Unit tests for logic, integration tests for boundaries (DB, APIs), a few end-to-end tests for critical user flows.
- Test the failure path, not just the happy path.
- Would this test actually fail if the code were broken? Break the code and check.
- Fix or delete flaky tests. A test nobody trusts is worse than none.
9. Vet your dependencies
Why: Every package is code you run but didn’t review.
- Does the package actually exist and is it the real one? Agents sometimes invent package names, and attackers register them (“slopsquatting”).
- Is it maintained? Check last release, open issues, and number of maintainers.
- Pin versions with a lockfile.
- Scan for known vulnerabilities (Dependabot,
npm audit,pip-audit). - Is the license compatible with your project?
Part 4: Ship
10. Trace one deployment
Commit → CI → tests → build → artifact → rollout → health checks → rollback
CI tests every change automatically. The artifact is the packaged build that ships. Rollout releases it. Health checks confirm it’s alive. Rollback is the undo button.
Why: “It deployed” is not an answer. How it deployed is.
- Follow your last merged PR from commit to production.
- Is the rollout gradual (canary, percentage-based) or all at once?
- If this release were bad, how would you know, and how fast could you undo it?
- Has anyone actually tested the rollback?
11. Separate config, secrets, and code
Why: Hardcoded config means redeploying to change a setting. Hardcoded secrets mean a breach.
- Config lives in environment variables or a config service, not in code.
- Secrets live in a secrets manager, never in git. Scan history, since deleted secrets stay in commits.
- Use feature flags to ship code turned off, then turn it on gradually. Remove flags once they’re done.
Part 5: Run
12. Make it observable
Why: You can’t fix what you can’t see.
- Logs tell you what happened. Make them structured (JSON) and include a request ID.
- Metrics tell you how much. Track the four golden signals: latency, traffic, errors, saturation.
- Traces tell you where. Follow one request across services.
- Set an SLO (e.g. “99.9% of requests under 300 ms”) and alert when you’re burning through it.
- Alert on user-facing symptoms, not every CPU spike. Every alert should require action.
13. Study the failure modes
Why: The happy path is easy. Production engineering starts when it breaks. Agents default to the happy path.
- For every dependency: what if it’s slow, down, or returning garbage?
- Every network call needs a timeout.
- Retry with exponential backoff and jitter, and only retry idempotent operations.
- Use circuit breakers to stop calling something that’s failing.
- Degrade gracefully: show cached data or a partial page instead of an error.
- Read your team’s postmortems. They’re the best free education you’ll get.
14. Measure performance before optimizing
Why: Intuition about what’s slow is usually wrong.
- Profile first. Optimize the slowest thing, not the most obvious one.
- Watch p95/p99 latency, not averages. Averages hide your unhappiest users.
- Look for N+1 queries (one query per item in a loop). Agents write these often.
- Load test before launch. Know your breaking point.
- Know which parts scale horizontally (add machines) and which don’t (a single database).
15. Understand cost
Why: Small choices compound fast.
- A simple-looking query can scan terabytes. Run
EXPLAIN. - One extra model call can double inference cost.
- A retry loop can become a traffic multiplier: one failure × ten retries × ten thousand users = a self-inflicted outage.
- Check the cloud cost dashboard. Know what your service costs per month and per request.
- For every loop the agent writes, ask: what’s the worst-case number of calls?
16. Be ready for incidents
Why: Outages are inevitable. Chaos during them is optional.
- Mitigate first, investigate second. Roll back, then find the root cause.
- Keep runbooks: step-by-step fixes for known problems.
- Know who’s on call, how to reach them, and how to escalate.
- Write blameless postmortems: what happened, why, and what changes so it can’t recur.
Part 6: Protect
17. Threat model your own service
Why: An agent will write code that works. It won’t always write code that’s safe.
Sketch the service, draw the data flows, and walk through STRIDE-LM:
| Threat | Ask yourself |
|---|---|
| Spoofing | Can someone pretend to be another user? |
| Tampering | Can data be altered in transit or in storage? |
| Repudiation | Can someone deny an action because nothing was logged? |
| Information disclosure | Can data leak to people who shouldn’t see it? |
| Denial of service | Can someone knock the service offline? |
| Elevation of privilege | Can a regular user gain admin powers? |
| Lateral Movement | If one piece is breached, can an attacker hop to the rest? |
18. Apply security basics everywhere
Why: Most breaches exploit well-known mistakes, not clever zero-days.
- Know the OWASP Top 10.
- Validate all input on the server. Never trust the client.
- Use parameterized queries. Never build SQL from strings.
- Authentication (who are you?) is not authorization (what can you do?). Check authorization on every request, for every object.
- Grant least privilege to users, services, and API keys.
- Encrypt in transit (TLS) and at rest.
19. Respect user data
Why: Privacy failures cost trust, money, and sometimes legal trouble.
- Know which fields are PII (personally identifiable information).
- Keep PII out of logs, error messages, and analytics.
- Collect only what you need. Define how long you keep it.
- Can you fully delete a user’s data on request, including backups and caches?
- Know which regulations apply to you (GDPR, CCPA, HIPAA, PCI).
Part 7: Sustain
20. Keep it maintainable
Why: Code is read far more often than it’s written, and agents can write it faster than you can read it.
- Clear names beat clever code.
- Small functions, small files, one responsibility each.
- Delete dead code. Git remembers it.
- Keep a visible list of tech debt, and pay some down every sprint.
21. Write down decisions
Why: Six months from now, nobody will remember why. Including you.
- README: what it is, how to run it, how to test it, how to deploy it.
- ADRs (architecture decision records): one page per big decision, covering the context, the options, and the choice.
- Comments explain why, not what.
22. Work well with humans
Why: Software is a team sport, even when agents write the code.
- Keep PRs small and focused. Write a description that explains the why.
- Write clear commit messages.
- In reviews, critique the code, not the person. Ask questions instead of issuing orders.
- Ask for help early. Being stuck for a day costs more than a five-minute question.
Before you merge agent code: the 60-second checklist
- I understand every line.
- It solves the actual problem, and nothing else changed.
- Tests cover the happy path and the failure path.
- Every network call has a timeout. Every retry has a limit.
- No secrets, PII, or debug output.
- Input is validated. Authorization is checked.
- Dependencies are real, pinned, and vetted.
- Schema changes are backward compatible.
- It’s observable: logs, metrics, errors surfaced.
- I know how to roll it back.
None of this works without the fundamentals
You can’t reason about a stale cache if you don’t know what a cache does. You can’t spot a runaway retry loop if you’ve never thought about how networks fail. Agents didn’t make fundamentals optional. They made them the job.
The bottom line
Your job is shifting from writing code to directing the agents that write it.
And you can’t direct what you don’t understand.