Systems thinking is still the work
Generated code makes syntax cheaper. It does not remove a system's obligations.
That is the distinction I use when I review an AI-assisted change. I need to know what the change can cause, what it assumes, what evidence supports it, and how the people operating the system recover when the evidence is incomplete.
Start with the effect
Before I ask whether code runs, I ask what it can do. An API call may create a shipment, change a record, or send a message. A browser agent may submit a form. A mapper may turn an omitted field into a value that changes where work goes.
Each is a systems decision, even when the implementation is small. A review needs four clear answers:
- What external effect can this change create?
- Which facts must be true before that effect is allowed?
- What evidence supports the proposed action?
- What happens when a fact or approval is no longer valid?
Those questions fit inside a pull request. They require more than a syntax check.
Tests establish a boundary
The free AI-assisted code review kit uses an illustrative shipment mapper. Its original fixtures pass when countryCode is CA or US. A proposed default also passes those fixtures when the field is absent.
The result shows that the supplied fixtures still work. It does not show that an omitted country means Canada, or that the downstream system should receive a value the request did not provide.
Tests remain useful evidence. Their value depends on reading their boundary accurately. Someone responsible for the workflow still has to decide whether a missing fact should cause rejection, correction, or a review queue.
Approval needs an object
The same issue appears when an agent proposes an action. Approval for one revision is not evidence for a later revision.
The revision-bound approval example compares an action identifier, a proposed revision, and a captured state identifier. If any one changes, it rejects the approval. The implementation is intentionally narrow. It does not claim to be a complete production control.
Its contribution is visible behaviour. The workflow can say why it stopped instead of carrying an old approval into a changed proposal.
Put the boundary in the tool
I used the same principle in Metron, a small local coding agent. It limits file slices to 121 lines, listed files to 60, search matches to 10 per request, and model round-trips to 10 per turn. It proposes a patch, checks it with git apply --check, and waits for explicit approval before applying it.
Those limits are a specific design, not universal defaults. They let another engineer examine what the tool can reach, what it proposes, and where it stops.
The durable skill is being able to explain what a system is allowed to do, what information it relied on, and what should happen when that information changes. AI can generate alternatives quickly. It cannot choose which assumptions a team should accept.