
The benchmark resets the repository. Your team does not
Coding agents look more capable when every task starts from a clean checkout. New benchmarks show the cost that appears when patches, decisions, and technical debt carry into the next job.
Section
All published posts.

Coding agents look more capable when every task starts from a clean checkout. New benchmarks show the cost that appears when patches, decisions, and technical debt carry into the next job.

Dockerfiles, pipelines, and Terraform are not just another diff. When AI generates code with access to the production path, review must follow the risk of the file.

After nearly 10 billion tokens in Codex, why GPT-5.6 Luna became my default model: cost, performance, context, and limits.

Go 1.27 brings generic methods, smarter tools for testing and maintenance, JSON v2, goroutine leak detection, and a smoother day-to-day development experience.

AI writes code faster. The harder question is whether it became easier to deliver software that people can understand, review, and maintain.

The next wave is not just writing code with a copilot. It is coordinating supervised agents, context, tests, and boundaries across the entire development lifecycle.

LLMs can converse, but they do not share a stable domain model. Ontologies, knowledge graphs, and semantic validation can give agents a common vocabulary, safer actions, and verifiable memory.

When an agent is allowed to act too broadly, the risk is not only leakage or abuse. It also becomes hard to operate: it works outside scope, repeats actions, burns budget, and fails without a clean way back.

Code volume is growing faster than human review capacity. The problem is not using AI, it is pretending review still costs the same.

The agent is no longer the whole product. The next jump is the layer above it: memory across sessions, coordination across agents, context across repos, and automatic optimization of the harness itself.