
The benchmark resets the repository. Your team does not
Coding agents look more capable when every task starts from a clean checkout. New benchmarks show the cost that appears when patches, decisions, and technical debt carry into the next job.
Developer Experience @ PicPay
Personal notes on software engineering, developer experience, Go, PHP, infrastructure, and AI.

Coding agents look more capable when every task starts from a clean checkout. New benchmarks show the cost that appears when patches, decisions, and technical debt carry into the next job.

Dockerfiles, pipelines, and Terraform are not just another diff. When AI generates code with access to the production path, review must follow the risk of the file.

After nearly 10 billion tokens in Codex, why GPT-5.6 Luna became my default model: cost, performance, context, and limits.

Go 1.27 brings generic methods, smarter tools for testing and maintenance, JSON v2, goroutine leak detection, and a smoother day-to-day development experience.

AI writes code faster. The harder question is whether it became easier to deliver software that people can understand, review, and maintain.

The next wave is not just writing code with a copilot. It is coordinating supervised agents, context, tests, and boundaries across the entire development lifecycle.

LLMs can converse, but they do not share a stable domain model. Ontologies, knowledge graphs, and semantic validation can give agents a common vocabulary, safer actions, and verifiable memory.

When an agent is allowed to act too broadly, the risk is not only leakage or abuse. It also becomes hard to operate: it works outside scope, repeats actions, burns budget, and fails without a clean way back.