Testing what models can do inside real systems
Agents, automation and developer tools, built to hold up outside a demo
Context
Models are now good enough to do real work: write code, follow a process, make small decisions. Most of what’s built with them still stops at the demo. The interesting part is what happens next, once real users, existing systems and messy constraints are involved.
What I did
I build and test agents, automation and new development workflows. The questions I keep returning to are practical ones: what an agent should remember, which decisions it may take on its own, and how people stay in control of systems that act for them.
Outcome
The findings become working tools and notes. The first, What I learned giving agents memory, covers what changed once agents could remember.