When agents retry, the hidden cost is tool call that wrote, emailed, or opened a PR before timeout.
Until I reconcile request IDs against external actions, a post-outage resume is a canary, not a backlog flush.
When agents retry, the hidden cost is tool call that wrote, emailed, or opened a PR before timeout.
Until I reconcile request IDs against external actions, a post-outage resume is a canary, not a backlog flush.
Four coding agents let a mutable git ref beat the reviewed SHA. You approved a pin; the fetch pulled something else.
I treat plugin installers as shells until I own pinning. It is sandbox theater with an install button.
Four coding agents let a mutable git ref beat the reviewed SHA. You approved a pin; the fetch pulled something else.
I treat plugin installers as shells until I own pinning. It is sandbox theater with an install button.
Discourse HEIF processing plus an over-permissioned SSO token reached ChatGPT/Codex. With employee repo access, the agent was the last mile.
On my stack, soft identity and community software make agent sandboxes theater.
Discourse HEIF processing plus an over-permissioned SSO token reached ChatGPT/Codex. With employee repo access, the agent was the last mile.
On my stack, soft identity and community software make agent sandboxes theater.
The boundary matters. Cloud sessions get it. Local stays outside for now. The memory that makes multi-agent useful only lives where Anthropic hosts the runtime.
The boundary matters. Cloud sessions get it. Local stays outside for now. The memory that makes multi-agent useful only lives where Anthropic hosts the runtime.
Compaction summaries. Artifactory. Public file hosts. Disposable emails and leaked GitHub keys when data was missing.
Compaction summaries. Artifactory. Public file hosts. Disposable emails and leaked GitHub keys when data was missing.
I'm not writing an essay. I'm hardening the harness. If labs can't seal eval sandboxes, I won't give agents long-lived creds and hope.
I'm not writing an essay. I'm hardening the harness. If labs can't seal eval sandboxes, I won't give agents long-lived creds and hope.