A reading of workflow distillation and why task success, constraint adherence, trajectory quality, and deployment cost belong in separate reported metrics.
Tailscale, tmux, mosh, and a mobile terminal let me steer long agent runs from my phone. Why that plain stack still outperforms Remote Control for me, for now.
Karpathy's LLM wiki: point a coding agent at a markdown folder and let it maintain the knowledge base. What it is, why it earns its place, and how to set it up.
Trillions committed, Nvidia the most valuable company ever, and the measured productivity gains are tiny. Notes on using the overbuild and surviving the crash.
Karpathy's autoresearch runs an AI agent in a keep-or-discard loop. I examined its techniques, then applied the loop to parameter search where a grid stalls.
An honest Codex CLI vs Claude Code comparison from running both daily on one codebase: strengths, config interop, and the two-model workflow that emerged.