A reading of workflow distillation and why task success, constraint adherence, trajectory quality, and deployment cost belong in separate reported metrics.
Tailscale, tmux, mosh, and a mobile terminal let me steer long agent runs from my phone. Why that plain stack still outperforms Remote Control for me, for now.
Trillions committed, Nvidia the most valuable company ever, and the measured productivity gains are tiny. Notes on using the overbuild and surviving the crash.
Karpathy's autoresearch runs an AI agent in a keep-or-discard loop. I examined its techniques, then applied the loop to parameter search where a grid stalls.
Fable 5 returned at double the price of Opus 4.8, with days-long autonomy as the headline capability. A comparison with Claude Code's /goal loop, sourced.