bergen, norwayvol. i · no. 22 · July 21, 2026rss feed

Hasan Arief

A lab notebook on agentic coding, open-weight models, and what they cost to run

Section

Open-Weight LLMs

Running open-weight models locally and on rented GPUs: quality, quantization, and serving stacks.

Note

Why agent distillation needs separate success, constraint, and cost metrics

A reading of workflow distillation and why task success, constraint adherence, trajectory quality, and deployment cost belong in separate reported metrics.

Guide

Open-weight LLMs for coding in 2026: hardware and real costs

Open-weight coding LLMs as of July 2026, GLM-5.2, Kimi K2.7, DeepSeek-V4, with the VRAM math, GPU rental prices, and when self-hosting beats an API.