Why a tiny model raises the eval stakes
Building a golden dataset
Deterministic checks for structured output
LLM-as-judge (cloud judge vs on-device judge)
Regression across Chrome versions and hardware
The latency/quality tradeoff
Wiring evals into CI
Next: Shipping & the compatibility matrix.