
I need to solve on-prem AI before I have to mortgage the house.
Why this note matters
This note is funny because it is also a strategy memo. The more agentic work moves from occasional chat into always-on infrastructure, the more local inference stops being a hobby and starts being cost control.
It connects directly to the site’s model-training and SLM thread: small models, NPUs, local harnesses, and practical experiments that make AI cheaper to run without making the workflow smaller.