Agent systems · July 2026 – present

A local agent system that works the nights I'm away.

One Windows PC, one 16 GB GPU. When I lock the machine or step away, it loads a local model, researches topics on the public web in read-only mode, drafts and builds code with tests in a sandbox, and makes some art. Nothing it produces counts until I review it.

commits, Jul 30 – Sep 29
545
tests in the suite
~1,900
research packs so far
888
packs in one 9.8-hour night, zero crashes
79
Fail closed

A web fetcher that assumes the worst

DNS is resolved before connecting, and private, loopback, and cloud-metadata addresses are refused. The live connection is re-checked to defeat DNS rebinding. Content types are allowlisted, and per-host pacing is shared across processes.

Sandboxed

Model-written code runs with no network

Patches are tested inside a Windows AppContainer created through ctypes, with no network access. A passing build re-runs the full suite against a cached baseline, and a failure must repeat before it closes the build.

Human gate

Only my verdict closes work

Every output lands in quarantine. One function writes verdicts, to a SHA-256 hash-chained log, and it re-checks the machine's own automatic rejects. I turned down an auto-merge feature on purpose.

Measuring it

  • The judge, on a holdout. A claim-support judge checks whether each research claim is backed by its quoted source. On a 200-pair holdout it scores 0.70–0.735 balanced accuracy across the local models I tested. The labels are built automatically from linked claim-quote pairs; a human-labelled set is still to do.
  • Models, A/B tested. Model and quantization choices are compared head to head. The current model decodes at 34 tokens a second on the 16 GB card, with a 0.55-second first token when cached.
  • Nightly scorecard. The best night so far: 85% of claims linked to a source, 66% judge-supported.

What's live

  • Research runs nightly. It is the lane that carries the numbers above.
  • Stable after a hardware hunt. Seven host crashes in 15 days across four stop codes traced to one corrupted physical memory page. Memory timing, voltage, and a bad-page exclusion fixed it, and the next full night ran 9.8 hours clean. One clean night is evidence, not proof.
  • Build is lightly used. About twenty sandboxed builds so far, each needing my accept.
  • Art lanes are mostly paused. They share the GPU and restart one at a time.
  • One user. Me. It is personal infrastructure in a private repo, and code walkthroughs are available on request.

Want the walkthrough?

The repo is private. Ask, and I'll walk you through the code.