Graded, source-mirrored research corpora for agents. Evidence strong enough to rest decisions on.
Experiments in evidence-grounded machinery for long-horizon coding agents.
Spec-to-shipped pipeline for coding agents: intent capture, decision tickets, model-routed worker harness, artifact-verified completion.
Spec-compiled ETL: profile the input, record every decision in an auditable spec, compile to deterministic Python. Bad rows quarantined and error-coded, never silently coerced.
Measurement-first toolkit for pushing local LLM inference on Apple M-series GPUs toward the hardware's actual limits.
Run a 35B MoE on a MacBook in ~2 GB of RAM: experts streamed from SSD, near-roofline Metal kernels, and a KV cache that reloads from disk instead of re-prefilling.