New research: Portable Computer is a local-first agent for private and cost-effective work. With an on-device 27B model, our harness scores 82.6% on real knowledge work, beating open-source harnesses Pi and Hermes. Our post-trained PPLX 27B reaches 85.4%.

The model and harness are designed together because small models fail in harnesses built for frontier models. Portable is a minimal system prompt, skills that load on demand, connectors as compact CLI tools instead of MCP servers, self-verification, and an always-on sandbox.

Web research: on 1,266 BrowseComp tasks, Computer reaches 66.7% accuracy vs 50.2% for Pi and 43.9% for Hermes on their respective search providers. Portable uses the least wall time and the fewest tokens. Inference and private documents stay local; only search touches the web.

Parsing runs entirely on-device, so sensitive documents never leave the machine. On ParseBench-100, Computer scores 65.1% vs 34.6% for Hermes and 13.9% for Pi, in least time with fewest tokens.

Shows that a 27B local model beats larger open harnesses when the is designed for it, and quantifies what user-gated escalation to a frontier model buys: 59.6% to 73.0% on Terminal Bench 2.1 at $0.415 per rollout.
Checking sign-in…
Loading comments…