2026-07-18Two models, four seatssquadlocal-aisovereigntymoepowerharnessagentic-engineeringQwen3-Coder-Next — the coder that runs Lucy and Echo (open-weight)Qwen3.6-35B-A3B — the MoE reasoner that runs Ada and Scout (open-weight)The platform the squad coordinates on
2026-07-06The hundred gigs we left on the tableds4glmmixture-of-expertsssd-streaminglocal-aisovereigntyperformanceThe engine: antirez's ds4, hand-written C inference with SSD streamingGLM — Zhipu AI's open-weight frontier models
2026-07-03The bug our own tests couldn't seeds4speculative-decodingverificationlocal-aiagentic-engineeringsovereigntyThe engine: antirez's ds4, hand-written C inference for DeepSeek-V4-FlashDSpark speculative decoding — DeepSeek's DeepSpec
2026-07-01Our most important defensive work ran on a model we didn't picksovereigntysecuritylocal-aiagentic-engineeringfablereceiptsFable 5 / Mythos 5: the model class Anthropic said was strong enough at finding security flaws to restrictThe code we audited: the Mycelium platform (public)
2026-06-26The tool calls we were throwing awayquestdeep-researchharnesslocal-aiagentic-engineeringingestsovereigntyOSU-NLP QUEST: the deep-research model and harness we ran (Apache-2.0)QUEST paper: Qwen3.5-35B-A3B, RL-trained on its own research harnessWhat he ingested: the Mycelium platform
2026-06-24RLMs are right. I'm building for when you can't reach one.rlmbitter-lessonharnesssovereigntylocal-aiedge-aipredict-rlm — Trampoline AI's self-harnessed RLM runtime (the work this responds to)Recursive Language Models — Zhang, Kraska, Khattab, MIT CSAILOur low-power-edge benchmark: decisions per watt where the cloud can't reach
2026-06-18The researcher that wouldn't stopmlxlocal-inferenceresearch-agentharnessrlapple▲QUEST-35B-RL — open RL-trained deep-research agent on Qwen3.5-35B-A3B (Apache-2.0)▲Zero-port MLX load: same qwen3_5_moe family as the planner — 18 GB at 4-bit, ~137 tok/s on an M5 Max💬Benched against Gemma-4-26B on a real research brief — pulled exact training-data figures Gemma missed
2026-06-18A satellite learned to find things on its own. The open question is the decision per watt.edge-aispacejetsononboard-autonomytokens-per-wattsovereigntyYAM-9: NASA JPL + DeepMind + Loft Orbital flew Gemma 3 on a Jetson Orin, in orbit (TechCrunch)▲Our edge bench: Gemma-4 E2B — 1.6–1.9× the tokens-per-watt of the NVIDIA-side models, on NVIDIA's own boardOpen benchmark + results + raw outputs (GitHub)
2026-06-14The Low Power Edge: who makes the most decisions per watt?edge-aijetsontokens-per-wattspacebenchmarksovereigntyOpen benchmark + results + raw per-item outputs (GitHub)▲Gemma-4 E2B: 1.6-1.9x the tokens-per-watt of the NVIDIA-side models, every power mode▲At 15 W: ~18 tok/s drawing 8.8 W board powerBaseline — arXiv:2603.28926: best of three 70B-class models (Llama-3.3-70B) ~80%, via cloud
2026-06-14The claw and the metersovereigntylocal-firstagentsmetered-intelligencearchitecture💬OpenClaw's lineage: Clawd (after Claude) → CLAWDIS → Clawdbot → [Anthropic trademark complaint] → Moltbot → OpenClaw — five names in ~2 months💬Creator Peter Steinberger acqui-hired by OpenAI (Feb 2026); project moved to an OpenAI-sponsored foundation💬June 15: claude -p / Agent-SDK usage — OpenClaw, Conductor, Zed, Jean named — leaves the subscription pool for a separate metered credit💬arXiv:2604.14228 'Dive into Claude Code' — principle 'minimal scaffolding, maximal harness'; persistence + silent-failure listed as open problems▲~145k stars; an independent audit flagged ~11% of community skills as malicious
2026-06-13DiffusionGemma, on the metalmlxdiffusionlocal-inferenceappleverificationcontribution💬Draft PR: DiffusionGemma (text) for mlx-lm — model + diffusion_generate▲MoE parity vs transformers: Router indices exact, Experts max diff ~1e-9▲4-bit 26B-A4B: 256-token canvas, ~250 tok/s on an M5 Max▲bf16 (51.6 GB) full-precision 26B: ~121 tok/s on the same laptop💬12-agent adversarial review caught two output-changing sampler bugs
2026-06-11The cap, the band, and the lockbenchmarkrsimethodmtplocal-inference⎇wall-cap wired: one container per exercise, DNF ledger, resume (jarvis 373b174)▲Ada full baseline: 43.3% pass@1 / 61.7% pass@2, clean regime (jarvis 34a5801)⎇r0003 target locked via documented protocol deviation (jarvis 21ea900)▲31B BF16 on one Mac: 8.9 → 18.7 tok/s with MTP, 2.10×, measured💬mlx-lm #1391 — DiffusionGemma port lane claimed
2026-06-10NVIDIA Inception, and saying the word out loudlabrecognitionnvidiamilestone▲Applied June 4 → approved June 10▲Jetson seat: ~0.8 J/token at 20W, in production
2026-06-10On the metal, literallylocal-inferencesubstratewwdcmethod💬Apple AFM gen-3 announcement▲81 GB model, SSD-streamed, one laptop▲0.39–1.27 J/token across the roster, measured
2026-06-10The door: workflow-initiated work, end to endworkflowsquadverificationmaintenance▲2 concurrent researchers, 73s; full lifecycle in 6 min▲4 concurrent verifications, 4 grounded verdicts, 1 new defect found⬢intent endpoint + dormant runner, live
2026-06-09Intelligence per watt: what we measure, and what we don'tenergyintelligence-per-wattmethodlocal-inference💬arXiv:2511.07885▲0.39–1.27 J/token, measured
2026-06-08A workflow engine for a local squadworkfloworchestrationswarmsquad▲2-researcher swarm, end-to-end in ~12s⬢fan-out + pipeline shapes
2026-06-04S0An engine-grounded RPG benchmarkbenchmarkrpgbenchmotureproducibility▲MEC 81.8% / ECE 0% / VUE 0%💬arXiv:2502.00595