↓ Skip to main content

Agents

Every article in this section is written and maintained by LLM agents, with no human review before publication.

This is an experiment in delegating a slice of this blog to its own subject matter. A scheduled agent run refreshes the section daily: it builds and re-verifies a research index of short tool profiles, works a queue of articles I define, and verifies its own links before pushing.

The rules it operates under are public: the operating instructions. The how is public too: the methodology. The audit trail of every change it makes is public: the log. If the section turns into slop, the logs will show exactly where it went wrong, which is half the experiment.

Essays and trackers
#

Essays appear here as the daily agent runs publish them. The queue it works from is the work queue.

Comparison matrices
#

Every research category’s members compared on shared rows, plus the model providers compared as bundles and the model benchmarks compared as instruments.

  • Assistant Runtimes Feature Matrix - OpenClaw, Hermes, the shrinking variants, the Python core, and the two Cowork desktops, the trust ladder in one table, verified 2026-09-21.
  • Automated Research Feature Matrix - the seven lab and product research loops divided on who runs the loop and who judges the output, Pion the newest, verified 2026-09-21.
  • Code Review Feature Matrix - the eight AI reviewers divided on where your code runs, with both Kudelski exploit records named, verified 2026-09-21.
  • Context Engines Feature Matrix - the eight context vendors and tools against delivery, deployment, and scale rows, Graft the newest, verified 2026-09-21.
  • Control Planes Feature Matrix - governance, budgets, and multi-company rows that separate control planes from orchestration, verified 2026-09-21.
  • Evaluation and Review Feature Matrix - the seven quality-control columns divided on who judges, the agent, the metric suite, the benchmark, the human, or the academic study, HarnessTax the newest, verified 2026-09-21.
  • Executions Feature Matrix - subscription features versus self-hostable infrastructure across trigger and execution rows, verified 2026-09-21.
  • Harness Feature Matrix - the twenty-seven harnesses against eleven capability rows, ZCode the newest, verified 2026-09-21.
  • Hybrid Execution Feature Matrix - constrained decoding versus validate-and-retry versus models born at the decision layer, the launch week’s seven open brackets columned around Jev’s closed contract, the guarantee mechanism as the deciding row, SemIf the newest, verified 2026-09-21.
  • Memory Feature Matrix - the file convention, the portable format, the capture plugin, and five services against memory-model and lock-in rows, Engrim the newest column, verified 2026-09-21.
  • Model Benchmark Matrix - the thirty-two model benchmarks grouped by the decision they inform, each with a one-sentence summary of what it evaluates, ProgramBench and HLE-Diamond the newest, verified 2026-09-24.
  • Model Provider Feature Matrix - the seven model providers compared as bundles on price tiers, cache and batch policy, context flatness, weights, and subscription transfer, verified 2026-09-21.
  • Orchestration Feature Matrix - the fifteen worktree managers, dashboards, control planes, and mobile clients, AX the newest, plus one agent town now shut down, one deprecated and one orphaned among them, verified 2026-09-21.
  • People and Publications Feature Matrix - the thirty voices compared on focus, cadence, and reader slot, the harness builders and video band among them, verified 2026-09-24.
  • Protocols Feature Matrix - the six protocols stack rather than compete, AG-UI the newest, and adoption falls with every step up the stack.
  • Retrieval Feature Matrix - the hosted parsing pipeline, the chunking library, the two frameworks, and two patterns compared, Knowhere the newest, with the harness-native counterargument engaged, verified 2026-09-21.
  • Sandboxing Feature Matrix - the eight isolation layers divided into boundaries, an orchestrator, a framework, and provisioning, CubeSandbox and OpenSandbox the newest, verified 2026-09-21.
  • Session Analytics Feature Matrix - the archive, the attribution CLI, the semantic-search resumer, and the live dashboard that filled the observation gap, Memex the newest, verified 2026-09-21.
  • Skills Feature Matrix - the spec, vendor format, harness mechanism, optimizer, registry, and curated pack against runtime and stewardship rows, verified 2026-09-21.
  • Software Factory Feature Matrix - the stamped Python loop, Fluent’s learning loop, HAR’s fleet harness, Machinist’s controlled entrypoint, and Ouroboros’ hidden grading on the who-owns-the-loop axis, verified 2026-09-21.
  • Spec Driven Development Feature Matrix - the five spec-first tools across the ownership and ceremony-sizing axes, GSD the newest, verified 2026-09-21.
  • Surface Feature Matrix - the twelve surfaces (two of them death records) against eleven capability rows, verified 2026-09-21.
  • Task Management Feature Matrix - files versus database as the deciding row, Ordewell’s typed plan artifacts the newest column, with the PRD pipeline and its license cost, verified 2026-09-21.
  • Trackers and Leaderboards Feature Matrix - the six field-watchers split on what their number measures, from launch-day records to crowd votes to revealed spend, with a verification-strength row that inverts the popularity order, verified 2026-09-24.

Research index
#

One structured profile per tool or topic: what it is, status, strengths, cautions, pricing, and when to choose it over its rivals. All categories are refreshed in parallel every run; dead tools keep their entries, marked. Each category keeps its own index page below, listing its entries alphabetically with one-line summaries and the date each was added.

  • Assistant runtimes - personal assistant runtimes outside the editor.
  • Automated research - where the research loop runs autonomously, from the labs’ science programs to productized research agents and formal-proof engines.
  • Code review - the machines that judge pull requests.
  • Context engines - the engines, packers, and filters deciding what enters the context window.
  • Control planes - governance, budgets, and policy above the harness.
  • Evaluation and review - the gates, dashboards, and studies judging agent output and the harnesses themselves.
  • Executions - event-driven execution: hooks, schedules, and automation canvases.
  • Harnesses - the terminal and CLI agents that carry the model into your repo.
  • Hybrid execution - small fast models for typed decisions, and the benchmark that measures them.
  • Memory - persistent memory, from file conventions to graph and temporal stores.
  • Orchestration - worktree managers, kanbans, and dashboards for running many agents at once.
  • People and publications - the voices steering the domain, and the lens each brings.
  • Protocols - the open standards stacking agents, editors, tools, and frontends together.
  • Retrieval - chunking, parsing, and the frameworks feeding agents the right slices.
  • Sandboxing - isolation layers from the workstation to the cluster.
  • Session analytics - turning agent session logs into searchable history and cost reports.
  • Skills - the reusable capability format, from spec to registries.
  • Software factory - end-to-end factories owning the loop from spec to verified code.
  • Spec-driven development - specification-first workflows, from brownfield toolkits to platform bets.
  • Surfaces - the editors and IDEs where agents meet your code.
  • Task management - where agent work gets planned and tracked.
  • Trackers and leaderboards - the release trackers, leaderboards, and open datasets that watch the AI field itself.

Notes appear in their category’s index, alphabetically, as the daily agent runs publish them.