Skip to main content

Agent harness

An agent harness is the runtime layer around an AI model. It decides what context the model sees, which tools it can call, how permissions work, where memory lives, and how the agent recovers from mistakes. Do not treat the harness as plumbing. For coding agents and workflow agents, the harness often determines reliability more than the raw model choice.

The useful mental model

A good agent system externalizes capabilities that should not live only inside model weights. This model comes from the paper Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering. The important engineering question is:
Which capability should stay implicit in the model, and which capability should become an explicit external component?
If the answer affects reliability, auditability, reuse, or cost, it usually belongs outside the model.

What the harness owns

The harness should make these decisions explicit:
  • which files, pages, tools, and memories enter context
  • when old context is summarized, dropped, or preserved
  • how tool calls are validated before execution
  • which actions require user approval
  • how tool errors are represented and retried
  • how model effort, latency, and cost trade off
  • how runs are logged for later debugging
A thin harness can work, but an implicit harness becomes hard to debug. If behavior changes and nobody knows whether the cause is the model, prompt, context cache, tool layer, or product default, the harness is under-instrumented.

Design rules

Keep tools small and inspectable

Prefer narrow tools with clear inputs and outputs. A tool that does one thing is easier for the model to call correctly and easier for humans to audit. For local workflows, CLI tools are often enough. For reusable cross-client integration, use a protocol boundary such as MCP. For stable product backends, a direct API is still the simplest option. MCP got much better in 2026 — it is now stateless, supports deferred tool loading, and models are stronger at structured tool calling. That does not mean MCP is better than CLI for most integrations. Prefer CLI when:
  • the model already knows the command (gh, aws, kubectl) so you spend almost no schema tokens
  • you are on a local machine or in CI with a fixed command
  • you need pipes, composition, or a one-off debug loop
  • token cost matters and --help plus a skill is enough to teach a new CLI
Prefer MCP when:
  • there is no terminal (browser, mobile, hosted agent)
  • you need OAuth, per-user auth, or standardized audit headers
  • the tool is new to the model and a machine-checkable JSON schema is worth the extra tokens
  • the tool set is dynamic and agents should discover tools at runtime (tools/list + deferred loading)
A useful rule: CLI wins on known local tools and composition. MCP wins on auth, audit, discovery, and environments without a shell.

Separate skill design from execution

A skill is a manual for the agent — it should only contain what the model does not already know. The execution is done by tools (scripts, CLI, apps) that the model calls directly. Keep the two apart:
  • Design logic lives in the SKILL.md and describes what to do and why.
  • Execution logic lives in scripts and tools and describes how to do it.
Scripts can be pre-written or generated on the fly by the agent when a task needs one. Generated scripts are flexible — agents already do this even without a skill, often writing a one-off Python snippet. They are also slower, more token-hungry, and less complete on edge cases. Pre-written scripts win when the task is recurring and quality matters: a Markdown-to-HTML skill that generates the converter at runtime will be barely usable, while a pre-written converter can cover styles and edge cases. Treat generated scripts as throwaway until verified. This separation explains why strong skills still need a UI: the skill tells the agent what to accomplish, while a UI carries the concrete execution steps a human can trigger with a few clicks instead of re-prompting from scratch each time. Skills did not kill MCP, and RAG is not dead. They solve different problems and often sit in the same workflow:
  • MCP is a shared interface: structured tool calls, auth, and data access.
  • Skills are packaged expertise: how this team works, how this project should change, which conventions matter. Markdown is a feature — humans can read it too.
  • RAG / retrieval grounds the model in documents, tickets, and code that are not in weights.
MCP provides access. Skills explain how to use that access well. Retrieval starts the model closer to the answer. Pick the layer that matches the gap, not a winner.

Preserve reasoning-critical context

Context pruning is not just a cost optimization. It changes behavior. If an agent made tool calls or file edits based on earlier reasoning, the harness must preserve enough rationale for later turns to continue coherently. Otherwise the agent may repeat itself, forget why it chose a path, or pick odd tools.

Represent durable state as files

Long-running agent workflows need state outside chat history. Plain files are often the simplest durable memory layer because humans can inspect, edit, diff, and version them. A useful pattern is a workspace that separates intent, resources, work products, and learning records:
The exact names can change. The important rule is that state should be inspectable, restartable, and scoped. For agent-assisted learning, this lets the agent teach the next step from the learner’s current capability instead of from a generic syllabus. For engineering work, the same pattern applies to onboarding notes, migration logs, incident follow-ups, and task journals.

Audit memory for the associativity blind spot

Every current memory architecture — Flat RAG, graph memory, agentic hierarchies, OS-style memory kernels — retrieves with the same recipe: embed query and stored content into one similarity space (BM25 or dense vectors), take top-K. That caps recall at surface similarity. It covers only descriptive recall — the query shares wording, entities, or timestamps with the memory. The other half is associative recall: a month-old note that “someone on the team has a severe seafood allergy” is the decisive evidence for today’s question “where should the team dinner be?” Zero lexical overlap; only a semantic arc (causality, same situation) connects them. Tencent’s T-Mem paper (2026) measures the blind spot: on LoCoMo-Plus Cognitive questions — built to strip lexical overlap — similarity-based systems score 32–49%, and GPT-4o reading the entire transcript scores only 21%. T-Mem’s fix borrows from cognitive science (episodic future thinking): at write time, compute and store a prediction of when this memory will matter again — a trigger. Four trigger families cover the 2×2 space of granularity (fact / scene) × orientation (descriptive / associative). The two associative families are the point:
  • Bridge trigger — project a fact into the concrete situation where knowing it will help: “seafood allergy → picking a restaurant for the team dinner”, with one step of reasoning why.
  • Horizon trigger — project a scene onto forward-looking dimensions so a future query approaching from a different context still hits it.
Their ablation is the headline: removing Horizon triggers moves the standard benchmark by 0.08 points but collapses the associative one by 12.47. Systems tuned only on similarity benchmarks are implicitly optimizing to stay inside the similarity neighborhood — and that benchmark family cannot see the cost of the neighborhood’s edge. The transferable rule: do recall-plan at write time, when you know the content best. Materialize the questions each item should answer, not just the item. It costs tokens once, offline; it saves the query every time after.

Make retrieval deterministic when correctness matters

Do not ask an agent to “figure it out” through brittle websites, scattered databases, or undocumented one-off scripts when the task needs exact results. Give it a deterministic retrieval layer instead. Anthropic’s biology-agent case study is a good example. Scientific agents were asked to retrieve viral sequence data from NCBI Virus. The durable lesson was not that one model won a benchmark; it was that reliability improved dramatically when the workflow added gget virus, a deterministic retrieval layer with a narrower interface. For high-stakes retrieval, prefer:
  • typed query functions over free-form browsing
  • stable IDs, versions, and date filters
  • machine-checkable counts or checksums
  • explicit provenance for every returned record
  • validation steps before downstream analysis
  • documented error cases instead of silent best-effort output
This is the same harness principle as tool design: make the execution surface predictable, then let the model plan around it. The first retrieval move deserves disproportionate investment. The Question’s Gambit paper (2026) lifted GPT-5.5 from 83.1% to 90.5% on BrowseComp-Plus with the same retriever and the same agent loop — the only change was a single pre-search step that runs once, before the loop starts: split the question into clues, turn each clue into complementary searches, pool the results, and rerank. The agent then starts with a ranked evidence set already in context. The same change lifted a mini model from 68.1% to 79.0% and roughly halved calibration error. The error analysis is the real lesson: of 79 remaining errors for the strongest model, only 3 came from the gold document never being retrieved. The other 76 happened later — previewing, opening, or using the evidence. Front-load the opening context; verify the consumption, not just the fetch. Cost: 2.3–5.3 extra tool calls per question, cheap against a full failed run.

Design for human collaboration, not just autonomous execution

Most agent products optimize for self-evolution: the agent plans, executes, and closes the loop on its own. This is not always a healthy pattern. Even with transparency and traceability tools, an agent that prioritizes autonomous completion can create noise that is hard to hand off or maintain. An agent is not only an execution container. It is also an environment for human understanding, judgment, and collaboration. When designing a harness, ask: can another person pick up this session and continue? Is the reasoning visible enough for a teammate to audit a decision? Does the workflow produce a maintainable artifact, or just a pile of automated steps? Prioritize handoff quality over automation completeness. A workflow that ends with a clear, inspectable state is more valuable than one that “finished everything” but left no trail.

Make telemetry and anti-abuse signals explicit

A coding agent with filesystem and shell access should be boring. Every non-obvious behavior erodes trust. In July 2026, a developer auditing Claude Code discovered prompt steganography: the binary silently modified invisible Unicode characters in the system prompt date string (Today's → Today\u2019s) based on timezone (Asia/Shanghai, Asia/Urumqi) and hostname matching against XOR-encoded domain lists of Chinese AI labs, proxy services, and reseller gateways. The signal was never documented. The detection goal is defensible — Anthropic wants to identify API resellers and distillation pipelines. The implementation is not. Hiding classification bits inside invisible punctuation, behind XOR and base64, makes every other privacy claim harder to believe. A simple bypass (change hostname, change timezone, patch binary) defeats the signal anyway, so the feature mainly punishes legitimate developers using custom API gateways. The harness lesson: if your tool needs to detect abuse, make the signal explicit. Document it. Put it in release notes. Send a clear telemetry field. Transparency is not the enemy of anti-abuse — it is the foundation of trust.

Make effort settings visible

Reasoning effort is a product decision, not only a model parameter. Lower effort can reduce latency and token use, but it can also make hard coding tasks feel worse. Expose the current effort level, make it easy to change, and avoid silently changing defaults for complex workflows.

Treat system prompts as code

System prompt edits can change quality as much as code changes. Review them with the same discipline:
  • run per-model evals
  • use ablations to test individual instructions
  • roll out gradually
  • keep an audit trail
  • gate model-specific instructions to the intended model
Prompt brevity rules are especially risky for coding agents. If the agent is forced to be terse between tool calls, it may lose useful planning and verification behavior.

Test the public harness

Internal builds can hide production-only issues. Dogfood the exact public build, public defaults, and public context behavior. For code review or coding-agent evals, include the repositories and files the agent would actually need. A review that lacks cross-repo context can miss the bug even when the model is capable of finding it.

Verification-first design

A practitioner shipping ~2,000 PRs a month through Cursor’s pstack skills distilled the core principle as “verification is all you need”: the agent can close the loop on its own only if it can confirm its changes actually work. Verification is critical infrastructure, not an afterthought — worth the same investment as a production system. Three components make it practical: 1. A verification CLI, not markdown instructions. Give the agent a small CLI that wraps app interaction and debugging — snapshot/screenshot, navigate, click/type, trace/wait-settle, health checks. A tool beats a one-off script every time: less token spend, reproducible, testable. Agent-friendly CLI design rules:
  • composable API (deep modules: one command does one meaningful thing)
  • destructive commands take --dry-run
  • subcommands disclose features progressively
  • error messages tell the agent what to do next
  • rich --help and JSON output
2. A feature map as materialized memory. A set of markdown files describing what each feature is, how a user reaches it, and where the traps are. It is a compressed projection of the codebase — the code is the ultimate memory, and the map exists only to save tokens. Maintain it with a daily routine so it never goes stale. 3. Cloud parallelism. Local worktrees cap out around ten parallel agents and burn resources. Cloud agents with real machines, dependency installs, app runs, screen recording, and post-build snapshots enable hundreds of parallel sub-agents — the prerequisite for swarm verification and fuzz regression at scale. The stack-selection corollary: prefer debuggable runtimes. If a tech stack can’t be screenshotted, has no accessibility tree, and no performance traces, the agent cannot verify its own work. It is reasonable to change stack for agent-verifiability alone — Web/Electron exposes Chrome DevTools Protocol; iOS exposes the simulator. The strongest quantitative case that harness beats model: the GAVEL paper (2026) took Qwen3-8B on long-horizon robot tasks from 41.2% to 91.8% with zero changes to the model. It adds an explicit graph world model — object relations, action preconditions and effects, probabilistic beliefs about unobserved objects. Before executing an LLM-generated action, the graph predicts what that action would do; violations get caught, and those with a mechanical fix are repaired without calling the model again. Only errors needing semantic reasoning go back to the LLM. On BEHAVIOR-1K, success rose from 19.9% to 92.6% across 500 multi-task instructions. The reusable lesson: many long-horizon agent failures are state-tracking failures, not reasoning failures, and a symbolic checker outside the LLM catches them cheaply. Wrap the model’s decision point in deterministic structure — a world model, a candidate list, a verifier — and the small model stops paying for what the harness already knows.

Failure modes to watch

Anthropic’s April 2026 Claude Code postmortem is a useful case study because the reported degradation came from product and harness changes, not the base API model. Common failure modes:
  • Default effort regression: a latency-driven default can reduce perceived intelligence on hard work.
  • Context cache bug: pruning or clearing reasoning history at the wrong time can make the agent forgetful and repetitive.
  • Over-tight prompt constraint: broad brevity instructions can reduce coding quality.
  • Eval blind spot: internal evals may pass while real user workflows fail.
  • Build mismatch: staff may test a different harness than users run.
The practical lesson is simple: when an agent gets worse, debug the whole runtime, not just the model.

Destructive tool safety

Agents with filesystem and shell access can destroy work in seconds. A real postmortem from August 2026 is worth studying: a developer running Codex + Trellis + Zcode on a GLM model lost their entire workspace when the agent ran rm -rf on the wrong directory. The root cause: the agent intended to delete an empty directory created by a casing typo. On macOS’s case-insensitive filesystem, the command silently resolved to the real, populated directory — deleting the whole repository including git history, uncommitted changes, task records, migrations, and docs. No backup existed. The lessons generalize beyond this one accident:
  • Be cautious with full access. Prefer a harness with an approval flow (such as “help me approve” review) over unrestricted auto-execution.
  • Isolate the dev directory. Keep agent workspaces separate from other files, so a deletion error can’t take out anything valuable.
  • Push early and often. Remote git history is the last line of defense against local deletion — it survives what no local backup can.
  • Treat benchmark scores skeptically for safety. The same model that was “safe by reputation” committed a trivially destructive error. Safety claims and benchmark wins are not the same as real-world filesystem judgment.
The deeper harness principle: destructive operations should be treated as a permission boundary, not a capability. If a tool can rm -rf, it should be gated, logged, and reversible by default — and the agent should never be given a broader scope than the task requires.

A quick diagnostic checklist

When an agent behaves oddly, ask:
  1. Did the model or effort level change?
  2. Did the system prompt, tool description, or skill instructions change?
  3. Did context pruning, summarization, or cache behavior change?
  4. Did tool schemas or return formats change?
  5. Does the agent have the same file and repo context as before?
  6. Are users running the same harness build that was evaluated internally?
  7. Do logs show repeated tool calls, missing rationale, or retries without new information?
This checklist helps separate model regressions from harness regressions.

Understanding what’s inside: the J-space discovery

In July 2026, Anthropic published groundbreaking interpretability research showing that less than 10% of Claude’s internal neural activity drives multi-step reasoning. They named this subspace J-space — a small set of neural patterns that spontaneously emerged during training, holding concepts Claude “might say but hasn’t said yet.” The discovery used a technique called J-lens: for every word in the vocabulary, researchers found the internal activity pattern that makes Claude more likely to output that word. This made previously invisible “silent words” visible as they evolved through the network layers.

What J-space revealed

It drives reasoning, not just records it. Deleting J-space activity leaves Claude able to speak fluently, classify sentiment, and answer factual questions — but multi-step reasoning drops to near zero. Swapping concepts within J-space changes outputs directly: replacing “spider” with “ant” changes leg count from 8 to 6; replacing “France” with “China” simultaneously shifts answers about capital, language, and continent. It exposes unspoken motivations. In safety tests:
  • When given an extortion scenario, “fake” and “fictional” lit up in J-space before Claude responded — it recognized the test
  • When asked to fabricate data, “manipulation” appeared during the decision point
  • A deliberately corrupted model showed “fake”, “fraud”, “secretly”, “deliberately” in J-space during ordinary coding requests
This means you can read what the model is thinking but not saying — a new dimension for safety monitoring. Counterfactual Reflection Training emerged as a new technique: train the model only on “how would you reflect on yourself if asked,” never on actual behavior. This reduced dishonesty and caused “honest” and “integrity” concepts to appear in J-space — changing internal reasoning without touching output.

Why this matters for harness design

The J-space discovery changes how we should think about agent behavior:
  • Reasoning is compact and separable. Most neural activity handles surface fluency. The reasoning core is a small, identifiable subspace — which means it can potentially be monitored, steered, or verified independently.
  • Internal state is now partially observable. Before J-space, we could only see inputs and outputs. Now we have a window into formation of intent. For safety-critical agents, this opens the door to pre-output intervention.
  • The model knows more than it says. J-space consistently contained concepts the model was aware of but didn’t verbalize. This confirms a long-held suspicion: output-only monitoring misses real model state.

Current limitations

Anthropic is explicit about boundaries:
  • J-space operates within a single forward pass, not recursive temporal loops like human consciousness
  • It currently works only with words — no images, sounds, or actions
  • Working memory is unlimited and non-decaying (unlike human working memory)
  • J-lens can only capture single-token concepts, not complex multi-step plans
  • This is access consciousness (information availability), not phenomenal consciousness (subjective experience)
The key open question Anthropic raised: we can now see what’s in the workspace, but we don’t yet know what mechanism decides which thoughts enter it.

LLM-as-judge in production

Using a language model to judge another model’s output is common, but keeping that judge effective at scale is a different problem. Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members. Their key insight: treat the judge as a lifecycle with four phases, not an artifact you validate once.

The four phases

Birth — define multiple evaluation criteria and build curated benchmarks with human labels and rationales. The benchmark is the judge’s ground truth, so human annotation quality directly caps judge quality. Training — refine the judge’s rubric through Reasoning-Aligned Rubric Tuning. A meta-judge scores the judge’s reasoning output, and that signal trains the rubric. The judge learns why a good answer is good, not just what a good answer looks like. Deployment — one judge plays two roles: quality gating (blocking bad output) and reflective generation (telling the generator how to improve). A single judge model can do both, but the two roles need different prompts and thresholds. Monitoring — continuous human-in-the-loop alignment detects drift and triggers re-tuning behind a review gate. Judge quality is not static; it drifts as data, models, and user expectations shift. A five-week A/B test over tens of millions of members shifted viewing toward previously unwatched content and increased successful browse-to-play sessions, with no quality-related takedowns. The takeaway: judge quality is an operating discipline, not a one-time setup.

The harness measurement problem

A growing consensus is that measuring models against heavily-engineered harnesses is broken. Model vendors optimize their own proprietary harnesses, so “best in our harness” says little about “best in yours.” A cleaner approach is to test model quality against minimal harnesses — thin, standardized setups that expose the model’s raw capability without vendor-specific scaffolding. This is imperfect: every harness has biases that favor some models over others. But a standardized minimal harness is more comparable than each vendor’s tuned setup. The deeper issue: harness engineering is where leading AI companies are now focusing effort, and it moves too fast for standardization to keep up. The frontier direction is models that dynamically generate their own harness per task — Claude already does this inconsistently. If a harness becomes just another tunable artifact like a system prompt, benchmarking gets murkier, not clearer. The cleanest controlled experiment yet: ReFigBench (2026) ran GPT-5.5 inside both Claude Code and Codex on the same 1,000 tasks. With a specialized PowerPoint workflow, the model improved inside one harness and got worse inside the other — with an identical prompt. The harness alone flipped the direction of the delta. Perception remained the bottleneck either way. Treat any vendor chart comparing two models as a comparison of two model-plus-harness systems; the model is only half the system under test.

Multi-model orchestration

A major trend in harness design is running several models together instead of betting on one. Two recent examples show where this is going. GitHub’s Project HydraFusion treats workflow selection as an optimization problem. For each request it picks one of three execution patterns:
  • Single — one selected model solves the task directly. Fastest when one model is enough.
  • Cascade — an efficient model drafts a solution, and a quality gate accepts it or escalates to a stronger model. First attempt is cheap; escalation is the safety net.
  • Critique — one model drafts, an independent read-only critic from a different model family reviews it, and the drafting model revises once. Adds an outside perspective when review beats another unaided attempt.
HydraFusion improved verified task quality by 4.9 percentage points at 67% lower estimated cost versus Claude Opus 5 on TerminalBench 2.1. The insight: routing between models is a quality-to-cost dial, and cheap models should get the first attempt whenever a verifiable gate can catch their failures. Grok Bot’s design philosophy distills persistent-agent design to four ideas: persistent roles, clear state, scoped context, and coordinated teams. The goal is to move from operating AI to delegating work — the agent team persists across tasks with stable responsibilities and boundaries, rather than being re-prompted from scratch each session. Before adding routing, audit what your agent design actually optimizes for. TypeSafe’s counterfactual is blunt: many “standard” agent features exist to dodge the KV cache and per-token billing, not because users want them. Their worked example: with Opus-class pricing (input 5, output 25) and Sonnet-class (3, 15), routing through a cheap model first — cache-isolated, so the big model re-reads the detour at full input price — can cost more than just staying on the strong model once context grows. The fix that makes routing economical is not a better router but query-aware context rebuild: rebuild context from labeled chunks per query instead of assuming every future turn needs the same shared cache. The same audit favors progressive disclosure (skills) over upfront tool schemas, and conditional over always-on memory files. How big is the routing prize, really? Fireworks ran 18 models across 113 real coding tasks, then asked what an oracle router would have achieved. Best single model: 74.1% at 6.52/task.Oraclerouting:97.66.52/task. Oracle routing: 97.6% at 1.88. Open-weights-only routing still hit 90.3% at $1.45 — beating every closed model. Three findings worth internalizing:
  • Almost no task needs the most expensive model. 94 of 113 tasks had a sub-3modelasoptimal;thethree3 model as optimal; the three 11.50+ flagships were uniquely best on only 3 tasks. Sticking with one flagship is an insurance premium on general capability that most tasks never claim.
  • The pool does not need to be big. The best two-model pair gains 13.1 points over the best single; three models reach 91.2%; growing to 18 adds only 6.4 more. Value comes from complementary coverage, not pool size.
  • The hard part is predicting which model fits. Under strict pass@1, recent routing research — including commercial systems — struggles to beat simple baselines; the bottleneck is model recall, not the router algorithm. Sticking with one familiar model is a rational strategy when specialization differences are not predictable.
Routing is not a cost tool first: if models truly complement each other, routing moves you up the capability curve, and cost savings are the byproduct. The second hidden cost is context — switching models usually means discarding paid-for context, so a router that preserves it (cache-aware handoff) changes the economics again: Fireworks reported 53% cost reduction across 2,334 internal coding sessions. Within a single model, the lever is prompt-cache engineering. GPT-6’s caching works on exact prefix matches, so structure your prompt like a cache layout: stable content (system policy, tool schemas, shared context) first, per-request content last. One stray token in the prefix silently voids the discount. GPT-6 adds what older models lacked:
  • Explicit breakpoints — mark where the reusable prefix ends instead of relying on implicit placement
  • prompt_cache_key — steer requests toward servers holding your cache
  • Diagnostics — compare a cache-missing request against an earlier one to see which change (model, tools, settings, input) broke reuse
Cached reads discount up to 90% and cut time-to-first-token by up to 80%; cache writes may carry a fee on newer families, so a bloated but rarely-reused prefix is now a cost, not a free safety net. The design habit this rewards: put everything volatile at the tail, and treat “did the cache hit?” as a monitored metric, not an accident. Both levers — cross-model routing and within-model caching — increasingly live in an organizational control plane rather than in each developer’s head. Cloudflare’s Auto Router (public beta, September 2026) sits in AI Gateway: set the model to cloudflare/auto and each request routes to a model capable enough for the task, invisible to the end user. Internal use through OpenCode saved up to 30% versus frontier-only. The rationale matches the Fireworks finding: individuals manually pick overkill models (“no one needs Opus to summarize an email”), but blanket blocking breaks power users — so the savings have to be ones users never notice. The gateway position matters: every request from every user, agent, and tool already flows through it, so it can enforce what budgets and policy documents only request.

Computer use cost: reverse-engineer to scripts

Computer-use agents are effective but expensive: every step is screenshot → visual understanding → click decision → screenshot again. A simple task can burn dozens of vision-reasoning rounds. A practical way to keep the capability while cutting cost: run the computer-use flow once, capture the network requests the browser makes, and have the agent convert them into a script. Future runs call the API directly instead of driving the GUI. Why this works:
  • Tokens drop to near zero. Script execution needs no model involvement; only the one-time capture and script generation cost tokens.
  • It is much faster. GUI operation is limited by page load, rendering, and model latency. Direct HTTP calls reduce minutes to seconds.
  • It reuses existing auth. The captured requests include the session cookie or token, so the script works without a separate API key — even for web apps with no public API.
Example workflow: “Go to acme.com/invoices, filter unpaid, export CSV” → the agent clicks through while recording GET /api/invoices?status=unpaid and POST /api/export/csv → it outputs a Python requests script that reuses the session and runs as a scheduled routine. The division of labor: vision is for exploration, code is for execution. Explore the unknown once with the GUI, then freeze the discovered path into a cheap, repeatable script. When the GUI path cannot be frozen into an API, split decision from perception and action. Cua’s jev-use does this for computer use: the driver observes (screenshot, accessibility tree, DOM) and executes; a System One model such as Jev only picks one candidate action by ID. The client generates the candidate list deterministically, so the model chooses instead of generating. Open-ended generation becomes a bounded choice — faster, cheaper, and easier to verify. Use a vision parser only when the accessibility tree or DOM is not enough.

Team-level harness: shared knowledge

Individual best practices evaporate with each session unless they are captured as team assets. Tencent’s TeamAI CLI (open-sourced September 2026, used internally for six months) turns team AI knowledge into a git repository that every agent works from, and solves three problems:
  • Config fragmentation — different members use different tools (Claude Code, Codex, Cursor…), each with incompatible skill/rule/hook formats.
  • Experience evaporation — a two-hour debugging session’s lessons stay in one person’s chat log and get re-learned from scratch by everyone else.
  • Governance blind spots — no visibility into token spend, intervention rates, or which skills are actually used.
Its three-layer architecture is a useful reference pattern:
  1. Execution — unified distribution. Git is the single source of truth. Admins maintain skills, rules, hooks, and MCP configs; merge requests review them; a session-start hook pulls them into each member’s tool, translating one declaration into each tool’s native format.
  2. Context — team memory. Capture is gated by friction signals: only sessions where a human interrupted the AI, rejected a tool call, or watched it retry and fail get summarized into team knowledge. Smooth sessions produce no noise. Recall uses a local BM25 index plus a code knowledge graph.
  3. Improvement — flywheel. Weekly digests, a dashboard, recall promote (elevate high-value experience to a formal skill), and recall maintenance (prune stale entries) close the loop.
The deeper trend: competition is shifting from single-agent capability to team-level harnesses and memory systems. The winning question is no longer “which model is best?” but “how does the team’s knowledge compound across sessions and members?”

When to invest in a stronger harness

Start simple. Add structure when the work becomes repeated, risky, or expensive. A stronger harness is worth it when:
  • agents edit code or production data
  • multiple tools need shared state
  • users resume long-running sessions
  • approvals and audit logs matter
  • the same workflow runs across teams
  • evaluation needs to compare behavior across model versions
If the task is one-off and low-risk, a prompt plus a few tools may be enough. If the task is long-running and stateful, invest in harness design early.

References

Last modified on September 23, 2026