Skip to content
FunDev
FunDev
ai

Harness Engineering: The Developer's New Role in the Age of AI Agents

Harness Engineering: The Developer's New Role in the Age of AI Agents
7 views
12 min read
#ai

From Writing Code to Enabling Agents to Work

OpenAI produced a million lines of production code in five months. The amount of application code written directly by a human was zero lines. Three engineers, later seven, built an internal beta product using only Codex agents and merged roughly 1,500 PRs.

What did the human engineers do in this experiment? They did not implement functions or fix bugs. Instead, they designed the repository structure, defined architectural rules, built linters and tests, and, when an agent got stuck, added new tools and constraints rather than rewriting the prompt.

This is **harness engineering**.

Prompt → Context → Harness: Three Stages of Evolution

To understand harness engineering, first look at how our use of AI has evolved.

StagePeriodCore questionAnalogy
Prompt Engineering2022~2024“How should I phrase it so it understands?”Giving a horse voice commands
Context Engineering2025“What should I show it so it works well?”Giving a horse a map and signposts
Harness Engineering2026“What environment lets it work on its own?”Putting reins, saddles, fences, and roads in place to run ten horses at once

Prompt engineering optimizes the text instructions sent in a single LLM call. It operates at the level of “Phrasing it this way gets a better answer.”

Context engineering goes one step further. It designs everything inside the context window: documents retrieved through RAG, tool definitions, message history, and memory. The question is how to curate the tokens the model sees.

Harness engineering designs everything outside the agent, from CI pipelines, evaluation loops, and linting rules to observability infrastructure and lifecycle management. It includes prompts and context, but starts from the recognition that those alone are insufficient.

Phil Schmid of Google DeepMind put it this way.

The model is the CPU, and the harness is the operating system. However powerful the CPU, a poorly designed OS limits its performance.

Mitchell Hashimoto, the creator of Terraform, established the term in his February 2026 post, “My AI Adoption Journey.” The central definition is: when AI makes a mistake, go beyond simply fixing it and design guidelines and tools that systematically prevent the same mistake from happening again.

The Three Pillars of a Harness

OpenAI's experiment revealed three core components of a harness.

1. Context Engineering

Structure knowledge inside the repository so the agent can find and read the context it needs for its work.

OpenAI initially created one enormous AGENTS.md file. It failed because it wasted the context window. The team ultimately placed 88 separate AGENTS.md files throughout the repository.

The key principle is to treat AGENTS.md as a table of contents rather than an encyclopedia. Keep it to around 100 lines, and put detailed design documents in docs/. The agent reads the contents and follows them to the documents it needs.

# AGENTS.md (목차 역할, ~100줄)
├── 프로젝트 개요
├── 아키텍처 계층 규칙 → docs/architecture.md
├── 코딩 컨벤션 → docs/conventions.md
├── 테스트 가이드 → docs/testing.md
├── 품질 지표 → docs/quality-scores.md
└── 기술 부채 목록 → docs/tech-debt.md

Knowledge outside the repository—in Slack, Notion, or verbal agreements—is also invisible to the agent. Every important decision and pattern needs to be incorporated into the repository as code or documentation. The repository must become the single source of truth.

2. Architectural Constraints

Mechanically enforce “You must not do this,” rather than merely telling the agent, “Do it this way.”

OpenAI defined strict dependency layers for each business domain.

Types → Config → Repo → Service → Runtime → UI
 
규칙: 상위 레이어는 하위 레이어만 참조할 수 있다.
       UI가 Types를 직접 참조하는 것은 허용되지만,
       Types가 Service를 참조하는 것은 금지된다.

Custom linters and structural tests enforce these rules in CI. PRs that violate them are blocked automatically.

There is another key design choice: the linter's error messages themselves act as repair instructions. The messages are designed so that when an agent encounters a violation, it can read the error and fix it independently.

❌ 기존: "Layer violation detected"
✅ 개선: "Layer violation: Runtime cannot import from UI.
         Move the shared type to Types layer (src/types/).
         See docs/architecture.md#layer-rules for details."

3. Garbage Collection

As agents generate code, they copy bad patterns too, gradually eroding the architecture. OpenAI called this **“AI slop.”**

The solution is not manual cleanup. It is to encode “golden principles” in the repository and let background Codex tasks detect drift and automatically create refactoring PRs. Instead of a human spending every Friday cleaning up technical debt left by AI, another AI cleans it up automatically every day.

A harness is not a one-time setup. It is closer to an operating system that codifies rules and applies them automatically every day.

What a Harness Engineer Actually Does

There are no official job postings with the title “Harness Engineer” yet, but functionally equivalent roles already exist. Here is what the work involves.

Environment Design

Make the repository agent-friendly. Use AGENTS.md as a table of contents and structured docs/ as the source of truth. Version execution plans, design documents, technical debt, and even quality scores inside the repository. The goal is to let the agent work from the repository alone, rather than from knowledge in people's heads.

Tool Design

Carefully design the number, names, and return values of the tools an agent will use. Anthropic outlined the following principles for good agent tools.

PrincipleDescription
An appropriate number of toolsToo many cause choice paralysis; too few limit capability
Clear namingA tool's name alone should reveal its purpose
High-signal return valuesReturn meaningful results
Token-efficient responsesReduce unnecessary information to conserve context

A Vercel experiment supports this. After removing 80% of the tools from a text-to-SQL agent and leaving only a single bash execution tool, its success rate rose from 80% to 100%. The number of steps, token use, and response time all fell too.

Constraint Design

Define architectural layers, dependency rules, and type, test, and lint rules, then enforce them through tools. The key is to enforce invariants without micromanaging the implementation. Set the boundaries of what is forbidden and let the agent decide how to implement the rest.

Feedback-Loop Design

Design how the agent responds to test failures, lint errors, and deteriorating observability metrics. Treat every agent failure as a signal to improve the harness. If the agent repeats the same mistake, add a lint rule or test rather than rewriting the prompt.

Memory and Handoffs for Long-Running Agents

Anthropic introduced an initializer-agent pattern for long-running agents.

1. 첫 컨텍스트 윈도우: 초기화 전용 프롬프트로 환경 셋업
2. init.sh 실행, 200개+ 기능 목록 생성 (모두 "failing" 표시)
3. progress 파일(JSON) 생성
4. 새 컨텍스트 윈도우가 열려도 progress 파일을 읽어 작업 이어감
5. git commit 로그로 진행 상황 추적

A key finding: JSON works better than Markdown for feature tracking because agents are less likely to make arbitrary edits to structured data. A harness engineer is effectively building **“an onboarding system for an AI team.”**

Entropy-Reduction Loops

Encode golden principles in the repository and let background agents automatically open quality-check and refactoring PRs. A harness is a continuously operated system, not something configured once and forgotten.

Practical Examples: The Difference a Harness Makes

OpenAI: One Million Lines, Zero Handwritten Code

MetricValue
Duration5 months
Code volume1 million+ lines
PR count~1,500
Team size3 → 7 engineers
PRs per engineer per day3.5 on average
Code written directly by humans0 lines

The central slogan: “Humans steer. Agents execute.” Humans set the direction; agents carry out the work.

Stripe: The Minions System

Stripe's Minions system produces more than 1,000 merged PRs per week. A developer posts a task in Slack, and an agent writes code in an isolated development environment (devbox), gets it through CI, and opens a PR for human review.

It can access more than 400 internal tools through an MCP server, but curates only about 15 tools per session to avoid “token paralysis.” The essence of a harness is providing the right tools, rather than simply providing many.

KRAFTON: The Difference Comes from the Harness

Korea's KRAFTON used its own “Terminus-KIRA” harness to achieve second place worldwide on TerminalBench 2.0 with 74.8% accuracy, just 0.3 percentage points behind OpenAI at 75.1%. It demonstrates that the same model can produce different results depending on the harness.

LangChain: Demonstrating It with a Benchmark

LangChain raised its benchmark score through harness improvements alone, without changing the model.

ChangeBeforeAfter
ModelSameSame
HarnessDefaultImproved
Score52.8%66.5%
RankOutside the top 305th

The key technique is the “reasoning sandwich” pattern: allocate more reasoning compute to planning and verification, and less to implementation. Think deeply; write code quickly.

Anthropic: The Claude Code Harness

It introduced a pattern separating the initializer agent from the coding agent. Even when a fresh context window opens during a long task, progress files and git logs allow the work to continue.

Core Principles

These harness-engineering principles appear across the examples.

The Harness Is the Competitive Advantage

Models are becoming more general-purpose and cheaper. The real differentiation is in the environment where the agent works. The KRAFTON and LangChain examples demonstrate this.

AGENTS.md Is a Map, Not a Manual

Do not try to explain everything in detail. Focus on its role as a table of contents that tells the agent where to find things. OpenAI used 88 separate files instead of one giant AGENTS.md.

Enforce Invariants and Leave the Implementation to the Agent

Linters and tests define only what must not be done. Leave the specific implementation to the agent's discretion. Micromanagement is inefficient whether it happens in the prompt or the harness.

Turn Failures into Assets

An agent's failure is a defect in the harness. Feed failures back into documentation, tests, and tool improvements so the same problem does not recur. Add a new constraint instead of “prompting harder.”

Make the Repository the Single Source of Truth

Agreements in Slack, conventions in Notion, decisions conveyed verbally—the agent cannot see any of them. Everything important must be incorporated into the repository.

Market Trends

AI-Generated Code at Big Tech Companies

CompanyShare of AI-generated codeSource
Google30%+ of new codeSundar Pichai
Microsoft20~30%Satya Nadella
MetaTargeting 50% within one yearMark Zuckerberg

The Hiring Market

As of March 2026, there are no job postings titled “Harness Engineer” yet. Functionally equivalent roles are being recruited under titles such as Agentic AI Specialist and AI Infrastructure Engineer. AI engineering positions are growing 300% faster than traditional software engineering roles.

Trends in Korea

**Toss** introduced harness engineering on its technical blog as a mechanism for raising the floor of organizational productivity. Its argument is that LLM use “cannot remain a matter of individual intuition; it must become a system that the team designs and deploys.”

At Korea's first Ralph-ton hackathon, 13 elite developers set up harnesses on the first day, let their agents run autonomously overnight, and reviewed the results the next day. The winning team generated 100,000 lines of code with agents, of which 70,000 were tests.

Limitations and Open Questions

Harness engineering is an interesting direction, but it has its critics.

CriticismDetails
Insufficient functional verificationThoughtworks' Birgitta Böckeler noted that OpenAI's example handles structural constraints well but does not discuss behavioral verification sufficiently
Limits of the metaphorThe “harness” metaphor can frame AI solely as something to control and underestimate the cognitive and social aspects of human-agent collaboration
Long-term maintainabilityIt remains unknown whether an entirely agent-generated codebase can maintain architectural consistency over several years
Difficulty applying it to legacy systemsBuilding a harness for a new repository differs in difficulty from retrofitting one to an old codebase
Self-reporting biasOpenAI's case is a self-report of its own experiment, with limited independent verification

Thoughtworks described it as a term that was “two weeks old.” It is not yet a standardized official job title. The direction, however, is clear.

Harness Tools and Frameworks

ToolFeaturesRole in a harness
AGENTS.mdAdopted by 60,000+ GitHub repositories; under the Linux FoundationStandardizes agent instructions
CLAUDE.mdThe agent harness behind Anthropic's Claude CodeManages long-running agents
LangGraphMiddleware-based agent orchestrationHarness middleware such as LoopDetection and ReasoningSandwich
CrewAIRole-based multi-agent orchestrationDual structure: Crews for dynamic collaboration and Flows for deterministic tasks
OpenAI CodexAgent-based coding + App ServerLong-running autonomy (6+ hours) + Reusable harness protocol

OpenAI designed a JSON-RPC server called Codex App Server to reuse the Codex harness across products. The CLI, VS Code extension, web app, and macOS app all run on the same harness. This shows that harness engineering goes beyond writing new prompts for each app to encompass designing a reusable agent execution environment.

What You Can Do Now

An article on Martin Fowler's site asked whether harnesses will become **“the new service templates.”** Just as teams spin up new services through golden paths today, tomorrow they may choose from a harness catalog bundling custom linters, structural tests, AGENTS.md, and cleanup agents.

You do not need to build a grand system all at once. Here are things you can start doing now.

  1. Write AGENTS.md (or CLAUDE.md)—summarize the repository's structure, rules, and main patterns in no more than 100 lines
  2. Make linter errors agent-friendly—include how to fix the problem, rather than just saying there is a violation
  3. Break work into verifiable units—split tasks into small units whose success or failure can be determined through tests
  4. Document agent failures—when a failure repeats, add a constraint instead of another prompt
  5. Move knowledge into the repository—document conventions and decisions that exist only in Slack or Notion under docs/

Wrapping Up

The central thesis of harness engineering is simple.

As models become increasingly general-purpose, the harness determines whether an agent succeeds or fails.

The ability to design an environment where agents produce correct code is becoming more important than writing clean code yourself. From writing code to enabling agents to work: that shift in the developer's role has already begun.


References

Related posts

Comments

Korean and English pages share this conversation.

Write a comment

0 / 5,000
You will need this password to edit or delete this comment.

Loading comments…