From Writing Code to Enabling Agents to Work
OpenAI produced a million lines of production code in five months. The amount of application code written directly by a human was zero lines. Three engineers, later seven, built an internal beta product using only Codex agents and merged roughly 1,500 PRs.
What did the human engineers do in this experiment? They did not implement functions or fix bugs. Instead, they designed the repository structure, defined architectural rules, built linters and tests, and, when an agent got stuck, added new tools and constraints rather than rewriting the prompt.
This is **harness engineering**.
Prompt → Context → Harness: Three Stages of Evolution
To understand harness engineering, first look at how our use of AI has evolved.
| Stage | Period | Core question | Analogy |
|---|---|---|---|
| Prompt Engineering | 2022~2024 | “How should I phrase it so it understands?” | Giving a horse voice commands |
| Context Engineering | 2025 | “What should I show it so it works well?” | Giving a horse a map and signposts |
| Harness Engineering | 2026 | “What environment lets it work on its own?” | Putting reins, saddles, fences, and roads in place to run ten horses at once |
Prompt engineering optimizes the text instructions sent in a single LLM call. It operates at the level of “Phrasing it this way gets a better answer.”
Context engineering goes one step further. It designs everything inside the context window: documents retrieved through RAG, tool definitions, message history, and memory. The question is how to curate the tokens the model sees.
Harness engineering designs everything outside the agent, from CI pipelines, evaluation loops, and linting rules to observability infrastructure and lifecycle management. It includes prompts and context, but starts from the recognition that those alone are insufficient.
Phil Schmid of Google DeepMind put it this way.
The model is the CPU, and the harness is the operating system. However powerful the CPU, a poorly designed OS limits its performance.
Mitchell Hashimoto, the creator of Terraform, established the term in his February 2026 post, “My AI Adoption Journey.” The central definition is: when AI makes a mistake, go beyond simply fixing it and design guidelines and tools that systematically prevent the same mistake from happening again.
The Three Pillars of a Harness
OpenAI's experiment revealed three core components of a harness.
1. Context Engineering
Structure knowledge inside the repository so the agent can find and read the context it needs for its work.
OpenAI initially created one enormous AGENTS.md file. It failed because it wasted the context window. The team ultimately placed 88 separate AGENTS.md files throughout the repository.
The key principle is to treat AGENTS.md as a table of contents rather than an encyclopedia. Keep it to around 100 lines, and put detailed design documents in docs/. The agent reads the contents and follows them to the documents it needs.
# AGENTS.md (목차 역할, ~100줄)
├── 프로젝트 개요
├── 아키텍처 계층 규칙 → docs/architecture.md
├── 코딩 컨벤션 → docs/conventions.md
├── 테스트 가이드 → docs/testing.md
├── 품질 지표 → docs/quality-scores.md
└── 기술 부채 목록 → docs/tech-debt.mdKnowledge outside the repository—in Slack, Notion, or verbal agreements—is also invisible to the agent. Every important decision and pattern needs to be incorporated into the repository as code or documentation. The repository must become the single source of truth.
2. Architectural Constraints
Mechanically enforce “You must not do this,” rather than merely telling the agent, “Do it this way.”
OpenAI defined strict dependency layers for each business domain.
Types → Config → Repo → Service → Runtime → UI
규칙: 상위 레이어는 하위 레이어만 참조할 수 있다.
UI가 Types를 직접 참조하는 것은 허용되지만,
Types가 Service를 참조하는 것은 금지된다.Custom linters and structural tests enforce these rules in CI. PRs that violate them are blocked automatically.
There is another key design choice: the linter's error messages themselves act as repair instructions. The messages are designed so that when an agent encounters a violation, it can read the error and fix it independently.
❌ 기존: "Layer violation detected"
✅ 개선: "Layer violation: Runtime cannot import from UI.
Move the shared type to Types layer (src/types/).
See docs/architecture.md#layer-rules for details."3. Garbage Collection
As agents generate code, they copy bad patterns too, gradually eroding the architecture. OpenAI called this **“AI slop.”**
The solution is not manual cleanup. It is to encode “golden principles” in the repository and let background Codex tasks detect drift and automatically create refactoring PRs. Instead of a human spending every Friday cleaning up technical debt left by AI, another AI cleans it up automatically every day.
A harness is not a one-time setup. It is closer to an operating system that codifies rules and applies them automatically every day.
What a Harness Engineer Actually Does
There are no official job postings with the title “Harness Engineer” yet, but functionally equivalent roles already exist. Here is what the work involves.
Environment Design
Make the repository agent-friendly. Use AGENTS.md as a table of contents and structured docs/ as the source of truth. Version execution plans, design documents, technical debt, and even quality scores inside the repository. The goal is to let the agent work from the repository alone, rather than from knowledge in people's heads.
Tool Design
Carefully design the number, names, and return values of the tools an agent will use. Anthropic outlined the following principles for good agent tools.
| Principle | Description |
|---|---|
| An appropriate number of tools | Too many cause choice paralysis; too few limit capability |
| Clear naming | A tool's name alone should reveal its purpose |
| High-signal return values | Return meaningful results |
| Token-efficient responses | Reduce unnecessary information to conserve context |
A Vercel experiment supports this. After removing 80% of the tools from a text-to-SQL agent and leaving only a single bash execution tool, its success rate rose from 80% to 100%. The number of steps, token use, and response time all fell too.
Constraint Design
Define architectural layers, dependency rules, and type, test, and lint rules, then enforce them through tools. The key is to enforce invariants without micromanaging the implementation. Set the boundaries of what is forbidden and let the agent decide how to implement the rest.
Feedback-Loop Design
Design how the agent responds to test failures, lint errors, and deteriorating observability metrics. Treat every agent failure as a signal to improve the harness. If the agent repeats the same mistake, add a lint rule or test rather than rewriting the prompt.
Memory and Handoffs for Long-Running Agents
Anthropic introduced an initializer-agent pattern for long-running agents.
1. 첫 컨텍스트 윈도우: 초기화 전용 프롬프트로 환경 셋업
2. init.sh 실행, 200개+ 기능 목록 생성 (모두 "failing" 표시)
3. progress 파일(JSON) 생성
4. 새 컨텍스트 윈도우가 열려도 progress 파일을 읽어 작업 이어감
5. git commit 로그로 진행 상황 추적A key finding: JSON works better than Markdown for feature tracking because agents are less likely to make arbitrary edits to structured data. A harness engineer is effectively building **“an onboarding system for an AI team.”**
Entropy-Reduction Loops
Encode golden principles in the repository and let background agents automatically open quality-check and refactoring PRs. A harness is a continuously operated system, not something configured once and forgotten.
Practical Examples: The Difference a Harness Makes
OpenAI: One Million Lines, Zero Handwritten Code
| Metric | Value |
|---|---|
| Duration | 5 months |
| Code volume | 1 million+ lines |
| PR count | ~1,500 |
| Team size | 3 → 7 engineers |
| PRs per engineer per day | 3.5 on average |
| Code written directly by humans | 0 lines |
The central slogan: “Humans steer. Agents execute.” Humans set the direction; agents carry out the work.
Stripe: The Minions System
Stripe's Minions system produces more than 1,000 merged PRs per week. A developer posts a task in Slack, and an agent writes code in an isolated development environment (devbox), gets it through CI, and opens a PR for human review.
It can access more than 400 internal tools through an MCP server, but curates only about 15 tools per session to avoid “token paralysis.” The essence of a harness is providing the right tools, rather than simply providing many.
KRAFTON: The Difference Comes from the Harness
Korea's KRAFTON used its own “Terminus-KIRA” harness to achieve second place worldwide on TerminalBench 2.0 with 74.8% accuracy, just 0.3 percentage points behind OpenAI at 75.1%. It demonstrates that the same model can produce different results depending on the harness.
LangChain: Demonstrating It with a Benchmark
LangChain raised its benchmark score through harness improvements alone, without changing the model.
| Change | Before | After |
|---|---|---|
| Model | Same | Same |
| Harness | Default | Improved |
| Score | 52.8% | 66.5% |
| Rank | Outside the top 30 | 5th |
The key technique is the “reasoning sandwich” pattern: allocate more reasoning compute to planning and verification, and less to implementation. Think deeply; write code quickly.
Anthropic: The Claude Code Harness
It introduced a pattern separating the initializer agent from the coding agent. Even when a fresh context window opens during a long task, progress files and git logs allow the work to continue.
Core Principles
These harness-engineering principles appear across the examples.
The Harness Is the Competitive Advantage
Models are becoming more general-purpose and cheaper. The real differentiation is in the environment where the agent works. The KRAFTON and LangChain examples demonstrate this.
AGENTS.md Is a Map, Not a Manual
Do not try to explain everything in detail. Focus on its role as a table of contents that tells the agent where to find things. OpenAI used 88 separate files instead of one giant AGENTS.md.
Enforce Invariants and Leave the Implementation to the Agent
Linters and tests define only what must not be done. Leave the specific implementation to the agent's discretion. Micromanagement is inefficient whether it happens in the prompt or the harness.
Turn Failures into Assets
An agent's failure is a defect in the harness. Feed failures back into documentation, tests, and tool improvements so the same problem does not recur. Add a new constraint instead of “prompting harder.”
Make the Repository the Single Source of Truth
Agreements in Slack, conventions in Notion, decisions conveyed verbally—the agent cannot see any of them. Everything important must be incorporated into the repository.
Market Trends
AI-Generated Code at Big Tech Companies
| Company | Share of AI-generated code | Source |
|---|---|---|
| 30%+ of new code | Sundar Pichai | |
| Microsoft | 20~30% | Satya Nadella |
| Meta | Targeting 50% within one year | Mark Zuckerberg |
The Hiring Market
As of March 2026, there are no job postings titled “Harness Engineer” yet. Functionally equivalent roles are being recruited under titles such as Agentic AI Specialist and AI Infrastructure Engineer. AI engineering positions are growing 300% faster than traditional software engineering roles.
Trends in Korea
**Toss** introduced harness engineering on its technical blog as a mechanism for raising the floor of organizational productivity. Its argument is that LLM use “cannot remain a matter of individual intuition; it must become a system that the team designs and deploys.”
At Korea's first Ralph-ton hackathon, 13 elite developers set up harnesses on the first day, let their agents run autonomously overnight, and reviewed the results the next day. The winning team generated 100,000 lines of code with agents, of which 70,000 were tests.
Limitations and Open Questions
Harness engineering is an interesting direction, but it has its critics.
| Criticism | Details |
|---|---|
| Insufficient functional verification | Thoughtworks' Birgitta Böckeler noted that OpenAI's example handles structural constraints well but does not discuss behavioral verification sufficiently |
| Limits of the metaphor | The “harness” metaphor can frame AI solely as something to control and underestimate the cognitive and social aspects of human-agent collaboration |
| Long-term maintainability | It remains unknown whether an entirely agent-generated codebase can maintain architectural consistency over several years |
| Difficulty applying it to legacy systems | Building a harness for a new repository differs in difficulty from retrofitting one to an old codebase |
| Self-reporting bias | OpenAI's case is a self-report of its own experiment, with limited independent verification |
Thoughtworks described it as a term that was “two weeks old.” It is not yet a standardized official job title. The direction, however, is clear.
Harness Tools and Frameworks
| Tool | Features | Role in a harness |
|---|---|---|
| AGENTS.md | Adopted by 60,000+ GitHub repositories; under the Linux Foundation | Standardizes agent instructions |
| CLAUDE.md | The agent harness behind Anthropic's Claude Code | Manages long-running agents |
| LangGraph | Middleware-based agent orchestration | Harness middleware such as LoopDetection and ReasoningSandwich |
| CrewAI | Role-based multi-agent orchestration | Dual structure: Crews for dynamic collaboration and Flows for deterministic tasks |
| OpenAI Codex | Agent-based coding + App Server | Long-running autonomy (6+ hours) + Reusable harness protocol |
OpenAI designed a JSON-RPC server called Codex App Server to reuse the Codex harness across products. The CLI, VS Code extension, web app, and macOS app all run on the same harness. This shows that harness engineering goes beyond writing new prompts for each app to encompass designing a reusable agent execution environment.
What You Can Do Now
An article on Martin Fowler's site asked whether harnesses will become **“the new service templates.”** Just as teams spin up new services through golden paths today, tomorrow they may choose from a harness catalog bundling custom linters, structural tests, AGENTS.md, and cleanup agents.
You do not need to build a grand system all at once. Here are things you can start doing now.
- Write AGENTS.md (or CLAUDE.md)—summarize the repository's structure, rules, and main patterns in no more than 100 lines
- Make linter errors agent-friendly—include how to fix the problem, rather than just saying there is a violation
- Break work into verifiable units—split tasks into small units whose success or failure can be determined through tests
- Document agent failures—when a failure repeats, add a constraint instead of another prompt
- Move knowledge into the repository—document conventions and decisions that exist only in Slack or Notion under
docs/
Wrapping Up
The central thesis of harness engineering is simple.
As models become increasingly general-purpose, the harness determines whether an agent succeeds or fails.
The ability to design an environment where agents produce correct code is becoming more important than writing clean code yourself. From writing code to enabling agents to work: that shift in the developer's role has already begun.





Comments
Korean and English pages share this conversation.
Loading comments…