A self-evolving external harness for RCA
OpsHarness: A Self-Evolving Harness for Root Cause Analysis
OpsHarness turns diagnosis experience into reusable expertise—then verifies every evolution before it reaches the live harness.
The Chinese University of Hong Kong · ByteDance
Agentstrong reasoning · limited system context
Harness
Expertsystem knowledge · verified evolution
Two public benchmarks + one industrial deployment
01 — The RCA gap
General models are capable.
The RCA gap is a harness that learns.
Modern models already provide strong reasoning, planning, and tool use. What RCA still lacks is an external harness that adapts those capabilities to each system—and keeps improving as incidents reveal new patterns.
General capability is no longer the bottleneck
Modern models can reason, plan, and use tools, but they arrive without the context and operating experience of the target system.
RCA adaptation belongs in the harness
Put RCA skills, system knowledge, observability, and diagnosis tools around the model instead of rebuilding its general machinery.
Static adaptation must become verified self-evolution
Distill successful and failed diagnoses into atomic updates, then test every proposal before it reaches the live harness.
RCA superpowers
Specialization belongs outside the model.
OpsHarness preserves the model’s general capabilities and supplies the external skills, knowledge, observability, verification, and learning loop that RCA requires.
Why self-evolution matters
One corrected incident should improve the next.
A misleading symptom first sends the investigation down the wrong path. Once the real propagation chain and its caveat are recorded, a later incident can be localized with fewer wrong turns.
OpsHarness turns this manual learning loop into a reviewable, verified system process.
02 — The architecture
A harness with
memory and control.
OpsHarness separates what the agent knows from how that knowledge is acquired, tested, and promoted.
Data Plane
System-specific expertise is organized for progressive disclosure, so the agent loads only what each diagnosis needs.
- K0General RCA knowledgePrinciples, constraints, and core workflows
- K1System profileSchema, components, telemetry, and SOPs
- K2Mined workflowsReusable diagnosis paths learned from use
- K3Operations & rulesAtomic skills, evidence patterns, and caveats
input → procedure → outputControl Plane
Four workflows carry a system from cold start to an expert harness while keeping every self-modification reviewable.
- 01SetupProfile a new telemetry system
- 02DiagnoseInvestigate and rank root causes
- 03EvolveMine correct and failed trajectories
- 04VerifyGate changes before promotion
Verified self-evolution
Learn from both success and failure—without learning the wrong lesson.
OpsHarness compares successful and failed trajectories, turns their evidence into atomic proposals, and evaluates a staged harness before any update reaches production.
- 01Mine trajectoriesContrast useful and failed diagnostic paths.
- 02Propose atomic updatesCreate reviewable skills, rules, and caveats.
- 03Pass two gatesImprove source cases and avoid regression on held-out cases.
03 — The evidence
It improves with use.
The gate makes it reliable.
Across OpenRCA, RCAEval, four model backbones, two agent frameworks, and an industrial deployment, the gains are consistent and practical.
Final A@1 for full OpsHarness, versus 41.4% without evolution and 36.1% for Direct.
Final-window A@1 for full OpsHarness versus the non-evolving harness.
of evolution proposals are rejected; unverified evolution finishes at only 0.33 A@1.
Average A@1 for OpsHarness versus Direct across six production configurations.
Complete benchmark results
24 configurations across two benchmarks and six sub-datasets.
OpsHarness is evaluated with four model backbones. Specialized RCA agents are marked †; within each backbone, the OpsHarness row is emphasized.
Scroll horizontally to view all metrics →
| Framework | OpenRCA | RCAEval | Final A@1 |
||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Telecom | Bank | Market | Online Boutique | Sock Shop | Train Ticket | ||||||||||||||
| A@1 | A@3 | Avg | A@1 | A@3 | Avg | A@1 | A@3 | Avg | A@1 | A@3 | Avg | A@1 | A@3 | Avg | A@1 | A@3 | Avg | ||
| GPT-5.5 | |||||||||||||||||||
| RCA-Agent† | 36.4 | 45.5 | 55.0 | 35.7 | 64.3 | 65.0 | 23.2 | 53.6 | 55.0 | 0.0 | 11.1 | 46.0 | 5.6 | 5.6 | 29.0 | 0.0 | 0.0 | 24.0 | 16.8 |
| mABC† | 9.1 | 18.2 | 14.0 | 10.7 | 21.4 | 35.0 | 3.1 | 9.8 | 32.0 | 16.7 | 33.3 | 76.0 | 11.1 | 22.2 | 61.0 | 5.6 | 11.1 | 61.0 | 9.4 |
| Codex (Direct) | 54.5 | 54.5 | 64.0 | 57.1 | 64.3 | 72.0 | 23.2 | 46.4 | 59.0 | 66.7 | 83.3 | 95.0 | 61.1 | 77.8 | 89.0 | 44.4 | 72.2 | 83.0 | 51.2 |
| Codex (ICL) | 27.3 | 36.4 | 50.0 | 46.4 | 53.6 | 65.0 | 20.0 | 40.0 | 49.0 | 55.6 | 77.8 | 93.0 | 61.1 | 61.1 | 85.0 | 61.1 | 88.9 | 94.0 | 45.3 |
| OpsHarness (no-evolve) | 54.5 | 54.5 | 64.0 | 46.4 | 57.1 | 62.0 | 26.8 | 43.3 | 58.0 | 66.7 | 72.2 | 91.0 | 72.2 | 72.2 | 91.0 | 50.0 | 77.8 | 89.0 | 52.8 |
| OpsHarness | 72.7 | 72.7 | 77.0 | 64.2 | 71.4 | 78.0 | 37.1 | 66.5 | 72.0 | 72.2 | 88.9 | 96.0 | 77.8 | 88.9 | 93.0 | 72.2 | 94.4 | 96.0 | 66.0 |
| Claude Sonnet 4.6 | |||||||||||||||||||
| RCA-Agent† | 9.1 | 18.2 | 32.0 | 53.6 | 67.9 | 79.0 | 16.5 | 33.5 | 55.0 | 11.1 | 22.2 | 59.0 | 5.6 | 5.6 | 29.0 | 0.0 | 5.6 | 44.0 | 16.0 |
| mABC† | 0.0 | 0.0 | 0.0 | 3.6 | 25.0 | 40.0 | 3.1 | 6.7 | 30.0 | 11.1 | 33.3 | 76.0 | 11.1 | 11.1 | 59.0 | 5.6 | 22.2 | 69.0 | 5.8 |
| Claude Code (Direct) | 36.4 | 45.5 | 55.0 | 24.2 | 36.3 | 61.0 | 17.0 | 26.3 | 38.0 | 38.9 | 61.1 | 85.0 | 27.8 | 44.4 | 76.0 | 16.7 | 22.2 | 76.0 | 26.8 |
| Claude Code (ICL) | 45.5 | 54.5 | 64.0 | 21.4 | 39.3 | 58.0 | 26.7 | 36.7 | 49.5 | 44.4 | 72.2 | 89.0 | 27.8 | 44.4 | 72.0 | 38.9 | 55.6 | 76.0 | 34.1 |
| OpsHarness (no-evolve) | 36.4 | 63.6 | 73.0 | 39.3 | 46.4 | 61.0 | 28.6 | 35.7 | 50.0 | 38.9 | 77.8 | 89.0 | 33.3 | 50.0 | 79.0 | 50.0 | 72.2 | 89.0 | 37.8 |
| OpsHarness | 63.6 | 63.6 | 73.0 | 57.1 | 63.7 | 81.9 | 35.7 | 64.2 | 72.0 | 61.1 | 83.3 | 95.0 | 55.6 | 77.8 | 89.0 | 61.1 | 77.8 | 93.0 | 55.7 |
| GLM-5.2 | |||||||||||||||||||
| RCA-Agent† | 36.4 | 54.5 | 59.0 | 60.7 | 64.3 | 77.0 | 13.4 | 23.7 | 40.5 | 16.7 | 22.2 | 56.0 | 16.7 | 16.7 | 33.0 | 0.0 | 0.0 | 28.0 | 24.0 |
| mABC† | 0.0 | 9.1 | 14.0 | 0.0 | 17.9 | 34.0 | 6.7 | 6.7 | 20.5 | 11.1 | 33.3 | 74.0 | 11.1 | 11.1 | 52.0 | 0.0 | 11.1 | 59.0 | 4.8 |
| Codex (Direct) | 54.5 | 63.6 | 68.0 | 57.1 | 57.1 | 67.0 | 26.8 | 50.0 | 65.0 | 38.9 | 77.8 | 93.0 | 44.4 | 72.2 | 89.0 | 55.6 | 72.2 | 87.0 | 46.2 |
| Codex (ICL) | 36.4 | 36.4 | 45.0 | 46.4 | 46.4 | 55.0 | 16.7 | 23.3 | 36.4 | 61.1 | 83.3 | 95.0 | 50.0 | 77.8 | 91.0 | 50.0 | 77.8 | 87.0 | 43.4 |
| OpsHarness (no-evolve) | 54.5 | 63.6 | 80.0 | 57.1 | 64.3 | 74.0 | 28.6 | 42.9 | 52.0 | 66.7 | 88.9 | 89.0 | 44.4 | 55.6 | 74.0 | 55.6 | 88.9 | 94.0 | 51.2 |
| OpsHarness | 72.7 | 81.8 | 87.0 | 57.1 | 64.3 | 74.0 | 42.9 | 50.0 | 61.0 | 77.8 | 88.9 | 89.0 | 55.6 | 66.7 | 76.0 | 88.9 | 94.4 | 98.0 | 65.8 |
| DeepSeek-V4 | |||||||||||||||||||
| RCA-Agent† | 45.5 | 45.5 | 50.0 | 28.6 | 35.7 | 47.0 | 9.8 | 9.8 | 26.0 | 5.6 | 5.6 | 35.0 | 0.0 | 5.6 | 26.0 | 0.0 | 0.0 | 22.0 | 14.9 |
| mABC† | 0.0 | 0.0 | 5.0 | 3.6 | 14.3 | 32.0 | 0.0 | 3.1 | 16.0 | 11.1 | 33.3 | 72.0 | 0.0 | 5.6 | 43.0 | 0.0 | 0.0 | 48.0 | 2.5 |
| Codex (Direct) | 18.2 | 27.3 | 41.0 | 32.1 | 39.3 | 52.0 | 10.3 | 23.2 | 39.0 | 27.8 | 44.4 | 76.0 | 22.2 | 27.8 | 63.0 | 11.1 | 22.2 | 59.0 | 20.3 |
| Codex (ICL) | 36.4 | 36.4 | 41.0 | 28.6 | 39.3 | 50.0 | 20.0 | 30.0 | 42.8 | 61.1 | 72.2 | 91.0 | 22.2 | 33.3 | 61.0 | 16.7 | 38.9 | 67.0 | 30.8 |
| OpsHarness (no-evolve) | 27.2 | 45.5 | 46.0 | 35.7 | 46.4 | 61.0 | 21.4 | 28.6 | 50.0 | 44.4 | 53.3 | 76.0 | 27.8 | 50.0 | 80.0 | 33.3 | 66.7 | 72.0 | 31.6 |
| OpsHarness | 45.5 | 63.6 | 68.0 | 46.4 | 50.0 | 64.0 | 28.6 | 42.9 | 60.0 | 53.3 | 73.3 | 80.0 | 50.0 | 72.2 | 90.0 | 66.7 | 77.8 | 89.0 | 48.4 |
04 — Usage preview
One harness.
Two agent interfaces.
OpsHarness keeps the same skills, tools, knowledge, and ground-truth guard across Codex and Claude Code. Only the invocation syntax changes.
The commands below preview the public workflow described in the paper and implementation.
Codex
# Adapt to a telemetry system
$setup
# Diagnose an incident
$diagnose "2021-03-04 18:00 to 18:30,
service latency spike"
# Learn from feedback
$evaluate <sessions> "/path/to/ground-truth"
$evolve "last 10 diagnoses"Claude Code
# Adapt to a telemetry system
/ops-harness:setup
# Diagnose an incident
/ops-harness:diagnose "2021-03-04 18:00 to 18:30,
service latency spike"
# Learn from feedback
/ops-harness:evaluate <sessions> "/path/to/ground-truth"
/ops-harness:evolve "last 10 diagnoses"05 — Paper
Abstract
Automated root cause analysis (RCA) with large language models (LLMs) has drawn growing attention. Today, SREs typically automate RCA with LLMs in one of two ways: directly using a general-purpose agent (e.g., Codex or Claude Code) for diagnosis, or building a specialized RCA agent from scratch. As mainstream general agents grow more capable and iterate quickly, our quantitative study finds that the former now often surpasses the latter.
Its accuracy, however, still falls short of production needs, and this gap stems mainly from the external adaptation layer outside the agent's general capabilities, namely the harness. We therefore argue that LLM-based RCA should focus on this external harness, reusing the strong general capabilities of a modern agent rather than rebuilding an agent from scratch.
A key capability of such a harness is to self-evolve, accumulating system-specific experience from past diagnoses so that it gets better the more it is used. We introduce OpsHarness, a self-evolving RCA harness that turns diagnosis experience into reusable expertise. Its data plane combines layered operational knowledge with an idea-card tool library, while its control plane coordinates setup, diagnosis, evolution, and verification.
During evolution, OpsHarness contrasts successful and failed trajectories, converts their evidence into atomic proposals, and admits updates only through a dual-gate verification process designed to prevent overfitting and regression. Across two public benchmarks and an industrial deployment, OpsHarness achieves 59.0% top-1 accuracy, improving over a bare general agent by 63.4% and over baseline RCA agents by 4.02×.
06 — Cite the work
BibTeX
If OpsHarness is useful in your research, please cite the arXiv preprint.
@misc{huang2026opsharness,
title={From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis},
author={Haiyu Huang and Jiewei Lyu and Zhihan Jiang and Jinyang Liu and Xiao He and Tieying Zhang and Wu Xiang and Michael R. Lyu},
year={2026},
eprint={2608.25661},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2608.25661}
}