[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94287":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":10,"rankLanguage":10,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":19,"topics":21,"createdAt":10,"pushedAt":10,"updatedAt":38,"readmeContent":39,"aiSummary":40,"trendingCount":15,"starSnapshotCount":15,"syncStatus":41,"lastSyncTime":42,"discoverSource":43},94287,"LongHorizon-Harness","AMAP-ML\u002FLongHorizon-Harness","AMAP-ML","The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code \u002F Codex \u002F OpenClaw integration.","https:\u002F\u002Flh-harness.pages.dev",null,"Python",532,65,104,0,385,59.46,"MIT License",false,"main",[22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37],"agent","claude","claude-code","claude-plugin","cli","codex","codex-desktop","codex-plugin","cua","gui","harness","long-horizon","long-horizon-agents","longhorizon-harness","loop","loop-engineering","2026-08-25 04:01:21","\u003Cdiv align=\"center\">\n\n# LongHorizon-Harness\n\n### Advancing Long-Horizon Agents for Real-World Tasks\n\n**Operate the whole computer like a human. Work across desktop apps and the command line for dozens of hours.**\n\n**No state drift. Verifiable progress. Complex tasks carried through to completion.**\n\n\u003Cp align=\"center\">\n\u003Ca href=\"https:\u002F\u002Flh-harness.pages.dev\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🌐-Website-1f6feb.svg?style=flat-square\" alt=\"Website\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.01964\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-2608.01964-b31b1b.svg?style=flat-square\" alt=\"arXiv 2608.01964\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FAMAP-ML\u002FLongHorizon-Harness\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FGitHub-Repository-181717.svg?style=flat-square&logo=github&logoColor=white\" alt=\"GitHub repository\" \u002F>\u003C\u002Fa>\n\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤗-Trajectory_Coming_Soon-ffce00.svg?style=flat-square\" alt=\"Hugging Face trajectory\" \u002F>\n\u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2608.01964\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤗_Daily_Papers-2608.01964-ff8800.svg?style=flat-square\" alt=\"Hugging Face Daily Papers\" \u002F>\u003C\u002Fa>\n\u003Ca href=\".\u002FLICENSE\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-MIT-2ea44f.svg?style=flat-square\" alt=\"MIT License\" \u002F>\u003C\u002Fa>\n\u003C\u002Fp>\n\n[![Python](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-≥3.10-blue?logo=python&logoColor=white)](https:\u002F\u002Fwww.python.org\u002F)\n[![Agents](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fbackends-Claude%20Code%20|%20Codex%20|%20OpenClaw-8A2BE2)](#any-model-any-agent-backend)\n[![Benchmarks](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fbenchmarks-WeaveBench%20|%20OSWorld%202.0%20|%20Terminal--Bench%202.1-orange)](#hundreds-of-real-tasks-measured-gains)\n\n[Usage](#one-command-full-visibility) · [What You Get](#desktop-apps-and-cli-one-continuous-task) · [How It Works](#three-roles-one-trusted-state) · [Results](#hundreds-of-real-tasks-measured-gains) · [Project Website](https:\u002F\u002Flh-harness.pages.dev) · [简体中文](README.zh-CN.md)\n\n\u003Cbr>\n\u003Cimg src=\"assets\u002Fquickstart.gif\" alt=\"Install and run LongHorizon-Harness from the command line\" width=\"720\">\n\n\u003C\u002Fdiv>\n\n> **The model determines what an agent can do in one round. LongHorizon-Harness determines whether that work can be verified, preserved, and continued until the task is actually complete.**\n\n**Works with Claude Code, Codex, and OpenClaw. One-command install, ready to run.**\n\nLongHorizon-Harness is an execution, state-management, and result-verification system for long-horizon tasks. It does not train a new model or replace an existing agent. It runs on top of systems such as Codex and Claude Code, helping agents operate autonomously in real computer environments for extended periods and continuously move complex tasks forward.\n\n## Video Demo\n\nhttps:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fca8b77ce-9220-4d85-a272-b346009b2454\n\n\u003Cp align=\"center\">\u003Ca href=\"assets\u002Fpromotional_video_1440p.mp4\">\u003Cstrong>Open the promotional video (1440p MP4)\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\n\n## Three roles. One trusted state.\n\nLongHorizon-Harness separates planning, execution, and verification so that one growing context is not responsible for everything.\n\n| | Role | One responsibility |\n|---|---|---|\n| 🧭 | **Manager** | Maintains the original goal, verified progress, and next step |\n| ⚡ | **Executor** | Starts each round with a fresh context and focuses on one clearly defined task |\n| 🔍 | **Auditor** | Independently inspects files, interfaces, logs, and tests in the real environment |\n\nOnly results that pass independent verification enter persistent task state. Even when the context is refreshed, an action fails, or a deliverable does not pass inspection, the system retains previously verified progress and continues from what remains.\n\n## Desktop apps and CLI. One continuous task.\n\nLongHorizon-Harness supports both GUI and CLI workflows.\n\n| 🖥️ Operate the desktop | ⌨️ Work in the terminal |\n|---|---|\n| 🌐 Click, type, scroll, and browse | 💻 Write and modify code |\n| 📊 Operate spreadsheets | ▶️ Run commands and scripts |\n| 📄 Edit documents | 📦 Install dependencies and environments |\n| 🎨 Use design software | 🔧 Configure and debug systems |\n| 🧊 Operate 3D tools | 📁 Process files and data |\n\nOne task can begin in a browser, move to the command line for data processing, continue in desktop software to produce an artifact, and return to the terminal for validation or debugging. The goal, progress, and evidence remain under the same state-management system throughout.\n\n\u003Cdetails>\n\u003Csummary>\u003Cstrong>🖱️ Connect a computer-use MCP server\u003C\u002Fstrong>\u003C\u002Fsummary>\n\nGUI interaction is supplied through a compatible external computer-use MCP server. LongHorizon-Harness does not bundle or enable a specific computer-use implementation by default.\n\n```bash\nlh-harness run --task @task.md --agent claude_code \\\n  --mcp-config \u002Fpath\u002Fto\u002Fyour\u002Fmcp.json \\\n  --mcp-add-dir \u002Fpath\u002Fto\u002Fyour\u002Fmcp\u002Ffiles\n```\n\nYou can also use `LH_HARNESS_CLAUDECODE_MCP_CONFIG` and `LH_HARNESS_CLAUDECODE_ADD_DIRS`. When no configuration is supplied, the Claude Code adapter does not add MCP arguments.\n\n\u003C\u002Fdetails>\n\n## Any model. Any agent backend.\n\nLongHorizon-Harness is not tied to a specific model or agent backend. Existing models and agents connect through configuration without changing their original workflows.\n\n| | Layer | Supported choices |\n|---|---|---|\n| 🧠 | **Models** | Claude, GPT, Qwen, and other models exposed by an agent backend |\n| 🤖 | **Agent backends** | Claude Code, Codex CLI, OpenClaw, and custom `AgentAdapter` implementations |\n| 🎛️ | **Role assignment** | The Manager, Executor, and Auditor can each use a different model or backend |\n| 🖥️ | **Execution environments** | Local, `ssh:\u002F\u002Fuser@host:port`, and `docker:\u002F\u002Fcontainer` |\n\nA lightweight `AgentAdapter` preserves each agent's native execution loop while LongHorizon-Harness coordinates role boundaries, verified task state, and cross-round progress around it.\n\nUse one model for all three roles, or combine different models and backends to balance quality, speed, and cost.\n\n## Hundreds of real tasks. Measured gains.\n\nLongHorizon-Harness is not demonstrated only on a handful of carefully selected success cases.\n\nWe ran it on hundreds of complex tasks across GUI, CLI, and mixed computer environments:\n\n| Task domain | What the tasks involve |\n|---|---|\n| 🌐 **Web Frontend** | Developing, fixing, and validating websites and web applications through browser interaction, developer tools, and code changes |\n| 📊 **Data Analysis & Visualization** | Processing data, producing charts and dashboards, and checking analytical results and visual deliverables |\n| 🛠️ **Operations & Debugging** | Investigating logs, networks, performance, and service failures; configuring, diagnosing, and repairing systems |\n| 🎨 **Design & Image Processing** | Editing visual assets, matching design references, processing images, and verifying final visual quality |\n| 🎮 **Games & Interaction** | Building, operating, and debugging games or interactive applications; checking interaction logic and runtime behavior |\n| 📄 **Documents & Presentations** | Editing documents and slide decks, including content, formatting, references, layout, and final delivery |\n| 🧊 **Spatial Reasoning** | Completing tasks involving spatial relationships, geometry, precise placement, and 3D operations |\n| 🖥️ **Desktop & System Settings** | Operating desktop applications, files, and system settings across multi-application workflows |\n| 🔬 **Research & Education** | Completing literature research, coursework, teaching materials, forms, and research-support workflows |\n| 🎬 **Creative Production** | Producing presentations, video, audio, and other media while coordinating assets across tools |\n| ⚙️ **Engineering & Computing** | Using CAD, EDA, scientific software, development tools, and cloud or DevOps toolchains |\n| 🎫 **Personal Services** | Handling event ticketing, everyday services, games, and visual-search workflows |\n| 🏛️ **Administration & Compliance** | Completing office, legal, policy-sensitive form, institutional, and safety-aware submission workflows |\n| 💼 **Business & Finance** | Handling market analysis, procurement, loans, sales, reimbursements, and cross-application enterprise workflows |\n| 🏥 **Healthcare** | Completing medical quality-control, insurance, immunization, and structured health-form workflows |\n\n### Same model. Same execution backend. Only the harness changes.\n\n\u003Ctable>\n\u003Ctr>\n\u003Ctd align=\"center\" width=\"33%\">\n\u003Ch2>~50% → ~80%\u003C\u002Fh2>\n\u003Cstrong>GUI + CLI completion\u003C\u002Fstrong>\u003Cbr>\n\u003Csub>WeaveBench\u003C\u002Fsub>\n\u003C\u002Ftd>\n\u003Ctd align=\"center\" width=\"33%\">\n\u003Ch2>3×\u003C\u002Fh2>\n\u003Cstrong>Full desktop-task completion\u003C\u002Fstrong>\u003Cbr>\n\u003Csub>OSWorld 2.0\u003C\u002Fsub>\n\u003C\u002Ftd>\n\u003Ctd align=\"center\" width=\"33%\">\n\u003Ch2>69.7% → 77.2%\u003C\u002Fh2>\n\u003Cstrong>Code + CLI success\u003C\u002Fstrong>\u003Cbr>\n\u003Csub>Terminal-Bench 2.1 · 24% fewer tokens\u003C\u002Fsub>\n\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftable>\n\n\u003Cdiv align=\"center\">\n\u003Cimg src=\"assets\u002Fharness_perf.png\" alt=\"Performance gains across benchmarks and backbones\" width=\"72%\">\n\u003C\u002Fdiv>\n\n\u003Cdetails>\n\u003Csummary>\u003Cstrong>📊 Full benchmark results and experimental settings\u003C\u002Fstrong>\u003C\u002Fsummary>\n\n| Benchmark | Metric | Claude Code | **LongHorizon-Harness** | Gain |\n|---|---|:-:|:-:|:-:|\n| **WeaveBench** (114 tasks) | PassRate | 51.8 | **80.7** | **+28.9** |\n| **WeaveBench** | Overall | 0.702 | **0.835** | +0.133 |\n| **OSWorld 2.0** (108 tasks) | Binary | 2.8 | **8.3** | **3.0×** |\n| **OSWorld 2.0** | Partial | 21.5 | **35.2** | **+13.7** |\n| **Terminal-Bench 2.1** | Success rate | 69.7 | **77.2** | **+7.5** |\n\n\u003Csub>All rows use Qwen 3.7-Plus as the backbone and Claude Code as the execution backend.\u003C\u002Fsub>\n\n\u003C\u002Fdetails>\n\nFull result tables and case trajectories are available on the [LongHorizon-Harness project website](https:\u002F\u002Flh-harness.pages.dev).\n\n## One command. Full visibility.\n\nInstall LongHorizon-Harness:\n\n```bash\nuv tool install lh-harness\n```\n\nLongHorizon-Harness requires Python 3.10+ and at least one agent runtime: `claude`, `codex`, or `openclaw`.\n\nRun a task:\n\n```bash\nlh-harness run \\\n  --task \"Inspect the current directory and summarize its files.\"\n```\n\nRun a longer task from a file and open the Dashboard:\n\n```bash\nlh-harness run --task @task.md --dashboard\n```\n\nThe Dashboard shows every round's plan, execution result, audit evidence, and reason for rework. It also provides human gates when a task completes, becomes blocked, needs input, or fails repeatedly.\n\n| 📋 Plan | ⚡ Execution | 🔍 Audit | ♻️ Rework |\n|:---:|:---:|:---:|:---:|\n| What happens next | What the agent did | What the environment proves | Why another round is needed |\n\nEvery run is stored in an isolated `runs\u002F\u003Crun-id>\u002F` directory. The complete task state and audit trail make the agent's progress inspectable, recoverable, and reproducible.\n\n| Run record | What it preserves |\n|---|---|\n| 📋 **Task state** | Original goal, requirements, verified progress, and remaining work |\n| 🧾 **Event stream** | What happened throughout the run |\n| 🔍 **Audit reports** | Evidence and acceptance decisions for every round |\n| 🧠 **Role trajectories** | Manager, Executor, and Auditor inputs and outputs |\n| 📁 **Workspace** | Files and artifacts produced during execution |\n| ✅ **Final report** | The verified outcome of the task |\n\n\u003Cdetails>\n\u003Csummary>\u003Cstrong>⚙️ Installation alternatives and common CLI options\u003C\u002Fstrong>\u003C\u002Fsummary>\n\nInstall with `pip`:\n\n```bash\npip install lh-harness\n```\n\nDashboard commands:\n\n```bash\nlh-harness run --task @task.md --dashboard      # Monitor a live run\nlh-harness dashboard --runs-root .\u002Fruns         # Browse completed and active runs\n```\n\n| Option | Description |\n|---|---|\n| `--task` | Task text or `@task.md` |\n| `--agent` | `claude_code`, `codex`, or `openclaw` |\n| `--env` | `local`, `ssh:\u002F\u002F...`, or `docker:\u002F\u002F...` |\n| `--max-rounds` | Maximum number of Manage-Execute-Audit rounds; the CLI default is 30 |\n| `--dashboard` | Start live monitoring and human intervention |\n\n\u003C\u002Fdetails>\n\n## Evaluation Reproduction\n\n`eval\u002F` provides frozen reproduction suites for two benchmarks:\n\n| Directory | Benchmark | Description |\n|---|---|---|\n| [`eval\u002FWeaveBench-harness\u002F`](eval\u002FWeaveBench-harness\u002F) | WeaveBench (114 tasks) | Hybrid GUI+CLI tasks and a reproduction skill |\n| [`eval\u002FOSWorldv2-harness\u002F`](eval\u002FOSWorldv2-harness\u002F) | OSWorld-V2 (108 tasks) | Hybrid runner aligned with the official release |\n\nSee each directory's `README.md` or `README.zh-CN.md` for environment setup, parameters, and launch commands. The nested `cua_harness` packages are frozen compatibility copies used for evaluation; new integrations should use `src\u002Flh_harness\u002F`.\n\n## Citation\n\n```bibtex\n@article{longhorizonharness2026,\n  title={LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks},\n  author={Ziyu Ma and Hailang Huang and Shun Zou and Yong Wang and Shidong Yang and Yiming Hu and Fei Wei and XiangXiang Chu},\n  journal={arXiv preprint arXiv:2608.01964},\n  year   = {2026},\n  url    = {https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.01964}\n}\n```\n\n---\n\n\u003Cdiv align=\"center\">\n\n**Operate the whole computer. Preserve verified progress. Keep working until the task is done.**\n\n\u003C\u002Fdiv>\n","LongHorizon-Harness 是一个面向长周期任务的AI智能体执行与状态管理框架，支持在桌面应用和命令行环境中持续运行数十小时，保障任务状态不漂移、进度可验证、中断可恢复。其核心技术包括新鲜上下文执行（fresh-context execution）、持久化可信状态（durable verified state）、独立审计机制、多后端集成（原生支持 Claude Code、Codex、OpenClaw）及闭环任务追踪。适用于需跨GUI\u002FCLI协同、长时间连续执行的复杂自动化场景，如软件测试流水线、跨工具数据处理、端到端IT运维任务等。",2,"2026-08-05 02:30:07","CREATED_QUERY"]