[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94295":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":9,"languages":9,"totalLinesOfCode":9,"stars":10,"forks":11,"watchers":12,"openIssues":13,"contributorsCount":14,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":15,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":16,"rankGlobal":9,"rankLanguage":9,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":20,"hasPages":18,"topics":21,"createdAt":9,"pushedAt":9,"updatedAt":22,"readmeContent":23,"aiSummary":24,"trendingCount":14,"starSnapshotCount":14,"syncStatus":12,"lastSyncTime":25,"discoverSource":26},94295,"Qwen-CUA","xlang-ai\u002FQwen-CUA","xlang-ai","Qwen-CUA: Native Computer Use for (Almost) Everything — a screenshot-driven agent that operates computers with keyboard and mouse, jointly developed by the Qwen Team and XLang Lab.",null,140,7,2,1,0,34,46.11,"Apache License 2.0",false,"main",true,[],"2026-08-24 04:01:21","\u003Cdiv align=\"center\">\n  \u003Cimg src=\".\u002Fassets\u002Freadme\u002Fqwen-cua-mascot.png\" width=\"156\" alt=\"Qwen-CUA mascot holding a mouse pointer and keyboard\">\n  \u003Ch1>Qwen-CUA: Native Computer Use for (almost) Everything\u003C\u002Fh1>\n  \u003Cp>\n    A Qwen-based computer-use model and agent that sees screenshots,\u003Cbr>\n    reasons over visible state, and acts through native keyboard and mouse events.\n  \u003C\u002Fp>\n  \u003Cp>\n    \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.02352\">\u003Cimg alt=\"arXiv\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-2608.02352-B31B1B?style=flat-square\">\u003C\u002Fa>\n    \u003Ca href=\".\u002Fpaper\u002FQwen-CUA.pdf\">\u003Cimg alt=\"Paper\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPaper-PDF-7457D6?style=flat-square\">\u003C\u002Fa>\n    \u003Ca href=\".\u002Fdemo\u002FREADME.md\">\u003Cimg alt=\"Demo\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDemo-Run%20locally-3A8D7C?style=flat-square\">\u003C\u002Fa>\n    \u003Ca href=\".\u002FLICENSE\">\u003Cimg alt=\"License\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache--2.0-4B5563?style=flat-square\">\u003C\u002Fa>\n  \u003C\u002Fp>\n\u003C\u002Fdiv>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\".\u002Fpaper\u002FQwen-CUA.pdf\">\n    \u003Cimg src=\".\u002Fassets\u002Freadme\u002Fmain-results.png\" width=\"100%\" alt=\"Qwen-CUA results across eight computer-use benchmarks\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Csub>Eight benchmark families, from desktop control and long-horizon work to web interaction and adversarial robustness. Click the figure for the technical report.\u003C\u002Fsub>\u003C\u002Fp>\n\n\u003Ctable>\n  \u003Ctr>\n    \u003Ctd align=\"center\" width=\"33%\">\u003Cstrong>86.2\u003C\u002Fstrong>\u003Cbr>\u003Csub>OSWorld-Verified\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd align=\"center\" width=\"33%\">\u003Cstrong>20 images \u002F turn\u003C\u002Fstrong>\u003Cbr>\u003Csub>active visual history\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd align=\"center\" width=\"33%\">\u003Cstrong>~100k vCPUs\u003C\u002Fstrong>\u003Cbr>\u003Csub>rollout infrastructure\u003C\u002Fsub>\u003C\u002Ftd>\n  \u003C\u002Ftr>\n\u003C\u002Ftable>\n\n## One model. One native interface. Almost any software.\n\nQwen-CUA operates from pixels rather than hidden application state. It receives the same visual evidence available to a person and produces actions in a shared keyboard-and-mouse space—without DOM trees, accessibility metadata, shell access, or task-specific APIs.\n\n| **Qwen-CUA model** | **Agent runtime** |\n| --- | --- |\n| Understands screenshots and instructions, tracks progress, reasons about the visible interface, and proposes grounded native actions. | Captures observations, manages multimodal history, validates and executes actions, requests approval, and preserves replay evidence. |\n\n## Remember what matters\n\nLong computer-use trajectories accumulate image-heavy context quickly. Qwen-CUA keeps a larger active visual history, then folds older screenshots in blocks so the agent can preserve task state while reusing a stable prefix.\n\n\u003Ctable>\n  \u003Ctr>\n    \u003Ctd width=\"46%\" valign=\"top\">\n      \u003Ca href=\".\u002Fassets\u002Freadme\u002Fvisual-history.png\">\u003Cimg src=\".\u002Fassets\u002Freadme\u002Fvisual-history.png\" width=\"100%\" alt=\"Scaling active visual history to 20 screenshots\">\u003C\u002Fa>\n      \u003Cp align=\"center\">\u003Cstrong>More visual memory\u003C\u002Fstrong>\u003Cbr>\u003Csub>Twenty recent screenshots stay active.\u003C\u002Fsub>\u003C\u002Fp>\n    \u003C\u002Ftd>\n    \u003Ctd width=\"54%\" valign=\"top\">\n      \u003Ca href=\".\u002Fassets\u002Freadme\u002Fcontext-folding.png\">\u003Cimg src=\".\u002Fassets\u002Freadme\u002Fcontext-folding.png\" width=\"100%\" alt=\"Blockwise visual prefix folding for stable cache reuse\">\u003C\u002Fa>\n      \u003Cp align=\"center\">\u003Cstrong>Stable long-horizon context\u003C\u002Fstrong>\u003Cbr>\u003Csub>Blockwise folding bounds growth and improves prefix reuse.\u003C\u002Fsub>\u003C\u002Fp>\n    \u003C\u002Ftd>\n  \u003C\u002Ftr>\n\u003C\u002Ftable>\n\n## Learn from verifiable experience\n\nQwen-CUA scales computer-use training along two axes: broader verifiable tasks and rollout capacity across model generations, followed by successive training runs in which the current policy exposes unresolved queries and weak domains. Those diagnostics refresh both the supervised data mixture and the verifiable RL task distribution before the next run.\n\n\u003Ctable>\n  \u003Ctr>\n    \u003Ctd width=\"50%\" valign=\"top\">\n      \u003Ca href=\".\u002Fassets\u002Freadme\u002Ftraining-scale.png\">\u003Cimg src=\".\u002Fassets\u002Freadme\u002Ftraining-scale.png\" width=\"100%\" alt=\"Computer-use performance scaling with model generation, verifiable tasks, and rollout infrastructure\">\u003C\u002Fa>\n      \u003Cp align=\"center\">\u003Cstrong>Scaling training resources\u003C\u002Fstrong>\u003Cbr>\u003Csub>From 1k to 40k tasks and nearly 100k vCPUs.\u003C\u002Fsub>\u003C\u002Fp>\n    \u003C\u002Ftd>\n    \u003Ctd width=\"50%\" valign=\"top\">\n      \u003Ca href=\".\u002Fassets\u002Freadme\u002Fsuccessive-training-runs.png\">\u003Cimg src=\".\u002Fassets\u002Freadme\u002Fsuccessive-training-runs.png\" width=\"100%\" alt=\"Evaluation scores across successive Qwen-CUA training runs on OSWorld-Verified, OSWorld 2.0, and ScienceBoard\">\u003C\u002Fa>\n      \u003Cp align=\"center\">\u003Cstrong>Successive training runs\u003C\u002Fstrong>\u003Cbr>\u003Csub>SFT data and verifiable RL tasks are refreshed between runs.\u003C\u002Fsub>\u003C\u002Fp>\n    \u003C\u002Ftd>\n  \u003C\u002Ftr>\n\u003C\u002Ftable>\n\nEach SFT run starts from the same mid-training checkpoint instead of continually fine-tuning the previous agent checkpoint. The resulting SFT model recalibrates the RL pool with eight trial rollouts per task, retaining queries that are neither unreachable nor already saturated. Because teacher policies, data mixtures, domain coverage, and task distributions change between runs, the plotted lines connect development checkpoints rather than measuring controlled convergence or scaling.\n\n## Scale the model, scale the ceiling\n\nThe same computer-use recipe extends from **Qwen-CUA (397B-A17B)** to **Qwen-CUA-Max (>1T)**. The larger model reaches **87.6 on OSWorld-Verified** and improves both binary and partial-credit performance on OSWorld 2.0.\n\n\u003Cp align=\"center\">\n  \u003Ca href=\".\u002Fassets\u002Freadme\u002Fmodel-scaling.png\">\n    \u003Cimg src=\".\u002Fassets\u002Freadme\u002Fmodel-scaling.png\" width=\"94%\" alt=\"Qwen-CUA and Qwen-CUA-Max capacity scaling results\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n## From benchmark to real workflows\n\nQwen-CUA is designed for the interaction loop users actually see: inspect the page, plan, act, recover when the interface changes, and verify the result.\n\n\u003Cp align=\"center\">\n  \u003Ca href=\".\u002Fassets\u002Freadme\u002Fchrome-showcase.png\">\n    \u003Cimg src=\".\u002Fassets\u002Freadme\u002Fchrome-showcase.png\" width=\"100%\" alt=\"Qwen computer-use agent operating in a Chrome side panel\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\nThe self-contained [`demo\u002F`](.\u002Fdemo\u002FREADME.md) turns that loop into a local, browser-first reference agent:\n\n| **Operator console** | **Safety gates** | **Replayable runs** |\n| --- | --- | --- |\n| Inspect screenshots, actions, approvals, and raw model responses. | Pause sensitive actions and isolate every Playwright browser session. | Save events, screenshots, downloads, and deterministic verification evidence. |\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002Fxlang-ai\u002FQwen-CUA.git\ncd Qwen-CUA\u002Fdemo\ncp .env.example .env\n```\n\nThen follow the [demo quick start](.\u002Fdemo\u002FREADME.md#native-quick-start) to connect an OpenAI-compatible multimodal endpoint and launch the operator console.\n\n## Repository\n\n```text\nQwen-CUA\u002F\n├── paper\u002F          # Technical report\n├── demo\u002F           # Runnable browser-agent reference implementation\n├── assets\u002Freadme\u002F  # Figures used in this overview\n├── LICENSE\n└── README.md\n```\n\n> [!NOTE]\n> This release contains the technical report and reference demo. Model weights are not included in the repository.\n\n## Citation\n\nIf you find Qwen-CUA useful in your work, please cite our technical report:\n\n```bibtex\n@misc{lu2026qwencuanativecomputeruse,\n      title={Qwen-CUA: Native Computer Use for (almost) Everything},\n      author={Dunjie Lu and Shuai Bai and Tianyi Bai and Sicheng Fan and Chang Gao and Jian Guan and Feng Hu and Mianqiu Huang and Xingyang Huang and Yizhen Jiang and Yuheng Jing and Dehui Kong and Ning Li and Dayiheng Liu and Shixuan Liu and Zheng Liu and Que Shen and Bowen Wang and Junli Wang and Chencan Wu and Rui Xie and Tianbao Xie and Zhihui Xie and Haiyang Xu and An Yang and Tao Yu and Wenzhen Yuan and Xi Zhang and Zhenru Zhang and Mingkang Zhu and Zhaoqing Zhu and Yizhong Cao and Kai Dang and Binyuan Hui and Kaixin Li and Junyang Lin and Haiquan Wang and Zekun Wang and Yiheng Xu and Fan Yan and Mengqi Yuan and Danyang Zhang and Jiajun Zhang and Zhipeng Zhang and Fan Zhou and Fan Zhou},\n      year={2026},\n      eprint={2608.02352},\n      archivePrefix={arXiv},\n      primaryClass={cs.LG},\n      url={https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.02352},\n}\n```\n\n## Safety\n\nComputer-use agents can make mistakes, encounter prompt injection, and trigger consequential interface actions. Use isolated browser contexts, avoid authenticated or high-stakes workflows, and require human approval for sensitive operations. A model declaring success is not proof that the intended real-world outcome was achieved.\n\n## License\n\nApache-2.0. See [LICENSE](.\u002FLICENSE) and [NOTICE](.\u002FNOTICE).\n","Qwen-CUA 是一个基于通义千问（Qwen）的原生计算机操作智能体，通过分析屏幕截图理解界面状态，并直接驱动键盘和鼠标执行真实操作。其核心特点是纯像素输入、无依赖（不依赖DOM、无障碍API或后台接口）、支持长程任务记忆与20帧活跃视觉历史管理，结合多模态推理与动作接地机制。适用于自动化GUI交互场景，如跨应用流程执行、软件测试、无障碍辅助操作及教育演示等需真实人机交互界面的领域。","2026-08-05 02:30:08","CREATED_QUERY"]