ClawGUI
0.2.0 · GitHub
More about this app
On-device GUI-agent runner deploying the full ClawGUI brain stack on one phone controlled via Shizuku.
|
ClawGUI-Agent controls a real phone via natural language |
ClawGUI-RL trains a GUI agent with online reinforcement learning |
News
- π [2026/4/14] Our paper is available on arXiv: ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents.
- π₯ [2026/4/13] ClawGUI is released β train with ClawGUI-RL (GiGPO), evaluate with ClawGUI-Eval, deploy with ClawGUI-Agent. ClawGUI-2B, a 2B agent trained end-to-end with this pipeline, hits 17.1 MobileWorld SR vs. the 11.1 baseline. See Quick Start.
Table of Contents
π‘ Overview
ClawGUI is a research framework for GUI agents, covering the complete lifecycle from online RL training and standardized evaluation to real-device deployment.
Building a capable GUI agent involves three tightly coupled problems that are rarely solved together: you need an environment to train the agent online, rigorous benchmarks to measure what it has learned, and a production system to deploy it on real devices. ClawGUI addresses all three.
| Module | Role |
|---|---|
| π ClawGUI-RL | Build β Train GUI agents online with scalable RL: parallel Docker environments, real Android devices, and GiGPO+PRM for fine-grained step-level rewards |
| π ClawGUI-Eval | Evaluate β Measure what the agent has learned: 6 benchmarks, 11+ models, 95.8% faithful reproduction of official results |
| π€ ClawGUI-Agent | Deploy β Use GUI agents in the real world: control mobile devices via natural language through 12+ chat platforms, with one-command evaluation built in |
| π§© ClawGUI-Skills | Self-evolving skills β Training-free skill evolution proposed and validated in our paper: structured packages, retrieval, failure diagnosis, restricted revision, and reuse |
| π± ClawGUI-APP | On-Device Deploy β Run the full brain + GUI agent stack directly on one Android phone, no desktop coordinator needed, powered by Shizuku |
| π ClawGUI-2B | End-to-end validation: trained entirely with ClawGUI-RL and GiGPO, achieving 17.1 MobileWorld SR vs. the 11.1 baseline |
ποΈ Architecture
π Quick Start
git clone https://github.com/ZJU-REAL/ClawGUI.git
cd ClawGUI
Each module is independent with its own environment. Click into each one for full installation and usage instructions.
π ClawGUI-RL β Build
π
clawgui-rl/Β· π Full Documentation
ClawGUI-RL trains GUI agents with online reinforcement learning. It runs dozens of Docker-based Android emulators in parallel or trains directly on physical devices β and replaces standard GRPO with GiGPO+PRM for fine-grained step-level rewards that drive stronger policy learning.
- Parallel multi-environment β Dozens of Docker-based virtual Android environments simultaneously
- Real-device training β Physical or cloud Android phones with the same API
- GiGPO + PRM β Fine-grained step-level reward for better policy optimization than standard GRPO
- Spare server rotation β Automatic failover keeps training running without interruption
- Episode visualization β Record and replay any training trajectory
β Get started with ClawGUI-RL
π ClawGUI-Eval β Evaluate
π
clawgui-eval/Β· π Full Documentation Β· π€ Dataset Β· π€ ModelScope
ClawGUI-Eval gives GUI grounding research a reliable measurement baseline. Its three-stage Infer β Judge β Metric pipeline covers 6 benchmarks and 11+ models, with a 95.8% reproduction rate against official results β so numbers across papers are actually comparable.
- 6 benchmarks β ScreenSpot-Pro, ScreenSpot-V2, UIVision, MMBench-GUI, OSWorld-G, AndroidControl
- 11+ models β Qwen3-VL, Qwen2.5-VL, UI-TARS, MAI-UI, GUI-G2, UI-Venus, Gemini, Seed 1.8, and more
- Dual backend β Local GPU (
transformers) or remote API (OpenAI-compatible) - Multi-GPU & multi-thread β Parallel inference with automatic resume
- ClawGUI-Agent integration β Pair with ClawGUI-Agent to run the full pipeline via natural language
β Get started with ClawGUI-Eval
π€ ClawGUI-Agent β Deploy
π
clawgui-agent/Β· π Full Documentation Β· δΈζ
ClawGUI-Agent closes the loop from training to production. Built on OpenClaw and powered by nanobot, it lets you control Android, HarmonyOS, or iOS devices with natural language from 12+ chat platforms β and trigger the full ClawGUI-Eval benchmark pipeline with a single sentence, no scripts required.
- Cross-platform β Android (ADB), HarmonyOS (HDC), iOS (XCTest)
- Multi-model β AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS via OpenAI-compatible API
- One-command evaluation β Say "benchmark qwen3vl on screenspot-pro" and it handles env check β multi-GPU inference β judging β metrics β result comparison
- Personalized memory β Automatically learns user preferences and injects context across tasks
- Episode recording β Every task saved as structured episodes for replay and dataset building
- Web UI β Gradio interface for device management, task execution, and memory inspection
β Get started with ClawGUI-Agent
π§© ClawGUI-Skills β Self-Evolving Skills
π
clawgui-skills/Β· π Full Documentation Β· δΈζ
ClawGUI-Skills implements the training-free self-evolving GUI skill architecture proposed and validated in our paper βReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents.β It stores procedural task knowledge as structured skill packages and lets PhoneAgent retrieve, inject, diagnose, and revise them on demand.
- Four modes β
off,trace,reuse, andevolve; disabled by default to avoid extra context cost - Structured packages β
meta_info.json,plan.md,backup.md,recover.md, andfailure_examples/ - Instant revision β failed runs are diagnosed by an isolated verifier and mapped to targeted skill-file edits
- Visual inspection β the Web UI shows matched skill name,
skill_id, injected context, revisions, and failure examples
β Get started with ClawGUI-Skills
π± ClawGUI-APP β On-Device Deploy
π
clawgui-app/Β· π Setup Guide
ClawGUI-APP runs the full ClawGUI "brain + GUI agent" stack directly on one Android phone, removing the old split architecture where a desktop host orchestrates tasks and the phone only executes them. Built on Shizuku for high-privilege, non-root device control.
- Phone-only workflow β No desktop coordinator required; a device with Shizuku is enough
- Two-agent design β Brain LLM handles planning and tool orchestration, phone agent handles screen understanding and actions
- Multi-model support β AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS and more via OpenAI-compatible API
- Voice input (STT) β Tap-to-record microphone with OpenAI-compatible speech-to-text transcription (SiliconFlow, Groq Whisper, etc.)
- Conversation + automation β Sessions, long-term memory, external channels (Feishu), and trace replay
- Built for real usage β Floating overlay status, built-in IME, session persistence, and diagnostics
π― Roadmap
- ClawGUI-Agent β GUI agent framework for phone control and evaluation via natural language
- ClawGUI-RL β Scalable mobile online RL training infrastructure with GiGPO + PRM
- ClawGUI-Eval β Standardized GUI grounding evaluation suite with 6 benchmarks and 95%+ reproduction rate
- ClawGUI-2B β 2B GUI agent trained with GiGPO, achieving 17.1 MobileWorld SR (vs. 11.1 baseline)
- On-device ClawGUI-Agent (ClawGUI-APP) β Deploy ClawGUI-Agent directly on real phones β no desktop coordinator, paving the way for fully on-device inference (brain/VLM still served via cloud API today)
- Desktop Online RL β Extend ClawGUI-RL to desktop environments for online reinforcement learning
- Web Online RL β Extend ClawGUI-RL to web environments for online reinforcement learning
- More Skills for ClawGUI-Agent β Add more pluggable skills to expand ClawGUI-Agent's capabilities
- Hybrid CLI & GUI Mechanism β Explore hybrid interaction combining command-line and GUI operations
- Real-time RL β Integrate real-time reinforcement learning based on the OPD algorithm for ClawGUI-RL and ClawGUI-Agent
π€ Contributing
We welcome contributions of all kinds β new model support, new RL environments, bug fixes, and documentation improvements. See CONTRIBUTING.md for how to get started, module-specific guidelines, and PR requirements.
π Acknowledgements
ClawGUI is built upon the following excellent open-source projects. We sincerely thank their contributors:
License
This project is licensed under the Apache License 2.0.
π Citation
If you find ClawGUI useful in your research, please consider citing our paper:
@article{tang2026clawgui,
title={ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents},
author={Tang, Fei and Lu, Zhiqiong and Zhang, Boxuan and Lu, Weiming and Xiao, Jun and Zhuang, Yueting and Shen, Yongliang},
journal={arXiv preprint arXiv:2604.11784},
year={2026}
}
Star History
How Shizuku is used
Can tap, swipe, type, launch apps, take screenshots and manage input methods via `sh -c` shell service through Shizuku.
This is an AI-assisted analysis of Shizuku-related usages in the app's public source code. It is best effort, so it may not catch every single usage.
How this app uses Shizuku
ClawGUI uses Shizuku to run its on-device GUI agent, letting an AI model see the screen and operate the phone with shell privileges.
- Tap and swipe screen: the agent taps, double taps, long presses and swipes at coordinates from its visual plan using the
input tapandinput swipeshell commands. - Type and clear text: the agent enters the text from its next step with the
input textshell command, clipboard paste plus paste keyevent, or a broadcast to its own input method, and clears fields with keyevents or a clear-text broadcast. - Press system keys: the agent goes back, home or confirms with enter using the
input keyeventshell command. - Launch and open apps: the agent starts apps by package name using package listing, launcher and activity start shell commands, and can open links or intents with an activity start command.
- Capture screenshots: each automation step captures the current screen with the
screencapshell command for the model to see, plus display size and orientation queries to map coordinates. - Manage input method: the agent checks, enables and switches the active keyboard, including its own input method, with settings and input method shell commands.
Android APIs or commands used
input tapinput swipeinput textinput keyeventam broadcastam startpm list packagesmonkey -pcmd package resolve-activitydumpsys activity activitiesdumpsys packagedumpsys window displaysscreencap -pchmod 666wm sizesettings get secure default_input_methodime listime enableime setime reset
Notable details
If the Shizuku shell service is not bound, the controller tries a wireless ADB session first and then a local unprivileged shell, so many touch, key and launch operations still attempt to run but privileged ones may fail.
Changelog
What's new for version 0.2.0
Fixes
- Feishu GUI runner now drives step-by-step. The chat bubble updates with the current action on every step instead of staying frozen on "ζ£ε¨ζ§θ‘δ»»ε‘β¦". Each step has a 5-minute timeout so a wedged LLM call can't hang the loop forever.
- Clean error surfacing for Feishu replies. SDK errors are decoded to
code: messageinstead of the previousError@<hash>. Pre-flight failures (missing Vision creds, no device-control auth) now reply to Feishu with a human-readable reason instead of silently bailing. Image-reply failures land in Settings β ε€ι¨ιι β θ°θ―ζ₯εΏ so users can see them without logcat. @_user_Nmention stripping. Was hardcoded to@_user_1; now regex-matches@_(user|chat)_Nso@bot εζεεreaches PhoneAgent as justεζεε.- Live screenshot fallback when image-reply is enabled but trace recording is off β users still get a final-state image instead of silent nothing.
- IME enabled check via the Android Framework API instead of
ime list -s. Users who'd already enabled "ClawGUI Input" via system settings would see "ζͺε―η¨" in our panel because the shell command needed Shizuku/wadb auth to succeed. Settings β θΎε ₯ζ³ also dropped the manual switcher button β agent handles the switching itself.
Chore
- Repo cleanup: retired the v1
clawgui-app/(Java client) and promotedclawgui-app-ng/into the naturalclawgui-app/path. Android applicationId stayscom.clawgui.ngso existing v0.3.0 installs upgrade in place.
Permissions
16 permissions requested