ShizuStore

ClawGUI

ZJU-REAL

0.2.0 · GitHub

Download APK
AI agents Android 8.0+ 4 months ago Apache-2.0
105 ShizuStore
352 GitHub
1.4k Stars
31 MB Size

More about this app

On-device GUI-agent runner deploying the full ClawGUI brain stack on one phone controlled via Shizuku.

ClawGUI Logo

ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

Python 3.12 License Stars arXiv Daily Paper

HuggingFace Model ModelScope Model Project Page

English | δΈ­ζ–‡

A full-stack framework for GUI agents, covering online RL training, standardized evaluation, and deployment.

ClawGUI-Agent controls a real phone
via natural language

ClawGUI-RL trains a GUI agent with online
reinforcement learning

News

Table of Contents

πŸ’‘ Overview

ClawGUI is a research framework for GUI agents, covering the complete lifecycle from online RL training and standardized evaluation to real-device deployment.

Building a capable GUI agent involves three tightly coupled problems that are rarely solved together: you need an environment to train the agent online, rigorous benchmarks to measure what it has learned, and a production system to deploy it on real devices. ClawGUI addresses all three.

Module Role
πŸš€ ClawGUI-RL Build β€” Train GUI agents online with scalable RL: parallel Docker environments, real Android devices, and GiGPO+PRM for fine-grained step-level rewards
πŸ“Š ClawGUI-Eval Evaluate β€” Measure what the agent has learned: 6 benchmarks, 11+ models, 95.8% faithful reproduction of official results
πŸ€– ClawGUI-Agent Deploy β€” Use GUI agents in the real world: control mobile devices via natural language through 12+ chat platforms, with one-command evaluation built in
🧩 ClawGUI-Skills Self-evolving skills β€” Training-free skill evolution proposed and validated in our paper: structured packages, retrieval, failure diagnosis, restricted revision, and reuse
πŸ“± ClawGUI-APP On-Device Deploy β€” Run the full brain + GUI agent stack directly on one Android phone, no desktop coordinator needed, powered by Shizuku
πŸ† ClawGUI-2B End-to-end validation: trained entirely with ClawGUI-RL and GiGPO, achieving 17.1 MobileWorld SR vs. the 11.1 baseline

πŸ—οΈ Architecture

ClawGUI System Architecture

πŸš€ Quick Start

git clone https://github.com/ZJU-REAL/ClawGUI.git
cd ClawGUI

Each module is independent with its own environment. Click into each one for full installation and usage instructions.

πŸš€ ClawGUI-RL β€” Build

πŸ“ clawgui-rl/ Β· πŸ“– Full Documentation

ClawGUI-RL trains GUI agents with online reinforcement learning. It runs dozens of Docker-based Android emulators in parallel or trains directly on physical devices β€” and replaces standard GRPO with GiGPO+PRM for fine-grained step-level rewards that drive stronger policy learning.

  • Parallel multi-environment β€” Dozens of Docker-based virtual Android environments simultaneously
  • Real-device training β€” Physical or cloud Android phones with the same API
  • GiGPO + PRM β€” Fine-grained step-level reward for better policy optimization than standard GRPO
  • Spare server rotation β€” Automatic failover keeps training running without interruption
  • Episode visualization β€” Record and replay any training trajectory
ClawGUI-RL Architecture

β†’ Get started with ClawGUI-RL

πŸ“Š ClawGUI-Eval β€” Evaluate

πŸ“ clawgui-eval/ Β· πŸ“– Full Documentation Β· πŸ€— Dataset Β· πŸ€– ModelScope

ClawGUI-Eval gives GUI grounding research a reliable measurement baseline. Its three-stage Infer β†’ Judge β†’ Metric pipeline covers 6 benchmarks and 11+ models, with a 95.8% reproduction rate against official results β€” so numbers across papers are actually comparable.

  • 6 benchmarks β€” ScreenSpot-Pro, ScreenSpot-V2, UIVision, MMBench-GUI, OSWorld-G, AndroidControl
  • 11+ models β€” Qwen3-VL, Qwen2.5-VL, UI-TARS, MAI-UI, GUI-G2, UI-Venus, Gemini, Seed 1.8, and more
  • Dual backend β€” Local GPU (transformers) or remote API (OpenAI-compatible)
  • Multi-GPU & multi-thread β€” Parallel inference with automatic resume
  • ClawGUI-Agent integration β€” Pair with ClawGUI-Agent to run the full pipeline via natural language
ClawGUI-Eval Architecture

β†’ Get started with ClawGUI-Eval

πŸ€– ClawGUI-Agent β€” Deploy

πŸ“ clawgui-agent/ Β· πŸ“– Full Documentation Β· δΈ­ζ–‡

ClawGUI-Agent closes the loop from training to production. Built on OpenClaw and powered by nanobot, it lets you control Android, HarmonyOS, or iOS devices with natural language from 12+ chat platforms β€” and trigger the full ClawGUI-Eval benchmark pipeline with a single sentence, no scripts required.

  • Cross-platform β€” Android (ADB), HarmonyOS (HDC), iOS (XCTest)
  • Multi-model β€” AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS via OpenAI-compatible API
  • One-command evaluation β€” Say "benchmark qwen3vl on screenspot-pro" and it handles env check β†’ multi-GPU inference β†’ judging β†’ metrics β†’ result comparison
  • Personalized memory β€” Automatically learns user preferences and injects context across tasks
  • Episode recording β€” Every task saved as structured episodes for replay and dataset building
  • Web UI β€” Gradio interface for device management, task execution, and memory inspection
ClawGUI-Agent

β†’ Get started with ClawGUI-Agent

🧩 ClawGUI-Skills β€” Self-Evolving Skills

πŸ“ clawgui-skills/ Β· πŸ“– Full Documentation Β· δΈ­ζ–‡

ClawGUI-Skills implements the training-free self-evolving GUI skill architecture proposed and validated in our paper β€œReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents.” It stores procedural task knowledge as structured skill packages and lets PhoneAgent retrieve, inject, diagnose, and revise them on demand.

  • Four modes β€” off, trace, reuse, and evolve; disabled by default to avoid extra context cost
  • Structured packages β€” meta_info.json, plan.md, backup.md, recover.md, and failure_examples/
  • Instant revision β€” failed runs are diagnosed by an isolated verifier and mapped to targeted skill-file edits
  • Visual inspection β€” the Web UI shows matched skill name, skill_id, injected context, revisions, and failure examples

β†’ Get started with ClawGUI-Skills

πŸ“± ClawGUI-APP β€” On-Device Deploy

πŸ“ clawgui-app/ Β· πŸ“– Setup Guide

ClawGUI-APP runs the full ClawGUI "brain + GUI agent" stack directly on one Android phone, removing the old split architecture where a desktop host orchestrates tasks and the phone only executes them. Built on Shizuku for high-privilege, non-root device control.

  • Phone-only workflow β€” No desktop coordinator required; a device with Shizuku is enough
  • Two-agent design β€” Brain LLM handles planning and tool orchestration, phone agent handles screen understanding and actions
  • Multi-model support β€” AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS and more via OpenAI-compatible API
  • Voice input (STT) β€” Tap-to-record microphone with OpenAI-compatible speech-to-text transcription (SiliconFlow, Groq Whisper, etc.)
  • Conversation + automation β€” Sessions, long-term memory, external channels (Feishu), and trace replay
  • Built for real usage β€” Floating overlay status, built-in IME, session persistence, and diagnostics

β†’ Build ClawGUI-APP

🎯 Roadmap

  • ClawGUI-Agent β€” GUI agent framework for phone control and evaluation via natural language
  • ClawGUI-RL β€” Scalable mobile online RL training infrastructure with GiGPO + PRM
  • ClawGUI-Eval β€” Standardized GUI grounding evaluation suite with 6 benchmarks and 95%+ reproduction rate
  • ClawGUI-2B β€” 2B GUI agent trained with GiGPO, achieving 17.1 MobileWorld SR (vs. 11.1 baseline)
  • On-device ClawGUI-Agent (ClawGUI-APP) β€” Deploy ClawGUI-Agent directly on real phones β€” no desktop coordinator, paving the way for fully on-device inference (brain/VLM still served via cloud API today)
  • Desktop Online RL β€” Extend ClawGUI-RL to desktop environments for online reinforcement learning
  • Web Online RL β€” Extend ClawGUI-RL to web environments for online reinforcement learning
  • More Skills for ClawGUI-Agent β€” Add more pluggable skills to expand ClawGUI-Agent's capabilities
  • Hybrid CLI & GUI Mechanism β€” Explore hybrid interaction combining command-line and GUI operations
  • Real-time RL β€” Integrate real-time reinforcement learning based on the OPD algorithm for ClawGUI-RL and ClawGUI-Agent

🀝 Contributing

We welcome contributions of all kinds β€” new model support, new RL environments, bug fixes, and documentation improvements. See CONTRIBUTING.md for how to get started, module-specific guidelines, and PR requirements.

πŸ™ Acknowledgements

ClawGUI is built upon the following excellent open-source projects. We sincerely thank their contributors:

License

This project is licensed under the Apache License 2.0.

πŸ“ Citation

If you find ClawGUI useful in your research, please consider citing our paper:

@article{tang2026clawgui,
  title={ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents},
  author={Tang, Fei and Lu, Zhiqiong and Zhang, Boxuan and Lu, Weiming and Xiao, Jun and Zhuang, Yueting and Shen, Yongliang},
  journal={arXiv preprint arXiv:2604.11784},
  year={2026}
}

Star History

Close

How Shizuku is used

Can tap, swipe, type, launch apps, take screenshots and manage input methods via `sh -c` shell service through Shizuku.

This is an AI-assisted analysis of Shizuku-related usages in the app's public source code. It is best effort, so it may not catch every single usage.

How this app uses Shizuku

ClawGUI uses Shizuku to run its on-device GUI agent, letting an AI model see the screen and operate the phone with shell privileges.

  • Tap and swipe screen: the agent taps, double taps, long presses and swipes at coordinates from its visual plan using the input tap and input swipe shell commands.
  • Type and clear text: the agent enters the text from its next step with the input text shell command, clipboard paste plus paste keyevent, or a broadcast to its own input method, and clears fields with keyevents or a clear-text broadcast.
  • Press system keys: the agent goes back, home or confirms with enter using the input keyevent shell command.
  • Launch and open apps: the agent starts apps by package name using package listing, launcher and activity start shell commands, and can open links or intents with an activity start command.
  • Capture screenshots: each automation step captures the current screen with the screencap shell command for the model to see, plus display size and orientation queries to map coordinates.
  • Manage input method: the agent checks, enables and switches the active keyboard, including its own input method, with settings and input method shell commands.

Android APIs or commands used

  • input tap
  • input swipe
  • input text
  • input keyevent
  • am broadcast
  • am start
  • pm list packages
  • monkey -p
  • cmd package resolve-activity
  • dumpsys activity activities
  • dumpsys package
  • dumpsys window displays
  • screencap -p
  • chmod 666
  • wm size
  • settings get secure default_input_method
  • ime list
  • ime enable
  • ime set
  • ime reset

Notable details

If the Shizuku shell service is not bound, the controller tries a wireless ADB session first and then a local unprivileged shell, so many touch, key and launch operations still attempt to run but privileged ones may fail.

Close

Changelog

What's new for version 0.2.0

Fixes

  • Feishu GUI runner now drives step-by-step. The chat bubble updates with the current action on every step instead of staying frozen on "ζ­£εœ¨ζ‰§θ‘Œδ»»εŠ‘β€¦". Each step has a 5-minute timeout so a wedged LLM call can't hang the loop forever.
  • Clean error surfacing for Feishu replies. SDK errors are decoded to code: message instead of the previous Error@<hash>. Pre-flight failures (missing Vision creds, no device-control auth) now reply to Feishu with a human-readable reason instead of silently bailing. Image-reply failures land in Settings β†’ ε€–ιƒ¨ι€šι“ β†’ θ°ƒθ―•ζ—₯εΏ— so users can see them without logcat.
  • @_user_N mention stripping. Was hardcoded to @_user_1 ; now regex-matches @_(user|chat)_N so @bot ε‘ζœ‹ε‹εœˆ reaches PhoneAgent as just ε‘ζœ‹ε‹εœˆ.
  • Live screenshot fallback when image-reply is enabled but trace recording is off β€” users still get a final-state image instead of silent nothing.
  • IME enabled check via the Android Framework API instead of ime list -s. Users who'd already enabled "ClawGUI Input" via system settings would see "ζœͺ启用" in our panel because the shell command needed Shizuku/wadb auth to succeed. Settings β†’ θΎ“ε…₯法 also dropped the manual switcher button β€” agent handles the switching itself.

Chore

  • Repo cleanup: retired the v1 clawgui-app/ (Java client) and promoted clawgui-app-ng/ into the natural clawgui-app/ path. Android applicationId stays com.clawgui.ng so existing v0.3.0 installs upgrade in place.
Close

Permissions

16 permissions requested

  • android.permission.INTERNET
  • android.permission.ACCESS_NETWORK_STATE
  • moe.shizuku.manager.permission.API_V23
  • android.permission.FOREGROUND_SERVICE
  • android.permission.FOREGROUND_SERVICE_SPECIAL_USE
  • android.permission.POST_NOTIFICATIONS
  • android.permission.WAKE_LOCK
  • android.permission.REQUEST_IGNORE_BATTERY_OPTIMIZATIONS
  • android.permission.SYSTEM_ALERT_WINDOW
  • android.permission.VIBRATE
  • android.permission.QUERY_ALL_PACKAGES
  • android.permission.CAMERA
  • com.clawgui.ng.DYNAMIC_RECEIVER_NOT_EXPORTED_PERMISSION
  • android.permission.WRITE_EXTERNAL_STORAGE
  • android.permission.READ_PHONE_STATE
  • android.permission.READ_EXTERNAL_STORAGE
Close