S
ShizuPortal
ShizuStore

Unofficial ShizuStore Web Portal

This is an unofficial web portal for ShizuStore self-hosted by rdevz-ph. Visit the official portal at https://shizustore.com/.

ClawGUI

ClawGUI

Featuredv0.2.0
by ZJU-REALβ€’AI agentsβ€’Apache-2.0
Need ShizuStore app? Download APK

About Application

ClawGUI Logo

ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

Python 3.12 License Stars arXiv Daily Paper

HuggingFace Model ModelScope Model Project Page

English | δΈ­ζ–‡

A full-stack framework for GUI agents, covering online RL training, standardized evaluation, and deployment.

ClawGUI-Agent controls a real phone
via natural language

ClawGUI-RL trains a GUI agent with online
reinforcement learning

News

Table of Contents

πŸ’‘ Overview

ClawGUI is a research framework for GUI agents, covering the complete lifecycle from online RL training and standardized evaluation to real-device deployment.

Building a capable GUI agent involves three tightly coupled problems that are rarely solved together: you need an environment to train the agent online, rigorous benchmarks to measure what it has learned, and a production system to deploy it on real devices. ClawGUI addresses all three.

ModuleRole
πŸš€ ClawGUI-RLBuild β€” Train GUI agents online with scalable RL: parallel Docker environments, real Android devices, and GiGPO+PRM for fine-grained step-level rewards
πŸ“Š ClawGUI-EvalEvaluate β€” Measure what the agent has learned: 6 benchmarks, 11+ models, 95.8% faithful reproduction of official results
πŸ€– ClawGUI-AgentDeploy β€” Use GUI agents in the real world: control mobile devices via natural language through 12+ chat platforms, with one-command evaluation built in
🧩 ClawGUI-SkillsSelf-evolving skills β€” Training-free skill evolution proposed and validated in our paper: structured packages, retrieval, failure diagnosis, restricted revision, and reuse
πŸ“± ClawGUI-APPOn-Device Deploy β€” Run the full brain + GUI agent stack directly on one Android phone, no desktop coordinator needed, powered by Shizuku
πŸ† ClawGUI-2BEnd-to-end validation: trained entirely with ClawGUI-RL and GiGPO, achieving 17.1 MobileWorld SR vs. the 11.1 baseline

πŸ—οΈ Architecture

ClawGUI System Architecture

πŸš€ Quick Start

git clone https://github.com/ZJU-REAL/ClawGUI.git
cd ClawGUI

Each module is independent with its own environment. Click into each one for full installation and usage instructions.

πŸš€ ClawGUI-RL β€” Build

πŸ“ clawgui-rl/ Β· πŸ“– Full Documentation

ClawGUI-RL trains GUI agents with online reinforcement learning. It runs dozens of Docker-based Android emulators in parallel or trains directly on physical devices β€” and replaces standard GRPO with GiGPO+PRM for fine-grained step-level rewards that drive stronger policy learning.

  • Parallel multi-environment β€” Dozens of Docker-based virtual Android environments simultaneously
  • Real-device training β€” Physical or cloud Android phones with the same API
  • GiGPO + PRM β€” Fine-grained step-level reward for better policy optimization than standard GRPO
  • Spare server rotation β€” Automatic failover keeps training running without interruption
  • Episode visualization β€” Record and replay any training trajectory
ClawGUI-RL Architecture

β†’ Get started with ClawGUI-RL

πŸ“Š ClawGUI-Eval β€” Evaluate

πŸ“ clawgui-eval/ Β· πŸ“– Full Documentation Β· πŸ€— Dataset Β· πŸ€– ModelScope

ClawGUI-Eval gives GUI grounding research a reliable measurement baseline. Its three-stage Infer β†’ Judge β†’ Metric pipeline covers 6 benchmarks and 11+ models, with a 95.8% reproduction rate against official results β€” so numbers across papers are actually comparable.

  • 6 benchmarks β€” ScreenSpot-Pro, ScreenSpot-V2, UIVision, MMBench-GUI, OSWorld-G, AndroidControl
  • 11+ models β€” Qwen3-VL, Qwen2.5-VL, UI-TARS, MAI-UI, GUI-G2, UI-Venus, Gemini, Seed 1.8, and more
  • Dual backend β€” Local GPU (transformers) or remote API (OpenAI-compatible)
  • Multi-GPU & multi-thread β€” Parallel inference with automatic resume
  • ClawGUI-Agent integration β€” Pair with ClawGUI-Agent to run the full pipeline via natural language
ClawGUI-Eval Architecture

β†’ Get started with ClawGUI-Eval

πŸ€– ClawGUI-Agent β€” Deploy

πŸ“ clawgui-agent/ Β· πŸ“– Full Documentation Β· δΈ­ζ–‡

ClawGUI-Agent closes the loop from training to production. Built on OpenClaw and powered by nanobot, it lets you control Android, HarmonyOS, or iOS devices with natural language from 12+ chat platforms β€” and trigger the full ClawGUI-Eval benchmark pipeline with a single sentence, no scripts required.

  • Cross-platform β€” Android (ADB), HarmonyOS (HDC), iOS (XCTest)
  • Multi-model β€” AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS via OpenAI-compatible API
  • One-command evaluation β€” Say "benchmark qwen3vl on screenspot-pro" and it handles env check β†’ multi-GPU inference β†’ judging β†’ metrics β†’ result comparison
  • Personalized memory β€” Automatically learns user preferences and injects context across tasks
  • Episode recording β€” Every task saved as structured episodes for replay and dataset building
  • Web UI β€” Gradio interface for device management, task execution, and memory inspection
ClawGUI-Agent

β†’ Get started with ClawGUI-Agent

🧩 ClawGUI-Skills β€” Self-Evolving Skills

πŸ“ clawgui-skills/ Β· πŸ“– Full Documentation Β· δΈ­ζ–‡

ClawGUI-Skills implements the training-free self-evolving GUI skill architecture proposed and validated in our paper β€œReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents.” It stores procedural task knowledge as structured skill packages and lets PhoneAgent retrieve, inject, diagnose, and revise them on demand.

  • Four modes β€” off, trace, reuse, and evolve; disabled by default to avoid extra context cost
  • Structured packages β€” meta_info.json, plan.md, backup.md, recover.md, and failure_examples/
  • Instant revision β€” failed runs are diagnosed by an isolated verifier and mapped to targeted skill-file edits
  • Visual inspection β€” the Web UI shows matched skill name, skill_id, injected context, revisions, and failure examples

β†’ Get started with ClawGUI-Skills

πŸ“± ClawGUI-APP β€” On-Device Deploy

πŸ“ clawgui-app/ Β· πŸ“– Setup Guide

ClawGUI-APP runs the full ClawGUI "brain + GUI agent" stack directly on one Android phone, removing the old split architecture where a desktop host orchestrates tasks and the phone only executes them. Built on Shizuku for high-privilege, non-root device control.

  • Phone-only workflow β€” No desktop coordinator required; a device with Shizuku is enough
  • Two-agent design β€” Brain LLM handles planning and tool orchestration, phone agent handles screen understanding and actions
  • Multi-model support β€” AutoGLM, MAI-UI, GUI-Owl, Qwen-VL, UI-TARS and more via OpenAI-compatible API
  • Voice input (STT) β€” Tap-to-record microphone with OpenAI-compatible speech-to-text transcription (SiliconFlow, Groq Whisper, etc.)
  • Conversation + automation β€” Sessions, long-term memory, external channels (Feishu), and trace replay
  • Built for real usage β€” Floating overlay status, built-in IME, session persistence, and diagnostics

β†’ Build ClawGUI-APP

🎯 Roadmap

  • ClawGUI-Agent β€” GUI agent framework for phone control and evaluation via natural language
  • ClawGUI-RL β€” Scalable mobile online RL training infrastructure with GiGPO + PRM
  • ClawGUI-Eval β€” Standardized GUI grounding evaluation suite with 6 benchmarks and 95%+ reproduction rate
  • ClawGUI-2B β€” 2B GUI agent trained with GiGPO, achieving 17.1 MobileWorld SR (vs. 11.1 baseline)
  • On-device ClawGUI-Agent (ClawGUI-APP) β€” Deploy ClawGUI-Agent directly on real phones β€” no desktop coordinator, paving the way for fully on-device inference (brain/VLM still served via cloud API today)
  • Desktop Online RL β€” Extend ClawGUI-RL to desktop environments for online reinforcement learning
  • Web Online RL β€” Extend ClawGUI-RL to web environments for online reinforcement learning
  • More Skills for ClawGUI-Agent β€” Add more pluggable skills to expand ClawGUI-Agent's capabilities
  • Hybrid CLI & GUI Mechanism β€” Explore hybrid interaction combining command-line and GUI operations
  • Real-time RL β€” Integrate real-time reinforcement learning based on the OPD algorithm for ClawGUI-RL and ClawGUI-Agent

🀝 Contributing

We welcome contributions of all kinds β€” new model support, new RL environments, bug fixes, and documentation improvements. See CONTRIBUTING.md for how to get started, module-specific guidelines, and PR requirements.

App image

πŸ™ Acknowledgements

ClawGUI is built upon the following excellent open-source projects. We sincerely thank their contributors:

License

This project is licensed under the Apache License 2.0.

πŸ“ Citation

If you find ClawGUI useful in your research, please consider citing our paper:

@article{tang2026clawgui,
  title={ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents},
  author={Tang, Fei and Lu, Zhiqiong and Zhang, Boxuan and Lu, Weiming and Xiao, Jun and Zhuang, Yueting and Shen, Yongliang},
  journal={arXiv preprint arXiv:2604.11784},
  year={2026}
}

Star History

Star History Chart

Release Notes & Changelog

Fixes

  • Feishu GUI runner now drives step-by-step. The chat bubble updates with the current action on every step instead of staying frozen on "ζ­£εœ¨ζ‰§θ‘Œδ»»εŠ‘β€¦". Each step has a 5-minute timeout so a wedged LLM call can't hang the loop forever.
  • Clean error surfacing for Feishu replies. SDK errors are decoded to code: message instead of the previous Error@<hash>. Pre-flight failures (missing Vision creds, no device-control auth) now reply to Feishu with a human-readable reason instead of silently bailing. Image-reply failures land in Settings β†’ ε€–ιƒ¨ι€šι“ β†’ θ°ƒθ―•ζ—₯εΏ— so users can see them without logcat.
  • @_user_N mention stripping. Was hardcoded to @_user_1 ; now regex-matches @_(user|chat)_N so @bot ε‘ζœ‹ε‹εœˆ reaches PhoneAgent as just ε‘ζœ‹ε‹εœˆ.
  • Live screenshot fallback when image-reply is enabled but trace recording is off β€” users still get a final-state image instead of silent nothing.
  • IME enabled check via the Android Framework API instead of ime list -s. Users who'd already enabled "ClawGUI Input" via system settings would see "ζœͺ启用" in our panel because the shell command needed Shizuku/wadb auth to succeed. Settings β†’ θΎ“ε…₯法 also dropped the manual switcher button β€” agent handles the switching itself.

Chore

  • Repo cleanup: retired the v1 clawgui-app/ (Java client) and promoted clawgui-app-ng/ into the natural clawgui-app/ path. Android applicationId stays com.clawgui.ng so existing v0.3.0 installs upgrade in place.

Declared Android Permissions

16 total

Shizuku API Permissions

moe.shizuku.manager.permission.API_V23

This application connects to the Shizuku service to execute system-level operations with ADB elevated privileges.

Standard Android Permissions

INTERNET
ACCESS_NETWORK_STATE
FOREGROUND_SERVICE
FOREGROUND_SERVICE_SPECIAL_USE
POST_NOTIFICATIONS
WAKE_LOCK
REQUEST_IGNORE_BATTERY_OPTIMIZATIONS
SYSTEM_ALERT_WINDOW
VIBRATE
QUERY_ALL_PACKAGES
CAMERA
com.clawgui.ng.DYNAMIC_RECEIVER_NOT_EXPORTED_PERMISSION
WRITE_EXTERNAL_STORAGE
READ_PHONE_STATE
READ_EXTERNAL_STORAGE

Specifications

CategoryAI agents
Packagecom.clawgui.ng
LicenseApache-2.0
GitHub Stars1,356
Installs105
Target SDKAndroid 14+ (API 36)
Last UpdatedOctober 2, 2026
Release DateOctober 2, 2026
Root RequirementRootless (Shizuku)
ShizuStore Installation

Install via ShizuStore to enable automatic updates and silent installations using Shizuku.

ClawGUI - Shizuku App | ShizuPortal