AI 手机替身 纯手机端侧运行
100% 模型驻留手机端侧运行

大模型直接跑在手机里!
电脑仅用于首发安装

电脑仅在第一次帮手机传输模型文件和运行环境,安装完成后彻底拔掉数据线!大模型和 Agent 完全在手机本体上独立离线跑,零 API 费用、零外网依赖。

真机端侧运行
127.0.0.1 闭环
电脑仅用一次
装好即拔线
完全纯离线
断网也能自动操作

真我 GT 8 Pro 为什么能跑端侧模型?

手机拥有 16GB 豪华运行内存 + 骁龙 8 至尊版。 量化后的 4-bit 视觉大模型(如 Qwen2.5-VL-3B-GGUF)仅占约 2.5GB 内存,7B 模型仅占约 4.8GB 内存,手机运存极度充裕! 电脑端 AI 仅需通过 ADB 执行一次文件推送与启动,手机本地即建立 http://127.0.0.1:8080/v1 独立推理服务。

手机端侧模型推荐 (GGUF 权重)

驻留手机内存
端侧首选 · 极速低功耗
Qwen2.5-VL-3B (GGUF) 约 2.5GB 内存

专为移动设备调优的视觉模型 · 响应速度最快 · 发热极低

端侧优势: 针对骁龙 8 至尊版硬件优化,单次截屏推理仅需几秒,对手机电量消耗小,不易触发系统后台压制。
识别精度: 能够准确识别常见 App 的搜索框、安装按钮、翻页与列表项。
端侧推荐指数 ★★★★★ 综合体验最佳
Qwen2.5-VL-7B (Q4_K_M) 约 4.8GB 内存

极致 UI 精度 · 细小文字与密集图标定位之王

端侧优势: 16GB 运存可完整加载,OCR 辨识和细微按键点击命中率极高;计算稍重,适合复杂任务。
端侧推荐指数 ★★★★☆ 高精度进阶选择

给电脑 AI 的首发安装提示词

100% On-Device
on_device_setup.txt

把这段英文发给电脑上的 AI(如 Cursor / Claude Code),它将通过 ADB 一次性把模型和环境部署到手机本体中:

You are an expert Android systems engineer specializing in on-device AI deployment, ADB automation, and mobile LLM/VLM runtimes.
Your mission is to perform a ONE-TIME INITIAL SETUP to deploy a 100% ON-DEVICE local Vision-Language Model (VLM) and an autonomous GUI agent ([AutoGLM-For-Android] + [Shizuku]) onto the user's Android phone.

★ CORE ARCHITECTURAL REQUIREMENT:
The AI model and the agent must run ENTIRELY ON THE PHONE ITSELF (On-Device Inference on localhost: 127.0.0.1:8080).
The host PC is used ONLY ONCE during this initial setup to assist with downloading binaries, pushing model weights, installing APKs, and bootstrapping the on-device daemon via ADB. Once setup is complete, the user will disconnect the cable and run everything standalone on the phone with ZERO PC dependency and ZERO Cloud API costs.

### [Target Device & Hardware Profile]
1. Device: Realme GT 8 Pro (Qualcomm Snapdragon 8 Elite, 16GB LPDDR5X RAM + 512GB UFS 4.0 Storage, Realme UI / ColorOS based on Android 14/15).
2. Hardware Capacity: With 16GB RAM and top-tier Snapdragon NPU/GPU, this device can comfortably run a 4-bit quantized vision-language model (e.g., Qwen2.5-VL-3B-Instruct ~2.5GB VRAM or Qwen2.5-VL-7B-Instruct-Q4 ~4.8GB VRAM) locally inside device memory.
3. User Profile: Non-technical user. Minimize manual steps. Execute ADB commands autonomously wherever possible.

### [On-Device Stack to Deploy onto the Phone]
1. Permission Bridge: Shizuku (RikkaApps/Shizuku)
2. Agent GUI & Action Controller: AutoGLM-For-Android (Luokavin/AutoGLM-For-Android)
3. On-Device Model Runtime:
   - Precompiled Android ARM64 llama-server (or Termux with llama.cpp/Ollama-for-Android)
   - Running as a background local daemon listening on: `http://127.0.0.1:8080/v1`
4. Quantized Vision Model Weight:
   - Primary: `Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf` + `mmproj.gguf` (Lightweight, fastest inference on mobile battery)
   - High-Precision: `Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf` (Highest OCR and touch coordinate accuracy)

### [Critical Realme UI / ColorOS Pitfalls]
Realme UI includes system-level security restrictions. You must guide or configure:
1. Enable Developer Options: Tap 'Build number' 7 times in Settings -> About Phone -> Version.
2. Developer Settings:
   - Toggle ON 'USB Debugging'.
   - ★ [CRITICAL]: Toggle ON 'Disable Permission Monitoring' (禁止权限监控) — otherwise simulated touches are silently blocked by ColorOS.
   - Check 'Always allow from this computer'.

### [Step-by-Step SOP Deployment Workflow (One-Time Setup)]

#### Phase 1: Device Verification & Storage Preparation
1. Run `adb devices` to confirm connection.
2. Verify available phone storage: `adb shell df -h /sdcard` (ensure at least 10GB free space for model weights).
3. Create the model directory on the phone: `adb shell mkdir -p /sdcard/Download/ai_models/`.

#### Phase 2: Acquire and Push Model & Runtimes to the Phone
1. Assist in downloading or locating the necessary files on the PC:
   - `Shizuku.apk`
   - `AutoGLM-For-Android.apk` (and virtual keyboard IME APK)
   - ARM64 `llama-server` binary and the quantized GGUF vision model (`qwen2.5-vl-3b-instruct-q4_k_m.gguf` + `mmproj`).
2. Push APKs and model weights to the phone via ADB:
   - `adb install -r `
   - `adb install -r `
   - `adb push  /sdcard/Download/ai_models/`

#### Phase 3: One-Click Privilege Activation & Permissions
1. Activate Shizuku via ADB:
   `adb shell sh /sdcard/Android/data/moe.shizuku.privileged.api/start.sh`
   Confirm output displays "shizuku_starter exit with 0".
2. Pre-grant overlay and background execution permissions:
   `adb shell appops set com.autoglm.phone SYSTEM_ALERT_WINDOW allow`
3. Instruct user to open AutoGLM-For-Android and accept Shizuku authorization.

#### Phase 4: Bootstrap the On-Device Local Model Server
1. Launch the local inference server on the phone's localhost:
   Run the on-device server listening on `http://127.0.0.1:8080/v1` (via background process / Termux / Shizuku runner).
2. Test on-device HTTP endpoint from phone:
   `adb shell curl -s http://127.0.0.1:8080/v1/models`

#### Phase 5: Configure AutoGLM-For-Android to Phone Localhost
1. In the AutoGLM-For-Android Settings:
   - Base URL: `http://127.0.0.1:8080/v1` (100% on-device loopback)
   - Model: `qwen2.5-vl` (or matching GGUF model tag)
   - API Key: `local` (placeholder)
2. Enable the floating action ball on the phone desktop.
3. Dispatch a test command: "打开应用商店,搜索网易云音乐并下载安装".
4. Verify that the on-device model performs perception and triggers clicks directly without internet access.

#### Phase 6: Handoff & Disconnection
1. Inform the user: "One-time setup complete! You can now safely unplug the USB cable from the PC. Your phone is running the AI model and agent 100% on-device."

Please confirm understanding and begin execution when the user says "Start phone deployment" or connects the device!

部署完成后:手机端独立日常使用

完全离线
1

拔掉电脑数据线

安装完成后,电脑使命结束。你的真我手机已经是一个完整的独立 AI 手机。

2

桌面悬浮球随叫随到

轻触屏幕悬浮球,输入或语音说出指令,手机本地大模型即刻分析屏幕并自动点击执行。

3

零花费 · 零流量消耗

模型推理全在真我手机的芯片上运行,即使在飞行模式或没有网络的环境下也能正常工作。

常用指示示范(点击复制)

点击后直接粘贴给 AI
“打开应用商店,搜索网易云音乐并下载安装”
端侧视觉识别 → 自动检索 → 本地触发点击安装
“打开美团,搜索附近的瑞幸咖啡”
本地多模态解析 → 模拟滑动 → 呈现周边结果
复制成功!