finds.dev← search

// the find

minitap-ai/mobile-use

★ 3,217 · Python · Apache-2.0 · updated Sep 2026

AI agents can now use real Android and iOS apps, just like a human.

A Python agent that drives a real Android phone or iOS simulator from a natural-language command, reading the accessibility tree and tapping, typing, and swiping until the task finishes. It is aimed at developers who need to automate an app they cannot reach through an API, such as pulling structured data out of a mobile-only service or running exploratory QA flows on a device.

- The agent is split into planner, orchestrator, contextor, cortex, executor, and outputter nodes on LangGraph instead of one monolithic loop. That makes it possible to see which step failed and to set a different model per node in llm-config.override.jsonc.

- A unified_controller.py sits over separate Android (UI Automator over ADB) and iOS (idb and a WebDriverAgent client) controllers, so the planning layer does not need to know which platform it is driving.

- The --output-description flag turns a task into structured output such as a JSON list, which makes it usable as a data extraction step and not only as a demo of phone control.

- Provider support is broad, covering OpenAI-compatible endpoints, Anthropic, Google, MiniMax, and OpenRouter, and the README links an arXiv paper that describes the task-decomposition approach.

- Physical iOS devices are not supported, and the Docker quickstart only handles Android. iOS users are limited to simulators on macOS with Xcode and idb-companion installed.

- The README admits the accessibility-tree approach works poorly on games, which do not expose that data. Anything drawn on a canvas is effectively out of scope.

- Each step goes through LLM calls across several agent nodes, so cost and latency per task are not small. The README gives no per-task numbers, so you will need to tune llm-config before you know what a run costs.

- The 100% AndroidWorld claim comes from the vendor's own benchmark and points to a linked leaderboard. Treat it as something to verify, not a settled result. The shown test tree is small and covers mainly clients and one provider, so most of the graph is only exercised by running it against a device.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →