Your agent taps by name.
Labeltap turns Android's accessibility tree into named, callable elements. Label a screen once, and your agent calls tap "Wi-Fi switch" instead of guessing coordinates from a screenshot.
Works with Claude Code and other agents. CLI and HTTP API.
Pixels drift. Names hold.
Screenshot agents guess where a button is and tap a coordinate. A banner, a rotation or a layout update later, the same coordinate hits something else. Labeltap resolves the element from the accessibility tree every time.
Guessing from a screenshot
tap(540, 1210) # where the switch was
- A vision call for every step, even ones the agent has done before.
- Breaks when the layout shifts by a few pixels.
- No way to tell a miss from a hit without another screenshot.
Calling a labeled element
labeltap tap "Wi-Fi switch"
- resource id
- description
- text
- class + path
- Names are set once per app and screen, by you or suggested.
- Screen fingerprints confirm the agent is where it thinks it is.
- List rows become templates:
"Chat row" {contact}.
Label once. Call by name. Replay without a model.
The model decides what to do. Labeltap makes sure the doing is deterministic, fast and cheap, and takes the model out of steps it already knows.
Name the parts
Open a screen in the web UI, click an element, give it a name. Lists become templates, canvases become named regions.
"Wi-Fi switch" → rid switch_widget
Your agent calls them
Read the screen as compact text, tap or type by name, wait for a condition instead of sleeping. From a CLI or over HTTP.
labeltap read · tap · wait
Known flows run alone
Record a flow once, tune it, and replay it with checkpoints and no model call. Plans add loops and conditions without code.
labeltap run "Open Wi-Fi settings"
Command syntax shown is illustrative.
Measured on a real phone.
Measured on a Pixel 6a over adb. The bars are proportional, which is the point.
Measured times
0 to 4 seconds, true scale. Times vary with model, network and device.
What the agent reads per screen
Same screen, raw tree dump versus Labeltap's compact format.
Tuning a recorded flow
A short settings flow, 0 to 1.2 seconds, true scale.
Spend tokens on results, not on screens.
Scrolling a list, reading each row, skipping duplicates and noticing the end is a loop, not a decision. Labeltap runs it beside the phone and hands your agent one result.
- Screenshot agenta screen per scroll
- Labeltapone structured result
# illustrative $ labeltap harvest "Inbox row" \ --fields sender,subject --until end 6 screens · 41 rows · 3 duplicates skipped end of list detected [ {"sender": "…", "subject": "…"}, … ]
Early access is open.
Labeltap is built by a team of three in Germany. If you are building an agent that needs to use Android apps, tell us what it should do.