Labeltap

Your agent taps by name.

Labeltap turns Android's accessibility tree into named, callable elements. Label a screen once, and your agent calls tap "Wi-Fi switch" instead of guessing coordinates from a screenshot.

Works with Claude Code and other agents. CLI and HTTP API.

Illustration: an Android settings screen drawn like an engineering part, with four numbered callouts naming the back button, the search field, the Wi-Fi row and the Wi-Fi switch. 9:41 Settings Search settings Network & internet Wi-Fi Connected Bluetooth Display Battery Notifications 1 2 3 4
View A — Labeled screenIllustrative
View B — Detail

Pixels drift. Names hold.

Screenshot agents guess where a button is and tap a coordinate. A banner, a rotation or a layout update later, the same coordinate hits something else. Labeltap resolves the element from the accessibility tree every time.

Guessing from a screenshot

tap(540, 1210)   # where the switch was
  • A vision call for every step, even ones the agent has done before.
  • Breaks when the layout shifts by a few pixels.
  • No way to tell a miss from a hit without another screenshot.

Calling a labeled element

labeltap tap "Wi-Fi switch"
  • resource id
  • description
  • text
  • class + path
  • Names are set once per app and screen, by you or suggested.
  • Screen fingerprints confirm the agent is where it thinks it is.
  • List rows become templates: "Chat row" {contact}.
View C — Sequence

Label once. Call by name. Replay without a model.

The model decides what to do. Labeltap makes sure the doing is deterministic, fast and cheap, and takes the model out of steps it already knows.

C1 — Label

Name the parts

Open a screen in the web UI, click an element, give it a name. Lists become templates, canvases become named regions.

"Wi-Fi switch" → rid switch_widget
C2 — Call

Your agent calls them

Read the screen as compact text, tap or type by name, wait for a condition instead of sleeping. From a CLI or over HTTP.

labeltap read · tap · wait
C3 — Replay

Known flows run alone

Record a flow once, tune it, and replay it with checkpoints and no model call. Plans add loops and conditions without code.

labeltap run "Open Wi-Fi settings"

Command syntax shown is illustrative.

View D — Drawn to scale

Measured on a real phone.

Measured on a Pixel 6a over adb. The bars are proportional, which is the point.

Measured times

0 to 4 seconds, true scale. Times vary with model, network and device.

Screenshot, guess, tap
~4 s
Replayed Labeltap step
~180 ms
Snapshot (image + tree)
~215 ms
Single tap
~90 ms

What the agent reads per screen

Same screen, raw tree dump versus Labeltap's compact format.

Raw screen dump
4.7 KB
Labeltap screen text
1.6 KB · −66%

Tuning a recorded flow

A short settings flow, 0 to 1.2 seconds, true scale.

Full verification
1,140 ms
Tuned replay
745 ms
View E — Loops next to the device

Spend tokens on results, not on screens.

Scrolling a list, reading each row, skipping duplicates and noticing the end is a loop, not a decision. Labeltap runs it beside the phone and hands your agent one result.

  • Screenshot agenta screen per scroll
  • Labeltapone structured result
# illustrative
$ labeltap harvest "Inbox row" \
    --fields sender,subject --until end
6 screens · 41 rows · 3 duplicates skipped
end of list detected
[
  {"sender": "…", "subject": "…"},
  …
]

Early access is open.

Labeltap is built by a team of three in Germany. If you are building an agent that needs to use Android apps, tell us what it should do.

Request early access

hello@labeltap.dev