Files
autofirmer-expanded/clicker/README.md
T
Brandon LiandClaude Opus 5 c5c946d3cc Add Linux/Raspberry Pi setup and port the clicker to X11
The clicker had no Linux support at all: pygetwindow has no X11 backend, so
_other_activate returned "window management is unsupported" and every step
failed before it could click. focus.py now has a third backend that raises the
browser with xdotool and reads the focused window's WM_CLASS to verify it came
forward - the same activate-then-confirm shape as the macOS and Windows paths.

setup-linux.sh mirrors setup-windows.bat with apt instead of winget. Three
things are specific to this platform rather than incidental:

- Node comes from NodeSource. Pi OS ships one too old for Next 16, and the LTS
  line is also what better-sqlite3 publishes prebuilt arm64 binaries for.
- Python dependencies go in a virtualenv. Pi OS Bookworm enforces PEP 668, so
  pip into the system interpreter fails with externally-managed-environment.
- Autostart uses an XDG ~/.config/autostart entry, the direct analogue of the
  Windows Startup folder: no sudo, and it runs inside the graphical session,
  which the clicker needs for DISPLAY.

Wayland is called out in four places - the session guard, diagnose.py, the
installer and both READMEs - because Pi OS on a Pi 5 defaults to it and the
clicker simply cannot work there. Wayland does not let one client synthesise
input into another, so this is a switch-to-X11 situation, not a bug to fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-30 18:53:25 -05:00

270 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Clicker
Moves the real desktop mouse and clicks an element you name with a CSS selector.
The extension can see the DOM but can't move the mouse; this script can move the
mouse but can't see inside Chrome. They meet at the dashboard API:
```
clicker --POST /api/autobuyer/locate------> "where is button.buy?"
extension --POST /api/autobuyer/locate/claim takes the request
extension raises the Chrome window, brings
the tab to the front, scrolls the
element to centre, measures it
extension --POST /api/autobuyer/locate/result desktop x,y
clicker --GET /api/autobuyer/locate?id=---> reads the answer
clicker moves the mouse, clicks
```
## Two ways to run it
**`clicker.py`** — one action at a time, from the shell. Use this to find
selectors and confirm coordinates before wiring anything up.
**`runner.py`** — the daemon behind the dashboard's buttons. Leave it running; it
polls for queued runs and executes the steps.
```bash
python runner.py # then press a button on the AutoBuyer page
python runner.py --dry-run # walks the steps, never presses or types
```
Automations are defined in `lib/automations.ts`, and the steps are sent to the
runner by the server — so adding a button means editing that file, and nothing in
the runner or the page changes. Each run reports step-by-step progress back to the
dashboard, and the page's Stop button halts a run between steps.
Only one run executes at a time: two processes driving the same physical mouse
would interleave clicks.
## Setup
```bash
pip install -r requirements.txt
```
The macOS-only backends are gated behind platform markers, so this installs the
right set on either OS.
### macOS
Grant Accessibility permission to whatever runs the script — Terminal, iTerm,
VS Code — under System Settings → Privacy & Security → Accessibility. Without it
`pyautogui` moves nothing and fails silently, which looks identical to a bad
selector.
### Linux (including Raspberry Pi)
`setup-linux.sh` installs `xdotool` and `scrot` and builds a virtualenv for the
Python side.
**X11 only.** pygetwindow has no X11 backend, so raising the browser goes
through `xdotool` instead. Neither that nor pyautogui works under Wayland — it
refuses by design to let one client drive another — and Raspberry Pi OS on a Pi
5 defaults to Wayland. Switch with `sudo raspi-config` -> Advanced Options ->
Wayland -> X11 and reboot. `diagnose.py` prints which session you are in.
### Windows
No extra permissions, but two things differ.
**Display scaling.** The runner declares itself DPI-aware at startup, and prints
which mode it got (`per-monitor DPI aware`). That stops Windows reporting a
virtualised screen size and rescaling the coordinates it accepts, which would
otherwise put clicks progressively further off as you move from the top-left.
**Foreground lock.** Windows refuses to raise a window for a process that doesn't
already own the foreground — it flashes the taskbar instead, and the activation
call returns as if it worked. The runner verifies by reading the foreground
window's process, so a failure is reported rather than clicked into. If runs stop
with *"the browser is not frontmost"*, click the Chrome window once by hand and
try again; something else is holding the foreground.
Browser windows are matched on the owning process (`chrome.exe`, `msedge.exe`,
`brave.exe`), not the window title, so an editor with `chrome.js` open won't be
mistaken for the browser.
### First run on a new machine
Verify in this order — each step is harmless on its own, and the first one that
looks wrong tells you where the problem is:
```bash
# 1. Can the extension see the page, and do the coordinates look sane?
python clicker.py locate "a.add_account_btn"
# 2. Does the cursor actually land on the element? Nothing is pressed.
python clicker.py click "a.add_account_btn" --dry-run
# 3. A real click.
python clicker.py click "a.add_account_btn"
```
At step 1, check the reported desktop x,y against where the element actually is
on screen. A consistent offset means display scaling was misdetected — pass
`--scale` explicitly. Only then start the daemon.
### Linux / X11
Not supported. `pygetwindow`, which raises the browser window, has no X11 backend,
so every click is refused with *"window management is unsupported on linux"*.
Making it work needs an `xdotool` or `wmctrl` path in `focus.py`, a real window
manager (Xvfb alone has no focus semantics), and X11 rather than Wayland.
## Use
The AutoBuyer switch on the dashboard is the master arm: the script refuses to run
while it's off.
```bash
# Measure only — no mouse movement. Start here.
python clicker.py locate "button.buy"
# Move the cursor to the target but don't press.
python clicker.py click "button.buy" --dry-run
# Actually click.
python clicker.py click "button.buy"
# Type into a field: clicks it to place the caret, then types.
python clicker.py type 'input[name="quantity"]' --text "5000" --clear
# Keep the text out of shell history.
echo "5000" | python clicker.py type 'input#qty' --stdin --clear
# Pin it to a specific tab, and pick the 3rd match.
python clicker.py click ".trade-btn" --index 2 --url "https://tradeify.co/*"
```
| Flag | Meaning |
|---|---|
| `--url` | Chrome match pattern for the tab. Without it, the hosts in the extension manifest. |
| `--open-url` | Page to open if no tab matches `--url`. |
| `--index` | Which match, when the selector hits several (default 0). |
| `--api` | Dashboard URL (default `http://localhost:3000`). |
| `--timeout` | Seconds to wait for the extension (default 45 — a cold page load takes time). |
| `--scale` | CSS-to-desktop pixel ratio. Auto-detected; override if clicks land off. |
| `--dry-run` | Move the cursor, don't press. |
| `--force` | Click even when something is covering the element. |
| `--robotic` | Straight-line move and instant click, skipping the motion model. |
| `--seed` | Seed the motion RNG so a run replays identically (debugging). |
| `--no-activate` | Don't raise the browser first. The click may then be swallowed. |
| `--text` | Text to type (action `type`). |
| `--stdin` | Read the text from stdin instead, keeping it out of shell history. |
| `--clear` | Select-all and delete before typing, instead of appending at the caret. |
| `--allow-enter` | Permit newlines. Each one presses Enter, which may submit the form. |
| `--no-verify` | Skip reading the field back after typing. |
## Typing
`type` clicks the field to place the caret, then types with an uneven human
cadence (45130ms between keys, a beat after each space, an occasional longer
pause). Real OS-level keystrokes are also the *correct* way to fill a React form:
setting `.value` directly is ignored by controlled components, while genuine key
events are not.
Four guards, all of which catch silent failures:
- **Non-ASCII is rejected.** `pyautogui.write()` has no keycode for `é` or `£` and
skips them without complaint, which would leave a quietly truncated value in the
field. Better to refuse than to submit `caf` where you meant `café`.
- **Newlines are rejected** unless `--allow-enter`, because Enter may submit the
form — and on this site that could mean placing an order.
- **Non-editable targets are refused** — disabled, read-only, or simply not an
input. The keystrokes would go nowhere.
- **The field is read back afterwards** and compared against what was typed
(`--no-verify` skips it). A field that never took focus, or that ignored the
input, otherwise looks exactly like success. Exit code 4 means the text did not
land.
`--clear` selects-all and deletes first. Without it the text is inserted at
wherever the caret landed, which for a field with existing contents usually
produces something like `50005000`.
Typos are deliberately *not* simulated. A mistyped digit in a trading form is a
real loss, and the backspace-and-correct step is exactly the part that can go
wrong — a field with input masking or autocomplete can swallow the correction and
leave the wrong number behind.
Don't pass credentials via `--text`: it lands in your shell history and in the
process list. `--stdin` avoids both, but nothing here is built to handle secrets
safely.
## Window focus
A click on a window that isn't focused is consumed by the window manager
*activating* that window — it never reaches the control underneath. That's why an
automated click against a background Chrome appears to do nothing the first time
and work the second: the first click only brought Chrome forward.
The extension calls `chrome.windows.update({focused: true})`, but that only orders
windows **within** Chrome. If the frontmost *application* is your terminal — which
it is, since that's where you launched this — Chrome is still in the background.
So `focus.py` raises the browser application itself immediately before the press,
then confirms it actually came forward before committing to the click. If it can't
verify, it refuses rather than firing a click that would be eaten (exit code 3).
On macOS this uses `NSWorkspace` via pyobjc, which needs no Automation permission —
activating an app is not scripting it. Without pyobjc it falls back to `osascript`,
which does prompt for Automation permission the first time.
On Windows it goes through `pygetwindow`, then confirms by reading the foreground
window's process. That confirmation is the important half: Windows silently
declines to raise a window for a background process, and the activation call
reports success either way.
## Cursor motion
`humanize.py` moves the pointer the way a hand does rather than teleporting:
- a curved (cubic Bézier) path instead of a straight line, bowing to one side
- eased velocity — accelerate out, coast, decelerate in
- sub-pixel tremor that decays near the target, so the landing stays exact
- long throws (>260px) sometimes overshoot slightly and pull back
- a 60170ms dwell after arriving, before the press
- the button held down 55120ms rather than an instant down/up
This is about reliability as much as appearance. Plenty of web controls only arm
once they have actually been hovered — dropdowns, tooltip-gated buttons, custom
widgets — and a cursor that arrives and presses in the same tick can outrun the
page's own `mousemove` handlers. The dwell is what lets those catch up.
Note `pyautogui.PAUSE` is set to 0 on import: it otherwise sleeps 0.1s after
*every* call, which would add tens of seconds across a stepped path.
Exit codes: `0` ok, `1` error, `2` capture switch off, `3` refused to click
(covered element, or coordinates off-screen).
## Safety
- **Failsafe**: slam the cursor into a screen corner to abort mid-run.
- **Covered elements**: before reporting, the extension checks
`document.elementFromPoint()` at the target centre. If a cookie banner or modal
is on top, the click is refused rather than sent into the overlay — `--force`
overrides.
- **Off-screen check**: coordinates outside the display bounds are refused, which
catches a Chrome window on a second monitor or partly off the edge.
- The element is scrolled to the centre of the viewport before measuring, so a
target below the fold is handled rather than mis-clicked.
## Known limits
- **Latency** is bounded by the extension's poll interval (default 3s), so a click
takes a few seconds to fire. Drop the interval in the extension popup if that
matters.
- **Multi-monitor**: coordinates come from `window.screenX/screenY`, which are
relative to the primary display's origin. A Chrome window on a secondary monitor
usually still works, but verify with `locate` before trusting `click`.
- **Display scaling**: the scale factor is inferred by comparing the screen width
the OS reports against the one the browser reports, and is only trusted when it
lands on a real scaling factor (1.0, 1.25, 1.5, …). Anything else falls back to
1:1 with a warning, because on a multi-monitor desktop the two sides describe
different displays and the ratio is meaningless. If clicks land at a consistent
offset, set `--scale` explicitly.
- **Focus is exclusive**: the browser must be frontmost at the moment of each
click, so the machine can't be used for anything else during a run, and a stray
dialog stealing focus fails the step.
- The measurement and the click are separate moments. If the page moves the element
in between (a re-render, a late-loading banner), the click lands where it *was*.