The clicker had no Linux support at all: pygetwindow has no X11 backend, so _other_activate returned "window management is unsupported" and every step failed before it could click. focus.py now has a third backend that raises the browser with xdotool and reads the focused window's WM_CLASS to verify it came forward - the same activate-then-confirm shape as the macOS and Windows paths. setup-linux.sh mirrors setup-windows.bat with apt instead of winget. Three things are specific to this platform rather than incidental: - Node comes from NodeSource. Pi OS ships one too old for Next 16, and the LTS line is also what better-sqlite3 publishes prebuilt arm64 binaries for. - Python dependencies go in a virtualenv. Pi OS Bookworm enforces PEP 668, so pip into the system interpreter fails with externally-managed-environment. - Autostart uses an XDG ~/.config/autostart entry, the direct analogue of the Windows Startup folder: no sudo, and it runs inside the graphical session, which the clicker needs for DISPLAY. Wayland is called out in four places - the session guard, diagnose.py, the installer and both READMEs - because Pi OS on a Pi 5 defaults to it and the clicker simply cannot work there. Wayland does not let one client synthesise input into another, so this is a switch-to-X11 situation, not a bug to fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
270 lines
12 KiB
Markdown
270 lines
12 KiB
Markdown
# Clicker
|
||
|
||
Moves the real desktop mouse and clicks an element you name with a CSS selector.
|
||
|
||
The extension can see the DOM but can't move the mouse; this script can move the
|
||
mouse but can't see inside Chrome. They meet at the dashboard API:
|
||
|
||
```
|
||
clicker --POST /api/autobuyer/locate------> "where is button.buy?"
|
||
extension --POST /api/autobuyer/locate/claim takes the request
|
||
extension raises the Chrome window, brings
|
||
the tab to the front, scrolls the
|
||
element to centre, measures it
|
||
extension --POST /api/autobuyer/locate/result desktop x,y
|
||
clicker --GET /api/autobuyer/locate?id=---> reads the answer
|
||
clicker moves the mouse, clicks
|
||
```
|
||
|
||
## Two ways to run it
|
||
|
||
**`clicker.py`** — one action at a time, from the shell. Use this to find
|
||
selectors and confirm coordinates before wiring anything up.
|
||
|
||
**`runner.py`** — the daemon behind the dashboard's buttons. Leave it running; it
|
||
polls for queued runs and executes the steps.
|
||
|
||
```bash
|
||
python runner.py # then press a button on the AutoBuyer page
|
||
python runner.py --dry-run # walks the steps, never presses or types
|
||
```
|
||
|
||
Automations are defined in `lib/automations.ts`, and the steps are sent to the
|
||
runner by the server — so adding a button means editing that file, and nothing in
|
||
the runner or the page changes. Each run reports step-by-step progress back to the
|
||
dashboard, and the page's Stop button halts a run between steps.
|
||
|
||
Only one run executes at a time: two processes driving the same physical mouse
|
||
would interleave clicks.
|
||
|
||
## Setup
|
||
|
||
```bash
|
||
pip install -r requirements.txt
|
||
```
|
||
|
||
The macOS-only backends are gated behind platform markers, so this installs the
|
||
right set on either OS.
|
||
|
||
### macOS
|
||
|
||
Grant Accessibility permission to whatever runs the script — Terminal, iTerm,
|
||
VS Code — under System Settings → Privacy & Security → Accessibility. Without it
|
||
`pyautogui` moves nothing and fails silently, which looks identical to a bad
|
||
selector.
|
||
|
||
### Linux (including Raspberry Pi)
|
||
|
||
`setup-linux.sh` installs `xdotool` and `scrot` and builds a virtualenv for the
|
||
Python side.
|
||
|
||
**X11 only.** pygetwindow has no X11 backend, so raising the browser goes
|
||
through `xdotool` instead. Neither that nor pyautogui works under Wayland — it
|
||
refuses by design to let one client drive another — and Raspberry Pi OS on a Pi
|
||
5 defaults to Wayland. Switch with `sudo raspi-config` -> Advanced Options ->
|
||
Wayland -> X11 and reboot. `diagnose.py` prints which session you are in.
|
||
|
||
### Windows
|
||
|
||
No extra permissions, but two things differ.
|
||
|
||
**Display scaling.** The runner declares itself DPI-aware at startup, and prints
|
||
which mode it got (`per-monitor DPI aware`). That stops Windows reporting a
|
||
virtualised screen size and rescaling the coordinates it accepts, which would
|
||
otherwise put clicks progressively further off as you move from the top-left.
|
||
|
||
**Foreground lock.** Windows refuses to raise a window for a process that doesn't
|
||
already own the foreground — it flashes the taskbar instead, and the activation
|
||
call returns as if it worked. The runner verifies by reading the foreground
|
||
window's process, so a failure is reported rather than clicked into. If runs stop
|
||
with *"the browser is not frontmost"*, click the Chrome window once by hand and
|
||
try again; something else is holding the foreground.
|
||
|
||
Browser windows are matched on the owning process (`chrome.exe`, `msedge.exe`,
|
||
`brave.exe`), not the window title, so an editor with `chrome.js` open won't be
|
||
mistaken for the browser.
|
||
|
||
### First run on a new machine
|
||
|
||
Verify in this order — each step is harmless on its own, and the first one that
|
||
looks wrong tells you where the problem is:
|
||
|
||
```bash
|
||
# 1. Can the extension see the page, and do the coordinates look sane?
|
||
python clicker.py locate "a.add_account_btn"
|
||
|
||
# 2. Does the cursor actually land on the element? Nothing is pressed.
|
||
python clicker.py click "a.add_account_btn" --dry-run
|
||
|
||
# 3. A real click.
|
||
python clicker.py click "a.add_account_btn"
|
||
```
|
||
|
||
At step 1, check the reported desktop x,y against where the element actually is
|
||
on screen. A consistent offset means display scaling was misdetected — pass
|
||
`--scale` explicitly. Only then start the daemon.
|
||
|
||
### Linux / X11
|
||
|
||
Not supported. `pygetwindow`, which raises the browser window, has no X11 backend,
|
||
so every click is refused with *"window management is unsupported on linux"*.
|
||
Making it work needs an `xdotool` or `wmctrl` path in `focus.py`, a real window
|
||
manager (Xvfb alone has no focus semantics), and X11 rather than Wayland.
|
||
|
||
## Use
|
||
|
||
The AutoBuyer switch on the dashboard is the master arm: the script refuses to run
|
||
while it's off.
|
||
|
||
```bash
|
||
# Measure only — no mouse movement. Start here.
|
||
python clicker.py locate "button.buy"
|
||
|
||
# Move the cursor to the target but don't press.
|
||
python clicker.py click "button.buy" --dry-run
|
||
|
||
# Actually click.
|
||
python clicker.py click "button.buy"
|
||
|
||
# Type into a field: clicks it to place the caret, then types.
|
||
python clicker.py type 'input[name="quantity"]' --text "5000" --clear
|
||
|
||
# Keep the text out of shell history.
|
||
echo "5000" | python clicker.py type 'input#qty' --stdin --clear
|
||
|
||
# Pin it to a specific tab, and pick the 3rd match.
|
||
python clicker.py click ".trade-btn" --index 2 --url "https://tradeify.co/*"
|
||
```
|
||
|
||
| Flag | Meaning |
|
||
|---|---|
|
||
| `--url` | Chrome match pattern for the tab. Without it, the hosts in the extension manifest. |
|
||
| `--open-url` | Page to open if no tab matches `--url`. |
|
||
| `--index` | Which match, when the selector hits several (default 0). |
|
||
| `--api` | Dashboard URL (default `http://localhost:3000`). |
|
||
| `--timeout` | Seconds to wait for the extension (default 45 — a cold page load takes time). |
|
||
| `--scale` | CSS-to-desktop pixel ratio. Auto-detected; override if clicks land off. |
|
||
| `--dry-run` | Move the cursor, don't press. |
|
||
| `--force` | Click even when something is covering the element. |
|
||
| `--robotic` | Straight-line move and instant click, skipping the motion model. |
|
||
| `--seed` | Seed the motion RNG so a run replays identically (debugging). |
|
||
| `--no-activate` | Don't raise the browser first. The click may then be swallowed. |
|
||
| `--text` | Text to type (action `type`). |
|
||
| `--stdin` | Read the text from stdin instead, keeping it out of shell history. |
|
||
| `--clear` | Select-all and delete before typing, instead of appending at the caret. |
|
||
| `--allow-enter` | Permit newlines. Each one presses Enter, which may submit the form. |
|
||
| `--no-verify` | Skip reading the field back after typing. |
|
||
|
||
## Typing
|
||
|
||
`type` clicks the field to place the caret, then types with an uneven human
|
||
cadence (45–130ms between keys, a beat after each space, an occasional longer
|
||
pause). Real OS-level keystrokes are also the *correct* way to fill a React form:
|
||
setting `.value` directly is ignored by controlled components, while genuine key
|
||
events are not.
|
||
|
||
Four guards, all of which catch silent failures:
|
||
|
||
- **Non-ASCII is rejected.** `pyautogui.write()` has no keycode for `é` or `£` and
|
||
skips them without complaint, which would leave a quietly truncated value in the
|
||
field. Better to refuse than to submit `caf` where you meant `café`.
|
||
- **Newlines are rejected** unless `--allow-enter`, because Enter may submit the
|
||
form — and on this site that could mean placing an order.
|
||
- **Non-editable targets are refused** — disabled, read-only, or simply not an
|
||
input. The keystrokes would go nowhere.
|
||
- **The field is read back afterwards** and compared against what was typed
|
||
(`--no-verify` skips it). A field that never took focus, or that ignored the
|
||
input, otherwise looks exactly like success. Exit code 4 means the text did not
|
||
land.
|
||
|
||
`--clear` selects-all and deletes first. Without it the text is inserted at
|
||
wherever the caret landed, which for a field with existing contents usually
|
||
produces something like `50005000`.
|
||
|
||
Typos are deliberately *not* simulated. A mistyped digit in a trading form is a
|
||
real loss, and the backspace-and-correct step is exactly the part that can go
|
||
wrong — a field with input masking or autocomplete can swallow the correction and
|
||
leave the wrong number behind.
|
||
|
||
Don't pass credentials via `--text`: it lands in your shell history and in the
|
||
process list. `--stdin` avoids both, but nothing here is built to handle secrets
|
||
safely.
|
||
|
||
## Window focus
|
||
|
||
A click on a window that isn't focused is consumed by the window manager
|
||
*activating* that window — it never reaches the control underneath. That's why an
|
||
automated click against a background Chrome appears to do nothing the first time
|
||
and work the second: the first click only brought Chrome forward.
|
||
|
||
The extension calls `chrome.windows.update({focused: true})`, but that only orders
|
||
windows **within** Chrome. If the frontmost *application* is your terminal — which
|
||
it is, since that's where you launched this — Chrome is still in the background.
|
||
|
||
So `focus.py` raises the browser application itself immediately before the press,
|
||
then confirms it actually came forward before committing to the click. If it can't
|
||
verify, it refuses rather than firing a click that would be eaten (exit code 3).
|
||
|
||
On macOS this uses `NSWorkspace` via pyobjc, which needs no Automation permission —
|
||
activating an app is not scripting it. Without pyobjc it falls back to `osascript`,
|
||
which does prompt for Automation permission the first time.
|
||
|
||
On Windows it goes through `pygetwindow`, then confirms by reading the foreground
|
||
window's process. That confirmation is the important half: Windows silently
|
||
declines to raise a window for a background process, and the activation call
|
||
reports success either way.
|
||
|
||
## Cursor motion
|
||
|
||
`humanize.py` moves the pointer the way a hand does rather than teleporting:
|
||
|
||
- a curved (cubic Bézier) path instead of a straight line, bowing to one side
|
||
- eased velocity — accelerate out, coast, decelerate in
|
||
- sub-pixel tremor that decays near the target, so the landing stays exact
|
||
- long throws (>260px) sometimes overshoot slightly and pull back
|
||
- a 60–170ms dwell after arriving, before the press
|
||
- the button held down 55–120ms rather than an instant down/up
|
||
|
||
This is about reliability as much as appearance. Plenty of web controls only arm
|
||
once they have actually been hovered — dropdowns, tooltip-gated buttons, custom
|
||
widgets — and a cursor that arrives and presses in the same tick can outrun the
|
||
page's own `mousemove` handlers. The dwell is what lets those catch up.
|
||
|
||
Note `pyautogui.PAUSE` is set to 0 on import: it otherwise sleeps 0.1s after
|
||
*every* call, which would add tens of seconds across a stepped path.
|
||
|
||
Exit codes: `0` ok, `1` error, `2` capture switch off, `3` refused to click
|
||
(covered element, or coordinates off-screen).
|
||
|
||
## Safety
|
||
|
||
- **Failsafe**: slam the cursor into a screen corner to abort mid-run.
|
||
- **Covered elements**: before reporting, the extension checks
|
||
`document.elementFromPoint()` at the target centre. If a cookie banner or modal
|
||
is on top, the click is refused rather than sent into the overlay — `--force`
|
||
overrides.
|
||
- **Off-screen check**: coordinates outside the display bounds are refused, which
|
||
catches a Chrome window on a second monitor or partly off the edge.
|
||
- The element is scrolled to the centre of the viewport before measuring, so a
|
||
target below the fold is handled rather than mis-clicked.
|
||
|
||
## Known limits
|
||
|
||
- **Latency** is bounded by the extension's poll interval (default 3s), so a click
|
||
takes a few seconds to fire. Drop the interval in the extension popup if that
|
||
matters.
|
||
- **Multi-monitor**: coordinates come from `window.screenX/screenY`, which are
|
||
relative to the primary display's origin. A Chrome window on a secondary monitor
|
||
usually still works, but verify with `locate` before trusting `click`.
|
||
- **Display scaling**: the scale factor is inferred by comparing the screen width
|
||
the OS reports against the one the browser reports, and is only trusted when it
|
||
lands on a real scaling factor (1.0, 1.25, 1.5, …). Anything else falls back to
|
||
1:1 with a warning, because on a multi-monitor desktop the two sides describe
|
||
different displays and the ratio is meaningless. If clicks land at a consistent
|
||
offset, set `--scale` explicitly.
|
||
- **Focus is exclusive**: the browser must be frontmost at the moment of each
|
||
click, so the machine can't be used for anything else during a run, and a stray
|
||
dialog stealing focus fails the step.
|
||
- The measurement and the click are separate moments. If the page moves the element
|
||
in between (a re-render, a late-loading banner), the click lands where it *was*.
|