Files
Brandon LiandClaude Opus 5 3cc7ddcc5c Add scrollToLoad and an absent variant of waitFor
scrollToLoad walks a progressively-loading list to the bottom before the steps
that act on its items run. Stopping is two-part: no new matches appeared AND the
container was already pinned to the bottom — counting alone stops early on a slow
fetch. Hitting the scroll cap is reported rather than passed off as done, so a
later step never works quietly on a partial list.

The scrolling element is usually not the window. Lists like this live in a div
with its own overflow, and scrolling the document does nothing at all, so the
step walks up from a matched item to the ancestor that actually scrolls —
overflow allows it and there is more content than fits — with containerSelector
to name one outright when the guess is wrong. Verified against a page whose
document also scrolls, which is the case that tells the two apart: it found the
inner div and pulled 12 items up to 60 in 7 scrolls.

waitFor gains `absent`, for waiting on something to go rather than arrive — a
modal closing after a reset. It only accepts a genuine "selector matched
nothing"; an unreachable extension looks the same from a distance and would
otherwise satisfy the gate for the wrong reason, sending the next iteration into
a page that still has the modal open.

The locate queue carries a free-form options blob now, so a new kind of request
stops meaning a new column each time.

Also fixes a missing comma in the reset flow that broke the build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-30 14:31:36 -05:00
..

Clicker

Moves the real desktop mouse and clicks an element you name with a CSS selector.

The extension can see the DOM but can't move the mouse; this script can move the mouse but can't see inside Chrome. They meet at the dashboard API:

clicker    --POST /api/autobuyer/locate------>  "where is button.buy?"
extension  --POST /api/autobuyer/locate/claim   takes the request
extension                                       raises the Chrome window, brings
                                                the tab to the front, scrolls the
                                                element to centre, measures it
extension  --POST /api/autobuyer/locate/result  desktop x,y
clicker    --GET  /api/autobuyer/locate?id=--->  reads the answer
clicker                                          moves the mouse, clicks

Two ways to run it

clicker.py — one action at a time, from the shell. Use this to find selectors and confirm coordinates before wiring anything up.

runner.py — the daemon behind the dashboard's buttons. Leave it running; it polls for queued runs and executes the steps.

python runner.py                 # then press a button on the AutoBuyer page
python runner.py --dry-run       # walks the steps, never presses or types

Automations are defined in lib/automations.ts, and the steps are sent to the runner by the server — so adding a button means editing that file, and nothing in the runner or the page changes. Each run reports step-by-step progress back to the dashboard, and the page's Stop button halts a run between steps.

Only one run executes at a time: two processes driving the same physical mouse would interleave clicks.

Setup

pip install -r requirements.txt

The macOS-only backends are gated behind platform markers, so this installs the right set on either OS.

macOS

Grant Accessibility permission to whatever runs the script — Terminal, iTerm, VS Code — under System Settings → Privacy & Security → Accessibility. Without it pyautogui moves nothing and fails silently, which looks identical to a bad selector.

Windows

No extra permissions, but two things differ.

Display scaling. The runner declares itself DPI-aware at startup, and prints which mode it got (per-monitor DPI aware). That stops Windows reporting a virtualised screen size and rescaling the coordinates it accepts, which would otherwise put clicks progressively further off as you move from the top-left.

Foreground lock. Windows refuses to raise a window for a process that doesn't already own the foreground — it flashes the taskbar instead, and the activation call returns as if it worked. The runner verifies by reading the foreground window's process, so a failure is reported rather than clicked into. If runs stop with "the browser is not frontmost", click the Chrome window once by hand and try again; something else is holding the foreground.

Browser windows are matched on the owning process (chrome.exe, msedge.exe, brave.exe), not the window title, so an editor with chrome.js open won't be mistaken for the browser.

First run on a new machine

Verify in this order — each step is harmless on its own, and the first one that looks wrong tells you where the problem is:

# 1. Can the extension see the page, and do the coordinates look sane?
python clicker.py locate "a.add_account_btn"

# 2. Does the cursor actually land on the element? Nothing is pressed.
python clicker.py click "a.add_account_btn" --dry-run

# 3. A real click.
python clicker.py click "a.add_account_btn"

At step 1, check the reported desktop x,y against where the element actually is on screen. A consistent offset means display scaling was misdetected — pass --scale explicitly. Only then start the daemon.

Linux / X11

Not supported. pygetwindow, which raises the browser window, has no X11 backend, so every click is refused with "window management is unsupported on linux". Making it work needs an xdotool or wmctrl path in focus.py, a real window manager (Xvfb alone has no focus semantics), and X11 rather than Wayland.

Use

The AutoBuyer switch on the dashboard is the master arm: the script refuses to run while it's off.

# Measure only — no mouse movement. Start here.
python clicker.py locate "button.buy"

# Move the cursor to the target but don't press.
python clicker.py click "button.buy" --dry-run

# Actually click.
python clicker.py click "button.buy"

# Type into a field: clicks it to place the caret, then types.
python clicker.py type 'input[name="quantity"]' --text "5000" --clear

# Keep the text out of shell history.
echo "5000" | python clicker.py type 'input#qty' --stdin --clear

# Pin it to a specific tab, and pick the 3rd match.
python clicker.py click ".trade-btn" --index 2 --url "https://tradeify.co/*"
Flag Meaning
--url Chrome match pattern for the tab. Without it, the hosts in the extension manifest.
--open-url Page to open if no tab matches --url.
--index Which match, when the selector hits several (default 0).
--api Dashboard URL (default http://localhost:3000).
--timeout Seconds to wait for the extension (default 45 — a cold page load takes time).
--scale CSS-to-desktop pixel ratio. Auto-detected; override if clicks land off.
--dry-run Move the cursor, don't press.
--force Click even when something is covering the element.
--robotic Straight-line move and instant click, skipping the motion model.
--seed Seed the motion RNG so a run replays identically (debugging).
--no-activate Don't raise the browser first. The click may then be swallowed.
--text Text to type (action type).
--stdin Read the text from stdin instead, keeping it out of shell history.
--clear Select-all and delete before typing, instead of appending at the caret.
--allow-enter Permit newlines. Each one presses Enter, which may submit the form.
--no-verify Skip reading the field back after typing.

Typing

type clicks the field to place the caret, then types with an uneven human cadence (45130ms between keys, a beat after each space, an occasional longer pause). Real OS-level keystrokes are also the correct way to fill a React form: setting .value directly is ignored by controlled components, while genuine key events are not.

Four guards, all of which catch silent failures:

  • Non-ASCII is rejected. pyautogui.write() has no keycode for é or £ and skips them without complaint, which would leave a quietly truncated value in the field. Better to refuse than to submit caf where you meant café.
  • Newlines are rejected unless --allow-enter, because Enter may submit the form — and on this site that could mean placing an order.
  • Non-editable targets are refused — disabled, read-only, or simply not an input. The keystrokes would go nowhere.
  • The field is read back afterwards and compared against what was typed (--no-verify skips it). A field that never took focus, or that ignored the input, otherwise looks exactly like success. Exit code 4 means the text did not land.

--clear selects-all and deletes first. Without it the text is inserted at wherever the caret landed, which for a field with existing contents usually produces something like 50005000.

Typos are deliberately not simulated. A mistyped digit in a trading form is a real loss, and the backspace-and-correct step is exactly the part that can go wrong — a field with input masking or autocomplete can swallow the correction and leave the wrong number behind.

Don't pass credentials via --text: it lands in your shell history and in the process list. --stdin avoids both, but nothing here is built to handle secrets safely.

Window focus

A click on a window that isn't focused is consumed by the window manager activating that window — it never reaches the control underneath. That's why an automated click against a background Chrome appears to do nothing the first time and work the second: the first click only brought Chrome forward.

The extension calls chrome.windows.update({focused: true}), but that only orders windows within Chrome. If the frontmost application is your terminal — which it is, since that's where you launched this — Chrome is still in the background.

So focus.py raises the browser application itself immediately before the press, then confirms it actually came forward before committing to the click. If it can't verify, it refuses rather than firing a click that would be eaten (exit code 3).

On macOS this uses NSWorkspace via pyobjc, which needs no Automation permission — activating an app is not scripting it. Without pyobjc it falls back to osascript, which does prompt for Automation permission the first time.

On Windows it goes through pygetwindow, then confirms by reading the foreground window's process. That confirmation is the important half: Windows silently declines to raise a window for a background process, and the activation call reports success either way.

Cursor motion

humanize.py moves the pointer the way a hand does rather than teleporting:

  • a curved (cubic Bézier) path instead of a straight line, bowing to one side
  • eased velocity — accelerate out, coast, decelerate in
  • sub-pixel tremor that decays near the target, so the landing stays exact
  • long throws (>260px) sometimes overshoot slightly and pull back
  • a 60170ms dwell after arriving, before the press
  • the button held down 55120ms rather than an instant down/up

This is about reliability as much as appearance. Plenty of web controls only arm once they have actually been hovered — dropdowns, tooltip-gated buttons, custom widgets — and a cursor that arrives and presses in the same tick can outrun the page's own mousemove handlers. The dwell is what lets those catch up.

Note pyautogui.PAUSE is set to 0 on import: it otherwise sleeps 0.1s after every call, which would add tens of seconds across a stepped path.

Exit codes: 0 ok, 1 error, 2 capture switch off, 3 refused to click (covered element, or coordinates off-screen).

Safety

  • Failsafe: slam the cursor into a screen corner to abort mid-run.
  • Covered elements: before reporting, the extension checks document.elementFromPoint() at the target centre. If a cookie banner or modal is on top, the click is refused rather than sent into the overlay — --force overrides.
  • Off-screen check: coordinates outside the display bounds are refused, which catches a Chrome window on a second monitor or partly off the edge.
  • The element is scrolled to the centre of the viewport before measuring, so a target below the fold is handled rather than mis-clicked.

Known limits

  • Latency is bounded by the extension's poll interval (default 3s), so a click takes a few seconds to fire. Drop the interval in the extension popup if that matters.
  • Multi-monitor: coordinates come from window.screenX/screenY, which are relative to the primary display's origin. A Chrome window on a secondary monitor usually still works, but verify with locate before trusting click.
  • Display scaling: the scale factor is inferred by comparing the screen width the OS reports against the one the browser reports, and is only trusted when it lands on a real scaling factor (1.0, 1.25, 1.5, …). Anything else falls back to 1:1 with a warning, because on a multi-monitor desktop the two sides describe different displays and the ratio is meaningless. If clicks land at a consistent offset, set --scale explicitly.
  • Focus is exclusive: the browser must be frontmost at the moment of each click, so the machine can't be used for anything else during a run, and a stray dialog stealing focus fails the step.
  • The measurement and the click are separate moments. If the page moves the element in between (a re-render, a late-loading banner), the click lands where it was.