Lessons from building browser automation in production

Browser automation becomes a service only after jobs, resource limits, persistence, partial results, and recovery are designed around it.

browser automationPuppeteerqueuesreliability
Lessons from building browser automation in production illustration

A browser script can look reliable during a local run and still fail as soon as it serves concurrent users. Pages load at different speeds, sessions encounter challenges, browser processes leak memory, and user requests may end before a search finishes.

The operational unit should be a job with a lifecycle, not a request that holds a browser open until the last result appears.

Move long work behind a job boundary

Validate the search request, create a durable job record, return its identifier, and perform browser work asynchronously. The job should record status, progress, current query, results accumulated so far, and a clear stop or error reason.

The interface can then poll or subscribe to progress without depending on a long-lived HTTP connection. It can let the user cancel work and show partial results when an external page fails after some leads have already been collected.

Make browser ownership explicit

Launching an unrestricted browser per request wastes memory and makes load hard to predict. A browser pool can cap concurrency and reuse healthy instances. Each job should have clear ownership of its page and context, and cleanup should run in a finally path so errors do not strand resources.

Long-running instances still need health policy. Track page counts or memory signals, recycle browsers when necessary, and make restart behavior part of normal job completion rather than an emergency maintenance task.

Persist for recovery, not only reporting

If an application restarts while a job is active, the product should be able to identify unfinished work and decide whether to resume, mark it interrupted, or safely retry. Persist progress at useful intervals and make state transitions explicit. Keep result writes idempotent so replay does not duplicate leads.

Persistence should include enough information to explain the job but avoid retaining more personal data than the product needs. Apply cleanup and retention rules to files and intermediate results as deliberately as to database records.

Expect external sites to push back

Public websites change markup, throttle requests, return partial pages, and sometimes present CAPTCHAs. Detect those states and represent them in job results. Do not promise that a scraper can always get every result. A product that can return useful partial output and explain why work stopped is more honest and easier to operate.

Rate limits, delays, page limits, and user-provided proxies can be part of a retry strategy, but they need bounds. An unbounded retry turns a transient failure into a resource problem and can increase the impact on an external service.

Keep policy in the API

The backend should own authentication, account limits, team policy, and billing decisions. A browser client can render progress and request an export, but it should not decide whether a user has enough quota or which leads they are entitled to retain.

This separation also gives operational services a single point to enforce concurrency and rate limits across multiple clients.

Measure the job lifecycle

Record time spent waiting for a browser, navigation duration, result counts, challenge rates, retries, and failure categories. Avoid logging complete pages or sensitive user settings. These measurements help distinguish a provider change from a worker saturation problem without exposing unnecessary content.

Useful browser automation depends on this entire lifecycle. The scraping selector is the smallest part; the product needs resource ownership, persistence, feedback, and limits around it.