On a per-GB product, every byte of a response is billed, whether or not you read it. A page that weighs several megabytes when you need a few kilobytes of it is a cost problem before it is a speed problem.
Measure first
curl -s -o /dev/null -x "$PROXY_URL" \
-w 'download=%{size_download} header=%{size_header}\n' \
https://example.com/Plain HTTP clients move only what the server sends for that URL. A headless browser also fetches every script, stylesheet, image and font the page references, and any that the scripts then request.
Block what you do not need in a browser
await page.route('**/*', (route) => {
const type = route.request().resourceType();
return ['image', 'media', 'font'].includes(type) ? route.abort() : route.continue();
});- Blocking images, media and fonts is usually the largest saving and rarely breaks extraction.
- Blocking scripts can break pages that render with JavaScript. Test it against your content check.
- Request compressed responses, which most clients do by default, and check that yours does.
- Prefer an API endpoint a page calls over rendering the page, where the terms of the site allow it.
Record bytes per success before and after a change. The cost tool turns that figure and the price you are quoted into a cost per thousand successes.