Screaming Frog

screamingfrog.co.ukcontributed by Samuele Ongaro

YES

Replaceable in one session with an AI coding agent.

A desktop crawler that walks a whole site and reports broken links, redirect chains, duplicate titles, missing metadata and everything else a technical audit needs.

Promptfree, for everyone, and the only version there is
Build me the site crawler I actually need instead of Screaming Frog: walk the whole site, and tell me what is broken before somebody else finds it.

STACK
- Node 20+ as a command-line tool, with a small server-rendered report
- SQLite through better-sqlite3 for the crawl store, WAL mode
- undici for fetching, a real HTML parser, and Playwright only for the pages that need JavaScript
- Runs locally or in CI; there is no reason for this to be a hosted service

THE DATA MODEL
- crawls: id, start_url, started_at, finished_at, pages, config_json
- pages: id, crawl_id, url, status, redirect_to, content_type, bytes, depth, found_at_url, title, title_length, meta_description, description_length, h1_json, canonical, robots_meta, lang, word_count, text_hash, load_ms, is_indexable, indexability_reason
- links: id, crawl_id, from_page_id, to_url, to_page_id, anchor_text, rel, is_internal, position
- resources: id, crawl_id, page_id, kind, url, status, bytes — images, scripts, styles
- images: page_id, url, alt, width_attr, height_attr, bytes, natural_width, natural_height
- issues: id, crawl_id, page_id, rule, severity, detail_json
- Every crawl kept, so two can be compared. The comparison is what turns a crawler into a monitor

CRAWLING POLITELY AND CORRECTLY
- Obey robots.txt for your own site too, because that is what a search engine will do and you want to see what it sees
- An identifiable user agent, a concurrency limit, a delay per host, and a crawl budget so a calendar with infinite pages does not consume the afternoon
- Normalise URLs consistently: resolve relative links, strip the fragment, decide about trailing slashes and case once and apply it everywhere. Inconsistent normalisation is what produces phantom duplicates
- Follow redirects but record every hop, because the chain is the finding
- Respect nofollow when configured, and record it either way
- Render with a browser only where the page is empty without it; note which pages needed it, because that list is itself a finding

WHAT TO CHECK — THE LIST IS THE PRODUCT
- Broken internal links, with the pages that link to them. This is the single most valuable output and the reason to run it weekly
- Broken external links, checked more gently and with the flakiness accounted for — a single timeout is not a broken link, so re-check before reporting
- Redirect chains and loops, with the number of hops
- Pages returning an error, and pages linked from the site but blocked by robots
- Duplicate and missing titles, descriptions and h1s; lengths outside the range that gets truncated
- Duplicate content by text hash, and near-duplicates
- Canonical problems: missing, self-referencing when it should not be, pointing at a redirect, pointing off-site, or conflicting with the sitemap
- Pages in the sitemap that are not linked from anywhere, and linked pages missing from the sitemap
- Orphan pages, reachable only from the sitemap
- Depth: pages more than a few clicks from the home page, which is a structural problem rather than a page problem
- Images without alt text, without dimensions, or served far larger than they are displayed
- Mixed content, insecure links, and links to the wrong protocol or hostname
- hreflang consistency if the site has it, which is almost always wrong somewhere
- Structured data parsed and validated per page
- Response times, and the slowest pages ranked

WHAT THE HOSTED TOOL DOES THAT THIS SHOULD NOT PRETEND TO
- The value of the original is a decade of accumulated checks, each one a lesson from a real site. Your list will be shorter
- So make it easy to add a rule: a rule is a small function over a parsed page and its links, with a severity and a message. Add one every time you find something the hard way, and the tool improves in the direction of your own sites

THE REPORT
- A summary of issues by severity, then each rule with the affected pages and the specific detail
- Exports as CSV and JSON so it can go into a spreadsheet or a ticket
- A diff against the previous crawl: what broke, what was fixed, what is new. That is the output somebody will actually read every week
- Run in CI on a schedule and fail the build on new issues above a severity, which is how a site stops rotting silently

OPERATIONS
- No server needed. A local database file per crawl, and a static report
- .env or flags: start URL, concurrency, delay, budget, user agent, whether to render
- Sitemaps and a URL list accepted as seeds as well as a crawl

WHAT MATTERS MOST
URL normalisation and the crawl-to-crawl diff. Get normalisation right or every report will be full of duplicates that do not exist, and build the diff early, because a list of four hundred issues is ignored while a list of six new ones gets fixed.

Give me the repository, the crawler, the rule set with each rule as its own small module, the report, the CI workflow, and a README explaining how to add a rule.

What you lose

  • A long list of checks accumulated over years, each one a lesson from a real site
  • Crawling at speed with JavaScript rendering when a site needs it
  • Integrations that pull analytics and search data into the crawl

If you would rather not build

  • A Playwright or Node crawler in CI

What it costs

read from their page 17 Aug 2026

PlanBilled monthlyBilled yearlyLast read
—$23.25/mo—17 Aug 2026

Their pricing page is where these came from. Seeing a different price? Tell us.

The escape hatch

open source · no votes, no paid placement

linkinator

$0

Crawls a site and reports broken links; runs happily in CI.

JustinBeckwith/linkinatorfree · open source

Lighthouse CI

$0

Runs Lighthouse audits on every build with budgets you set.

GoogleChrome/lighthouse-cifree · open source

Why this verdict

our own opinion · changed only by a person

80/100

Verdict yes at 80. A crawler and a dozen checks are a weekend, and running them in CI on every pull request beats a manual audit.

History

tracked since 10 Aug 2026 · nothing is ever overwritten

Interest · last 30 dayspeak 1/day
views0130 Aug4 Sept9 Sept14 Sept19 Sept24 Sept28 Sept
— views— prompt copies none yet— votes none yet

Questions about Screaming Frog

answered from the record above

Is Screaming Frog free?

No — the plan we track is $23.25 a month. $279 per licence per year, which is about $23.25/month; the free tier crawls up to 500 URLs.

Can you replace Screaming Frog by building your own?

YES. Replaceable in one session with an AI coding agent. Replacement score 80 out of 100, build time one session. Read what you lose before you decide.

How much does Screaming Frog cost?

$23.25 a month on Licence — $279 a year. Recorded 17 Aug 2026.

What do you lose by replacing Screaming Frog?

A long list of checks accumulated over years, each one a lesson from a real site; Crawling at speed with JavaScript rendering when a site needs it; Integrations that pull analytics and search data into the crawl. If any of those carry weight for you, keep paying.

Is there an open-source alternative to Screaming Frog?

Yes: linkinator, Lighthouse CI. The prompt on this page is for when you want it your way instead.

Related entries

same category first, most replaced first

All 36 in Analytics

Not sending yet

Every week, something stops being worth paying for.

New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.

free forever · no tracking pixel · stored here, never passed to anyone

Esc