API documentation · 5 of 9
Project settings
The fields POST /projects and
PATCH /projects/{id} take. Every
project response returns them under settings, with url and
user_agent at the top level.
| Field | Type | Default | What it does |
|---|---|---|---|
url | string | required | Where the crawl starts. An http or https address. |
crawl_sitemap | bool | false | Also crawl the URLs listed in the site's sitemaps, not only the ones linked from its pages. |
check_external_links | bool | false | Request links to other sites to find the broken ones. |
allow_subdomains | bool | false | Follow links onto subdomains, such as blog.example.com. |
ignore_robots_txt | bool | false | Crawl the pages robots.txt disallows. |
follow_nofollow | bool | false | Follow links marked rel="nofollow". |
include_noindex | bool | false | Include pages marked noindex in the report. |
archive | bool | false | Keep a WACZ archive of every response, to replay and download. Turning it off deletes the stored archive. |
basic_auth | bool | false | The site sits behind HTTP basic authentication. Starting its crawl then needs credentials. |
render_javascript | bool | false | Render every HTML page in a headless browser before checking it, for sites that build their content with JavaScript. Links, headings, structured data and custom rules then see what a browser sees. Slower. |
user_agent | string | the default | The user agent the crawler sends. On PATCH, "" resets it. |
Strict on purpose
A field the API does not know is refused with 400 invalid_body rather
than ignored, so a typo such as crawlSitemap is reported instead of
silently dropped. A wrong type is refused the same way. There is no page
limit to set: every crawl fetches the whole site.