API documentation · 7 of 9
Results
What the latest crawl found. These read the project's most recent crawl, so wait for it to finish or you get a partial picture. Before a crawl has recorded any pages they answer 404 no_crawl.
GET/projects/{id}/reportread
The overview: the health score (the share of pages with no critical issue), issues by severity, and top_issues, the ranked list of what to fix first.
{
"crawl_id": 3,
"health": {
"score": 68, "band": "fair",
"pages_total": 87, "pages_failed": 28, "pages_clean": 59,
"checks_total": 6873, "checks_failed": 41, "checks_passed": 6832
},
"critical": [{ "type": "ERROR_40x", "priority": 1, "count": 2 }],
"alert": [...], "warning": [...], "passed": [...],
"top_issues": [
{ "type": "ERROR_40x", "severity": "critical", "priority": 1, "count": 2, "percent": 2.3 }
]
}
band is excellent, good, fair or weak. percent is the share of pages an issue is on.
GET/projects/{id}/issuesread
Issue counts grouped into critical, alert, warning and passed. An empty group is [], never null.
{ "crawl_id": 3, "critical": [{ "type": "ERROR_40x", "priority": 1, "count": 2 }], "alert": [], "warning": [], "passed": [] }
GET/projects/{id}/issues/{type}read
The pages one issue was found on, paginated with ?page=. {type} is an issue type from the calls above. evidence is the value that failed the check, when there is one to quote.
{
"crawl_id": 3,
"issue_type": "ERROR_40x",
"pager": { "page": 1, "total_pages": 1 },
"pages": [{ "id": 5610, "url": "https://example.com/gone", "status_code": 404, "title": "", "evidence": "404", ... }]
}
GET/projects/{id}/pagesread
Every URL the crawl recorded, paginated with ?page=. Add ?term= to filter by URL.
{
"crawl_id": 3,
"pager": { "page": 1, "total_pages": 9 },
"pages": [{
"id": 5501, "url": "https://example.com/", "status_code": 200,
"media_type": "text/html", "title": "Home", "depth": 0,
"crawled": true, "indexable": true, "in_sitemap": true, "blocked_by_robots": false
}]
}
GET/projects/{id}/pages/{rid}read
Everything recorded about one page: title, description, robots, canonical, headings, word count, size, time to first byte, keywords, hreflang and its issues. {rid} is a page id from the list above.
Its links and resources come one kind at a time with ?tab=, paginated with ?page=:
tab | Returns |
|---|---|
internal, external | internal_links / external_links: url, rel, text, nofollow and status code. |
inlinks | The pages that link to this one. |
redirections | The pages that redirect to this one. |
images | Images with their alt text. |
scripts, styles, iframes, audios | Their URLs. |
videos | Videos with their poster image. |
structured_data | The JSON-LD, Microdata and RDFa on the page: format, types, and for the types Google has rich results for, errors (missing required properties) and warnings (missing recommended ones). parse_error is set when a block could not be read. |
curl "http://localhost:9000/api/v1/projects/1/pages/5501?tab=internal&page=2" -H "Authorization: Bearer $KEY"
GET/projects/{id}/statsread
The figures the dashboard charts are drawn from.
{
"crawl_id": 3,
"media_types": [{ "key": "text/html", "value": 80 }],
"status_codes": [{ "key": "200", "value": 85 }, { "key": "404", "value": 2 }],
"canonical": { "canonical": 80, "non_canonical": 7 },
"scheme": { "http": 0, "https": 87 },
"image_alt": { "with_alt": 120, "without_alt": 14 },
"status_by_depth": [{ "depth": 0, "status_2xx": 1, "status_4xx": 0, ... }]
}
GET/projects/{id}/pagespeedread
What Google PageSpeed Insights measured for the front page, mobile and desktop: the score, field data from real Chrome users and lab data from one Lighthouse run. A metric with vital: true is a Core Web Vital. Either is null when the installation has no PageSpeed key.
{
"crawl_id": 3,
"mobile": {
"strategy": "mobile", "score": 64, "verdict": "average",
"field": [{ "key": "LCP", "value": 2400, "display": "2.4 s", "rating": "average", "vital": true }],
"lab": [...], "fetched": "2026-09-29T12:35:00Z"
},
"desktop": null
}
GET/projects/{id}/sitemapread
The sitemap files and their state, how many of their URLs the crawl reached, and the ones it never did, paginated with ?page=. enabled is the project's crawl_sitemap setting.
{
"crawl_id": 3, "enabled": true,
"url_total": 500, "reached": 431, "unreached": 69,
"files": [{ "url": "https://example.com/sitemap.xml", "source": "robots.txt", "status_code": 200, "url_count": 500, ... }],
"pager": { "page": 1, "total_pages": 2 },
"unreached_urls": [{ "sitemap": "https://example.com/sitemap.xml", "url": "https://example.com/orphan", "last_mod": "2026-01-30" }]
}