API documentation · 7 of 9

Results

What the latest crawl found. These read the project's most recent crawl, so wait for it to finish or you get a partial picture. Before a crawl has recorded any pages they answer 404 no_crawl.

GET/projects/{id}/reportread

The overview: the health score (the share of pages with no critical issue), issues by severity, and top_issues, the ranked list of what to fix first.

{
  "crawl_id": 3,
  "health": {
    "score": 68, "band": "fair",
    "pages_total": 87, "pages_failed": 28, "pages_clean": 59,
    "checks_total": 6873, "checks_failed": 41, "checks_passed": 6832
  },
  "critical": [{ "type": "ERROR_40x", "priority": 1, "count": 2 }],
  "alert": [...], "warning": [...], "passed": [...],
  "top_issues": [
    { "type": "ERROR_40x", "severity": "critical", "priority": 1, "count": 2, "percent": 2.3 }
  ]
}

band is excellent, good, fair or weak. percent is the share of pages an issue is on.

GET/projects/{id}/issuesread

Issue counts grouped into critical, alert, warning and passed. An empty group is [], never null.

{ "crawl_id": 3, "critical": [{ "type": "ERROR_40x", "priority": 1, "count": 2 }], "alert": [], "warning": [], "passed": [] }

GET/projects/{id}/issues/{type}read

The pages one issue was found on, paginated with ?page=. {type} is an issue type from the calls above. evidence is the value that failed the check, when there is one to quote.

{
  "crawl_id": 3,
  "issue_type": "ERROR_40x",
  "pager": { "page": 1, "total_pages": 1 },
  "pages": [{ "id": 5610, "url": "https://example.com/gone", "status_code": 404, "title": "", "evidence": "404", ... }]
}

GET/projects/{id}/pagesread

Every URL the crawl recorded, paginated with ?page=. Add ?term= to filter by URL.

{
  "crawl_id": 3,
  "pager": { "page": 1, "total_pages": 9 },
  "pages": [{
    "id": 5501, "url": "https://example.com/", "status_code": 200,
    "media_type": "text/html", "title": "Home", "depth": 0,
    "crawled": true, "indexable": true, "in_sitemap": true, "blocked_by_robots": false
  }]
}

GET/projects/{id}/pages/{rid}read

Everything recorded about one page: title, description, robots, canonical, headings, word count, size, time to first byte, keywords, hreflang and its issues. {rid} is a page id from the list above.

Its links and resources come one kind at a time with ?tab=, paginated with ?page=:

tabReturns
internal, externalinternal_links / external_links: url, rel, text, nofollow and status code.
inlinksThe pages that link to this one.
redirectionsThe pages that redirect to this one.
imagesImages with their alt text.
scripts, styles, iframes, audiosTheir URLs.
videosVideos with their poster image.
structured_dataThe JSON-LD, Microdata and RDFa on the page: format, types, and for the types Google has rich results for, errors (missing required properties) and warnings (missing recommended ones). parse_error is set when a block could not be read.
curl "http://localhost:9000/api/v1/projects/1/pages/5501?tab=internal&page=2" -H "Authorization: Bearer $KEY"

GET/projects/{id}/statsread

The figures the dashboard charts are drawn from.

{
  "crawl_id": 3,
  "media_types": [{ "key": "text/html", "value": 80 }],
  "status_codes": [{ "key": "200", "value": 85 }, { "key": "404", "value": 2 }],
  "canonical": { "canonical": 80, "non_canonical": 7 },
  "scheme": { "http": 0, "https": 87 },
  "image_alt": { "with_alt": 120, "without_alt": 14 },
  "status_by_depth": [{ "depth": 0, "status_2xx": 1, "status_4xx": 0, ... }]
}

GET/projects/{id}/pagespeedread

What Google PageSpeed Insights measured for the front page, mobile and desktop: the score, field data from real Chrome users and lab data from one Lighthouse run. A metric with vital: true is a Core Web Vital. Either is null when the installation has no PageSpeed key.

{
  "crawl_id": 3,
  "mobile": {
    "strategy": "mobile", "score": 64, "verdict": "average",
    "field": [{ "key": "LCP", "value": 2400, "display": "2.4 s", "rating": "average", "vital": true }],
    "lab": [...], "fetched": "2026-09-29T12:35:00Z"
  },
  "desktop": null
}

GET/projects/{id}/sitemapread

The sitemap files and their state, how many of their URLs the crawl reached, and the ones it never did, paginated with ?page=. enabled is the project's crawl_sitemap setting.

{
  "crawl_id": 3, "enabled": true,
  "url_total": 500, "reached": 431, "unreached": 69,
  "files": [{ "url": "https://example.com/sitemap.xml", "source": "robots.txt", "status_code": 200, "url_count": 500, ... }],
  "pager": { "page": 1, "total_pages": 2 },
  "unreached_urls": [{ "sitemap": "https://example.com/sitemap.xml", "url": "https://example.com/orphan", "last_mod": "2026-01-30" }]
}