Sitemap generation
Walk public board resources and serve a bounded sitemap index and XML buckets.
A
JGenerate the same eight-bucket sitemap model as Cavuno’s hosted board. The walker handles feature gating, pagination, canonical paths, and the minimum-content rule for job taxonomy pages.
Prerequisites
- A stateless public
boardclient. - A canonical public origin.
- Server routes for
/sitemap.xmland/sitemap/:filename. - A cache independent from candidate or employer sessions.
Build the sitemap service
Return these strings with Content-Type: application/xml; charset=utf-8. Return your framework’s 404 for an unknown filename or chunk.
Understand what is emitted
The buckets cover marketing routes, job categories, skills, locations and details, companies, salaries, and blog content. The blog bucket is removed when the feature is disabled. Other empty buckets produce valid empty URL sets.
Job taxonomy pages need at least five matching jobs before they are emitted. This avoids indexing thin category, skill, and location pages. Location counts include only jobs in the place itself, never jobs within a search distance, and distance views (?within=) are never emitted.
Expected result
/sitemap.xml is a sitemap index whose entries point to valid bucket files. Each bucket route returns a standards-compliant URL set, including a valid empty set when an enabled bucket has no URLs.
Verify the result
Confirm that:
- Every
<loc>uses the production HTTPS origin. - Disabled blog routes are absent.
- Job detail URLs include both company and job slugs.
- A non-existent bucket or chunk returns 404.
- XML reserved characters are escaped by the renderer.
Submit /sitemap.xml to your search-console tooling only after it works at the final public domain.
Errors and edge cases
- A catalog above the offset ceiling logs a truncation warning until bulk support is available. Monitor sitemap generation logs.
- Cross-axis salary pages and combined location/taxonomy pages are not currently emitted because the public API cannot enumerate them without an N+1 crawl.
- The walker can make many public reads. Do not run it during every ordinary page render.
- An origin with a path or the wrong environment creates invalid canonical URLs.
Production cautions
- Cache generated XML and revalidate often enough to remove expired jobs.
- Generate with a stateless client; never attach candidate authorization.
- Alert on truncation warnings and generation failures.
- Keep
robots.txt’sSitemap:line aligned withboard.seo().canonicalBase.
Next, format board-owned data with Styling and formatting.