# emprego.com # # The public job catalog is open to search engines and to AI search/citation # crawlers. Private areas (admin, accounts, dashboards, messages, …) are # protected by authentication and served noindex — not hidden by this file, so # they are deliberately not listed here. # Crawlers that collect content to TRAIN AI models are not permitted. This does # NOT reduce search/citation visibility: ChatGPT, Perplexity, Google and others # use separate search crawlers, which are allowed by the group below. # # This list must stay a SUPERSET of Cloudflare's managed robots.txt block, which # is currently prepended to this file at the edge. The moment that injection is # turned off (docs/robots-llms-cloudflare-followup.md section 1), whatever it # covered and this file does not is silently unblocked. Amazonbot and # CloudflareBrowserRenderingCrawler were exactly that gap on 2026-08-04 and are # listed here so the switch-off loses nothing. User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Bytespider User-agent: meta-externalagent User-agent: Amazonbot User-agent: CloudflareBrowserRenderingCrawler Disallow: / # Applebot is off the job and company trees, because it was taking the whole # site and it does not honour a crawl budget. # # Measured 2026-08-02 → 2026-08-04, every hour of every day: Applebot fetches a # steady ~23,000 HTML pages/hour — 60-72% of ALL the HTML emprego.com serves, # around the clock, never a spike. Googlebot over the same window takes ~180/h. # Human traffic is a few hundred per network. Those fetches are the long tail, # one URL each, so almost every one is an edge-cache miss and a full render # against D1's single reader — the load that leaves real visitors queueing. # # Crawl-delay was tried first and FAILED, measured rather than assumed: it # shipped 2026-08-04 22:18 UTC, Applebot re-read this file at 00:00 (HTTP 200, # the new copy), and its rate over the next two hours was 5,229-5,421 per 15 # minutes against 5,222-5,380 before. Not a drop, not even a downward drift. # Crawl-delay is not in RFC 9309; Disallow is, so it is the directive with an # actual chance of being honoured. Enforcement that does not depend on Applebot # co-operating at all is the WAF rule in scripts/cf-waf-challenge-applebot.mjs, # which challenges the same paths. # # WHY THESE PATHS. The jobs and companies trees are the whole cost: every # listing root, every facet under it, and every detail page. The homepage, # /about, /pricing, /guides, /insights and the sitemaps stay crawlable, so # Apple keeps an entry for the site rather than losing it entirely. Prefix # matching covers the facets and detail pages under each root. # # The list is generated from the app's own locale tables (JOBS_ROUTE_SLUGS in # @pinta/db, routeDict.companies in lib/routes.ts) and robots-txt.test.ts # asserts it still covers every locale — a new locale must not quietly open a # new door. # # Google IGNORES Crawl-delay and honours Disallow, so a Disallow here would # have cost Google's crawl if it named Google — it does not. Applebot-Extended # (AI training) stays fully disallowed above; different agent, different # question. # # A named group REPLACES `*` for that agent — robots.txt groups do not inherit — # so the /api rules below are repeated verbatim. Content-Signal is deliberately # not repeated: the served file must carry exactly one (see robots-txt.test.ts), # and the signal it states for `*` is already `search=yes`. User-agent: Applebot Crawl-delay: 2 Disallow: /ar/wazaif Disallow: /de/jobs Disallow: /es/empleos Disallow: /fr/emplois Disallow: /hi/jobs Disallow: /id/pekerjaan Disallow: /it/lavori Disallow: /jobs Disallow: /ms/pekerjaan Disallow: /ne/jobs Disallow: /nl/banen Disallow: /pt-BR/vagas Disallow: /pt-PT/vagas Disallow: /ta/jobs Disallow: /th/jobs Disallow: /sv/jobb Disallow: /nb/jobber Disallow: /fi/tyopaikat Disallow: /uk/robota Disallow: /zh/jobs Disallow: /ar/sharakat Disallow: /companies Disallow: /de/unternehmen Disallow: /es/empresas Disallow: /fr/entreprises Disallow: /hi/companies Disallow: /id/perusahaan Disallow: /it/aziende Disallow: /ms/syarikat Disallow: /ne/companies Disallow: /nl/bedrijven Disallow: /pt-BR/empresas Disallow: /pt-PT/empresas Disallow: /ta/companies Disallow: /th/companies Disallow: /sv/foretag Disallow: /nb/bedrifter Disallow: /fi/yritykset Disallow: /uk/kompanii Disallow: /zh/companies Disallow: /claim-organization Disallow: /*/claim-organization # Repeated from the `*` group below, where the reasoning is. Named groups do not # inherit — and Applebot is 60-72% of all HTML this site serves against # Googlebot's ~180/hour, so a rule that exists only in `*` misses the crawler # that matters most. /people is 60 real names, professions and cities per fetch. Disallow: /people$ Disallow: /people? Disallow: /es/personas$ Disallow: /es/personas? Disallow: /pt-BR/pessoas$ Disallow: /pt-BR/pessoas? Disallow: /pt-PT/pessoas$ Disallow: /pt-PT/pessoas? Disallow: /de/leute$ Disallow: /de/leute? Disallow: /nl/mensen$ Disallow: /nl/mensen? Disallow: /it/persone$ Disallow: /it/persone? Disallow: /fr/personnes$ Disallow: /fr/personnes? Disallow: /id/orang$ Disallow: /id/orang? Disallow: /ar/ashkhas$ Disallow: /ar/ashkhas? Disallow: /ms/orang$ Disallow: /ms/orang? Disallow: /zh/people$ Disallow: /zh/people? Disallow: /ta/people$ Disallow: /ta/people? Disallow: /uk/liudy$ Disallow: /uk/liudy? Disallow: /hi/people$ Disallow: /hi/people? Disallow: /ne/people$ Disallow: /ne/people? Disallow: /th/people$ Disallow: /th/people? Disallow: /sv/personer$ Disallow: /sv/personer? Disallow: /nb/personer$ Disallow: /nb/personer? Disallow: /fi/henkilot$ Disallow: /fi/henkilot? # Repeated from the `*` group below, where the reasoning is. Named groups do not # inherit, and this is the agent whose crawl rate made the cost worth blocking. Disallow: /*setLang= Allow: /api/agent/jobs Allow: /api/jobs/search Allow: /api/client-session Allow: /api/og/ Allow: /api/storage/organizations/ Allow: /api/storage/brands/ Disallow: /api/ User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no # The employer claim page is for one named employer, reached from an emailed # token or from the band on their own job pages. It is not a browsable surface # and it already renders noindex, so nothing here costs search visibility. # # Measured on a 90-second sample of live traffic, 2026-08-07 16:50 UTC: 123 # arrivals — ~5,000/hour — of which 89 carried a crawler user-agent (Applebot # 43, an assortment of smaller bots 42, Bingbot 2, Claude-SearchBot 3). None # were edge-cached: middleware caches only the anonymous allowlist, and this # route is deliberately off it, because pages/claim-organization/[orgId].astro # records the arrival during SSR and so must render every time. p90 was 935ms # of worker time, and an unclaimed org costs a two-query batch that reads every # job row and every application row the employer has. That landed on D1's # single reader through the afternoon it was already refusing queries — the # maintenance page went from 0% of real requests at 07:00 to 4.63% of the 16:00 # hour, peaking at 7.23% in one 15-minute window. # # Wildcard rather than one line per locale: unlike /jobs and /companies this # route keeps the same slug in every language, so only the prefix varies # (/hi/claim-organization/…, /ta/claim-organization/… were both in the sample). # `*` in a path is RFC 9309 §2.2.3 and is honoured by Google, Bing and Applebot. Disallow: /claim-organization Disallow: /*/claim-organization # The language switcher emits one ?setLang=true link per locale on every page, # and each href inherits the current query string — so /jobs?q=engenheiro& # location=lisboa yields 17 more crawlable URLs, and so does every other filter # and pagination permutation. Each is a 302 served no-store, and each # destination render is private/no-cache: an uncacheable worker render, forced # onto D1's PRIMARY reader whenever `q` is present, because replica FTS reads # hang. One crawled listing becomes 17 uncacheable primary-DB renders, on the # same reader whose saturation the notes above are about. # # The links stay in the HTML on purpose. They are how a human switches language # without losing their filters — middleware preserves the query across the # switch deliberately. They are simply not a crawl surface. Disallow: /*setLang= Allow: /api/agent/jobs Allow: /api/jobs/search # Ad configuration (publisher + slot ids) is delivered here rather than baked # into the HTML, because pages are edge-cached. Blocking it would make Google's # renderer see every page with no ad unit at all, while real users see ads. Allow: /api/client-session # Images that public pages point crawlers AT, rather than endpoints. `Disallow: # /api/` was blanket-blocking both, on 100% of job and company pages: # /api/og/ — the og:image on every job and company page. # /api/storage/organizations/ — employer logos. This is the path our own JSON-LD # names in hiringOrganization.logo and Organization.logo, so the markup # was pointing Google at a URL robots.txt forbade it to fetch. Costs the # logo in the Google Jobs card plus an "Image not crawlable or indexable" # warning, and the preview on every shared link (facebookexternalhit, # Twitterbot, LinkedInBot and Slackbot all honour robots.txt). # Scoped to organizations/ + brands/ on purpose — api/storage/[...path].ts gates # public serving on exactly those prefixes (brands/ carries brand logos, the # asset the brand pages' Organization.logo JSON-LD points crawlers at). users/ # and guest_cvs/ are auth-gated there and must stay disallowed; never widen # this to /api/storage/. Allow: /api/og/ Allow: /api/storage/organizations/ Allow: /api/storage/brands/ # /people is the candidate directory: real names, professions and cities of # members who opted in to being discoverable. It renders noindex, but noindex # stops INDEXING, not fetching — and the `*` group above states # `ai-input=yes`, so without this the page would be fetched and ingested by # every AI crawler the site invites. It is also linked from the footer of every # page in every locale, so discovery is immediate and in 20 variants. # # `$` and `?` rather than a bare prefix, and this is the whole point: the # directory moved to the HEAD of the /people tree in the 2026-08-25 rename, so # `Disallow: /people` would also close /people/ under it. Individual # profiles are deliberately reachable and indexable — sitemap-index.xml.ts # leaves people.xml out on purpose, "profiles stay reachable, just not # enumerated". Closing the list without closing the people on it needs the # end-anchor, and `?` catches the filtered forms the anchor alone would miss. # # Generated from routeDict.people (lib/routes.ts) and asserted by # robots-txt.test.ts, exactly like the jobs and companies trees above: a new # locale must not quietly open a new door onto people's names. Disallow: /people$ Disallow: /people? Disallow: /es/personas$ Disallow: /es/personas? Disallow: /pt-BR/pessoas$ Disallow: /pt-BR/pessoas? Disallow: /pt-PT/pessoas$ Disallow: /pt-PT/pessoas? Disallow: /de/leute$ Disallow: /de/leute? Disallow: /nl/mensen$ Disallow: /nl/mensen? Disallow: /it/persone$ Disallow: /it/persone? Disallow: /fr/personnes$ Disallow: /fr/personnes? Disallow: /id/orang$ Disallow: /id/orang? Disallow: /ar/ashkhas$ Disallow: /ar/ashkhas? Disallow: /ms/orang$ Disallow: /ms/orang? Disallow: /zh/people$ Disallow: /zh/people? Disallow: /ta/people$ Disallow: /ta/people? Disallow: /uk/liudy$ Disallow: /uk/liudy? Disallow: /hi/people$ Disallow: /hi/people? Disallow: /ne/people$ Disallow: /ne/people? Disallow: /th/people$ Disallow: /th/people? Disallow: /sv/personer$ Disallow: /sv/personer? Disallow: /nb/personer$ Disallow: /nb/personer? Disallow: /fi/henkilot$ Disallow: /fi/henkilot? Disallow: /api/ Sitemap: https://emprego.com/sitemap-index.xml