About our crawler
If you have arrived here from a line in your access log, this page tells you exactly what we were doing and how to make us stop.
Identity
MyNewsFeedBot/1.0 (+https://mynewsfeed.site/bot/)The crawler is not yet operating publicly. Our published IP ranges will be listed on this page before it is, so you can allowlist or block them by address as well as by user agent.
How it behaves
We identify ourselves honestly
Every request carries a user agent naming the crawler and linking to this page. We do not disguise ourselves as a browser, and we do not rotate identities to avoid being noticed.
We obey robots.txt
Fetched and cached per host, honoured both for our own user agent and for the wildcard. Crawl-delay is treated as a floor, never as a suggestion.
One budget per site, not one per customer
However many of our customers follow your site, it is fetched once and shared between them. A hundred subscribers does not become a hundred times the traffic.
We back off when asked
A 429 or a Retry-After is respected. Repeated errors slow us down, then stop us, and the customer is told their source is failing rather than us hammering you about it.
We ask for as little as possible
Conditional requests every time, so an unchanged page costs you a 304 and almost no bandwidth. Sites that go quiet are polled progressively less often.
We do not render your pages
No headless browser, no executing your JavaScript, no loading your assets. We read the document and nothing else.
We take the headline and the link
Not your article. We store the title, the URL, the timestamp and your name as the source, and every item on every page we host links back to you.
Add two lines to robots.txt
This is honoured within 24 hours, which is how long we cache your robots.txt for.
User-agent: MyNewsFeedBot
Disallow: /Stop us, and stop our customers
Blocking robots.txt stops us fetching you. An opt-out goes further: your domain is added to a global blocklist, existing pages stop showing your headlines, and no customer can add you again.
We need to know the request is really from the site owner, so we will ask you to prove it with a DNS TXT record or an email from an address at the domain. Once proven, it takes effect immediately and permanently.
Request an opt-outOr write to optout@mynewsfeed.site. We would rather hear from you than be blocked at the firewall.