Bingbot is the automated program Microsoft uses to visit web pages, download them and pass them back for indexing in Bing. If Bingbot cannot reach a page, that page cannot appear in Bing results or in the products that rely on Bing’s index.
How Bingbot works
Like other crawlers, Bingbot starts from pages it already knows about, follows the links on them, and reads sitemaps and URL submissions to find new ones. This process is called crawling. Before fetching anything on your site it checks your robots.txt file for rules aimed at it, either under its own name or under the general “all crawlers” group.
In your server logs Bingbot identifies itself with a user agent string containing “bingbot/2.0”, and it uses both a desktop and a mobile version. It renders pages with a modern browser engine, so it can process JavaScript, although content that appears only after heavy client-side scripting is still slower and less reliable to index than content in the initial HTML.
One practical difference from Googlebot: Bingbot respects the crawl-delay directive in robots.txt, which asks it to wait a set number of seconds between requests. Google ignores that line. Bing Webmaster Tools also has a crawl control setting that lets you shift Bingbot’s activity to quieter hours.
Why it matters
Bingbot is the only way into Bing’s index, and that index feeds more than bing.com, as the entry on Bing Webmaster Tools explains. Blocking Bingbot, even by accident, removes you from every product built on it. This happens more often than you might expect, usually through a firewall, a security plugin or a bot-management service that challenges unfamiliar crawlers and was only tested against Google.
Crawl behaviour is also a useful diagnostic. If Bingbot is spending its visits on filter URLs, old campaign pages or endless calendar links, the same waste is probably affecting other crawlers. For a small UK business site this rarely causes a problem; for a large shop with thousands of product and filter combinations it can mean new products take far longer to appear.
Common mistakes
- Blocking anything that claims to be Bingbot without checking whether it really is. Plenty of scrapers fake the user agent; the genuine crawler can be confirmed.
- Setting a long crawl-delay “to be safe”, which slows how quickly Bing finds new and updated pages.
- Writing robots.txt rules for Googlebot only, then adding a blanket disallow for every other crawler.
- Serving Bingbot an error or a cookie-consent wall that hides the page content, so the page is indexed with almost nothing on it.
How to act on it
To confirm a visitor really is Bingbot, run a reverse DNS lookup on its IP address. A genuine request resolves to a hostname ending in search.msn.com, and a forward lookup on that hostname returns the same IP. Bing Webmaster Tools also has a tool that checks an IP for you.
Then look at your robots.txt, firewall and any security or CDN settings to make sure Bingbot is allowed through. A short log file analysis shows which URLs it is actually requesting and whether it is getting errors. If it is wasting visits on low-value URLs, tidy up the internal links and parameters that generate them. Checking crawler access across Google and Bing is a standard part of my technical SEO work.
